N Neurarch Architectures Models Checks Data Docs Open the app

Models / phimoe

Phi-tiny-MoE-instruct

Reconstructed from its own config.json with no weights read. 145K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
3.75B
3,754,724,672 parameters
In the published checkpoint
3.76B
3,755,220,288 scalars · safetensors.total, read 2025-12-10
Delta
-0.01%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `configuration_slimmoe.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
196
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$218.62
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsfits

Structure

198 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 4096
2EmbeddingEmbedding1 × 4096 × 4096
3RoPERoPE1 × 4096 × 4096
4RMSNorm_1_1RMSNorm1 × 4096 × 4096
5Attention_1Grouped Query Attn1 × 4096 × 4096
6Add_1_attnAdd1 × 4096 × 4096
7RMSNorm_1_2RMSNorm1 × 4096 × 4096
8MoE_1MoE Layer1 × 4096 × 4096
9Add_1_ffnAdd1 × 4096 × 4096
10RMSNorm_2_1RMSNorm1 × 4096 × 4096
11Attention_2Grouped Query Attn1 × 4096 × 4096
12Add_2_attnAdd1 × 4096 × 4096
13RMSNorm_2_2RMSNorm1 × 4096 × 4096
14MoE_2MoE Layer1 × 4096 × 4096
15Add_2_ffnAdd1 × 4096 × 4096
16RMSNorm_3_1RMSNorm1 × 4096 × 4096
17Attention_3Grouped Query Attn1 × 4096 × 4096
18Add_3_attnAdd1 × 4096 × 4096
19RMSNorm_3_2RMSNorm1 × 4096 × 4096
20MoE_3MoE Layer1 × 4096 × 4096
21Add_3_ffnAdd1 × 4096 × 4096
22RMSNorm_4_1RMSNorm1 × 4096 × 4096
23Attention_4Grouped Query Attn1 × 4096 × 4096
24Add_4_attnAdd1 × 4096 × 4096
25RMSNorm_4_2RMSNorm1 × 4096 × 4096
26MoE_4MoE Layer1 × 4096 × 4096
27Add_4_ffnAdd1 × 4096 × 4096
28RMSNorm_5_1RMSNorm1 × 4096 × 4096
29Attention_5Grouped Query Attn1 × 4096 × 4096
30Add_5_attnAdd1 × 4096 × 4096
31RMSNorm_5_2RMSNorm1 × 4096 × 4096
32MoE_5MoE Layer1 × 4096 × 4096
33Add_5_ffnAdd1 × 4096 × 4096
34RMSNorm_6_1RMSNorm1 × 4096 × 4096
35Attention_6Grouped Query Attn1 × 4096 × 4096
36Add_6_attnAdd1 × 4096 × 4096
37RMSNorm_6_2RMSNorm1 × 4096 × 4096
38MoE_6MoE Layer1 × 4096 × 4096
39Add_6_ffnAdd1 × 4096 × 4096
40RMSNorm_7_1RMSNorm1 × 4096 × 4096
41Attention_7Grouped Query Attn1 × 4096 × 4096
42Add_7_attnAdd1 × 4096 × 4096
43RMSNorm_7_2RMSNorm1 × 4096 × 4096
44MoE_7MoE Layer1 × 4096 × 4096
45Add_7_ffnAdd1 × 4096 × 4096
46RMSNorm_8_1RMSNorm1 × 4096 × 4096
47Attention_8Grouped Query Attn1 × 4096 × 4096
48Add_8_attnAdd1 × 4096 × 4096
49RMSNorm_8_2RMSNorm1 × 4096 × 4096
50MoE_8MoE Layer1 × 4096 × 4096
51Add_8_ffnAdd1 × 4096 × 4096
52RMSNorm_9_1RMSNorm1 × 4096 × 4096
53Attention_9Grouped Query Attn1 × 4096 × 4096
54Add_9_attnAdd1 × 4096 × 4096
55RMSNorm_9_2RMSNorm1 × 4096 × 4096
56MoE_9MoE Layer1 × 4096 × 4096
57Add_9_ffnAdd1 × 4096 × 4096
58RMSNorm_10_1RMSNorm1 × 4096 × 4096
59Attention_10Grouped Query Attn1 × 4096 × 4096
60Add_10_attnAdd1 × 4096 × 4096
61RMSNorm_10_2RMSNorm1 × 4096 × 4096
62MoE_10MoE Layer1 × 4096 × 4096
63Add_10_ffnAdd1 × 4096 × 4096
64RMSNorm_11_1RMSNorm1 × 4096 × 4096
65Attention_11Grouped Query Attn1 × 4096 × 4096
66Add_11_attnAdd1 × 4096 × 4096
67RMSNorm_11_2RMSNorm1 × 4096 × 4096
68MoE_11MoE Layer1 × 4096 × 4096
69Add_11_ffnAdd1 × 4096 × 4096
70RMSNorm_12_1RMSNorm1 × 4096 × 4096
71Attention_12Grouped Query Attn1 × 4096 × 4096
72Add_12_attnAdd1 × 4096 × 4096
73RMSNorm_12_2RMSNorm1 × 4096 × 4096
74MoE_12MoE Layer1 × 4096 × 4096
75Add_12_ffnAdd1 × 4096 × 4096
76RMSNorm_13_1RMSNorm1 × 4096 × 4096
77Attention_13Grouped Query Attn1 × 4096 × 4096
78Add_13_attnAdd1 × 4096 × 4096
79RMSNorm_13_2RMSNorm1 × 4096 × 4096
80MoE_13MoE Layer1 × 4096 × 4096
81Add_13_ffnAdd1 × 4096 × 4096
82RMSNorm_14_1RMSNorm1 × 4096 × 4096
83Attention_14Grouped Query Attn1 × 4096 × 4096
84Add_14_attnAdd1 × 4096 × 4096
85RMSNorm_14_2RMSNorm1 × 4096 × 4096
86MoE_14MoE Layer1 × 4096 × 4096
87Add_14_ffnAdd1 × 4096 × 4096
88RMSNorm_15_1RMSNorm1 × 4096 × 4096
89Attention_15Grouped Query Attn1 × 4096 × 4096
90Add_15_attnAdd1 × 4096 × 4096
91RMSNorm_15_2RMSNorm1 × 4096 × 4096
92MoE_15MoE Layer1 × 4096 × 4096
93Add_15_ffnAdd1 × 4096 × 4096
94RMSNorm_16_1RMSNorm1 × 4096 × 4096
95Attention_16Grouped Query Attn1 × 4096 × 4096
96Add_16_attnAdd1 × 4096 × 4096
97RMSNorm_16_2RMSNorm1 × 4096 × 4096
98MoE_16MoE Layer1 × 4096 × 4096
99Add_16_ffnAdd1 × 4096 × 4096
100RMSNorm_17_1RMSNorm1 × 4096 × 4096
101Attention_17Grouped Query Attn1 × 4096 × 4096
102Add_17_attnAdd1 × 4096 × 4096
103RMSNorm_17_2RMSNorm1 × 4096 × 4096
104MoE_17MoE Layer1 × 4096 × 4096
105Add_17_ffnAdd1 × 4096 × 4096
106RMSNorm_18_1RMSNorm1 × 4096 × 4096
107Attention_18Grouped Query Attn1 × 4096 × 4096
108Add_18_attnAdd1 × 4096 × 4096
109RMSNorm_18_2RMSNorm1 × 4096 × 4096
110MoE_18MoE Layer1 × 4096 × 4096
111Add_18_ffnAdd1 × 4096 × 4096
112RMSNorm_19_1RMSNorm1 × 4096 × 4096
113Attention_19Grouped Query Attn1 × 4096 × 4096
114Add_19_attnAdd1 × 4096 × 4096
115RMSNorm_19_2RMSNorm1 × 4096 × 4096
116MoE_19MoE Layer1 × 4096 × 4096
117Add_19_ffnAdd1 × 4096 × 4096
118RMSNorm_20_1RMSNorm1 × 4096 × 4096
119Attention_20Grouped Query Attn1 × 4096 × 4096
120Add_20_attnAdd1 × 4096 × 4096
121RMSNorm_20_2RMSNorm1 × 4096 × 4096
122MoE_20MoE Layer1 × 4096 × 4096
123Add_20_ffnAdd1 × 4096 × 4096
124RMSNorm_21_1RMSNorm1 × 4096 × 4096
125Attention_21Grouped Query Attn1 × 4096 × 4096
126Add_21_attnAdd1 × 4096 × 4096
127RMSNorm_21_2RMSNorm1 × 4096 × 4096
128MoE_21MoE Layer1 × 4096 × 4096
129Add_21_ffnAdd1 × 4096 × 4096
130RMSNorm_22_1RMSNorm1 × 4096 × 4096
131Attention_22Grouped Query Attn1 × 4096 × 4096
132Add_22_attnAdd1 × 4096 × 4096
133RMSNorm_22_2RMSNorm1 × 4096 × 4096
134MoE_22MoE Layer1 × 4096 × 4096
135Add_22_ffnAdd1 × 4096 × 4096
136RMSNorm_23_1RMSNorm1 × 4096 × 4096
137Attention_23Grouped Query Attn1 × 4096 × 4096
138Add_23_attnAdd1 × 4096 × 4096
139RMSNorm_23_2RMSNorm1 × 4096 × 4096
140MoE_23MoE Layer1 × 4096 × 4096
141Add_23_ffnAdd1 × 4096 × 4096
142RMSNorm_24_1RMSNorm1 × 4096 × 4096
143Attention_24Grouped Query Attn1 × 4096 × 4096
144Add_24_attnAdd1 × 4096 × 4096
145RMSNorm_24_2RMSNorm1 × 4096 × 4096
146MoE_24MoE Layer1 × 4096 × 4096
147Add_24_ffnAdd1 × 4096 × 4096
148RMSNorm_25_1RMSNorm1 × 4096 × 4096
149Attention_25Grouped Query Attn1 × 4096 × 4096
150Add_25_attnAdd1 × 4096 × 4096
151RMSNorm_25_2RMSNorm1 × 4096 × 4096
152MoE_25MoE Layer1 × 4096 × 4096
153Add_25_ffnAdd1 × 4096 × 4096
154RMSNorm_26_1RMSNorm1 × 4096 × 4096
155Attention_26Grouped Query Attn1 × 4096 × 4096
156Add_26_attnAdd1 × 4096 × 4096
157RMSNorm_26_2RMSNorm1 × 4096 × 4096
158MoE_26MoE Layer1 × 4096 × 4096
159Add_26_ffnAdd1 × 4096 × 4096
160RMSNorm_27_1RMSNorm1 × 4096 × 4096
161Attention_27Grouped Query Attn1 × 4096 × 4096
162Add_27_attnAdd1 × 4096 × 4096
163RMSNorm_27_2RMSNorm1 × 4096 × 4096
164MoE_27MoE Layer1 × 4096 × 4096
165Add_27_ffnAdd1 × 4096 × 4096
166RMSNorm_28_1RMSNorm1 × 4096 × 4096
167Attention_28Grouped Query Attn1 × 4096 × 4096
168Add_28_attnAdd1 × 4096 × 4096
169RMSNorm_28_2RMSNorm1 × 4096 × 4096
170MoE_28MoE Layer1 × 4096 × 4096
171Add_28_ffnAdd1 × 4096 × 4096
172RMSNorm_29_1RMSNorm1 × 4096 × 4096
173Attention_29Grouped Query Attn1 × 4096 × 4096
174Add_29_attnAdd1 × 4096 × 4096
175RMSNorm_29_2RMSNorm1 × 4096 × 4096
176MoE_29MoE Layer1 × 4096 × 4096
177Add_29_ffnAdd1 × 4096 × 4096
178RMSNorm_30_1RMSNorm1 × 4096 × 4096
179Attention_30Grouped Query Attn1 × 4096 × 4096
180Add_30_attnAdd1 × 4096 × 4096
181RMSNorm_30_2RMSNorm1 × 4096 × 4096
182MoE_30MoE Layer1 × 4096 × 4096
183Add_30_ffnAdd1 × 4096 × 4096
184RMSNorm_31_1RMSNorm1 × 4096 × 4096
185Attention_31Grouped Query Attn1 × 4096 × 4096
186Add_31_attnAdd1 × 4096 × 4096
187RMSNorm_31_2RMSNorm1 × 4096 × 4096
188MoE_31MoE Layer1 × 4096 × 4096
189Add_31_ffnAdd1 × 4096 × 4096
190RMSNorm_32_1RMSNorm1 × 4096 × 4096
191Attention_32Grouped Query Attn1 × 4096 × 4096
192Add_32_attnAdd1 × 4096 × 4096
193RMSNorm_32_2RMSNorm1 × 4096 × 4096
194MoE_32MoE Layer1 × 4096 × 4096
195Add_32_ffnAdd1 × 4096 × 4096
196Final_RMSNormRMSNorm1 × 4096 × 4096
197LM_HeadLinear1 × 4096 × 32064
198OutputOutput1 × 4096 × 32064

What the verifier says

infoMoE layers require an auxiliary router z-loss + load-balance loss during training to prevent expert collapse. This is not visible in the architecture diagram but must be in the training loop. Applies to all 32: MoE_1, MoE_2, MoE_3, MoE_4, MoE_5, MoE_6, MoE_7, MoE_8, MoE_9, MoE_10, MoE_11, MoE_12, MoE_13, MoE_14, MoE_15, MoE_16, MoE_17, MoE_18, MoE_19, MoE_20, MoE_21, MoE_22, MoE_23, MoE_24, MoE_25, MoE_26, MoE_27, MoE_28, MoE_29, MoE_30, MoE_31, MoE_32. Fix: Add a note on these layers. Typical aux_loss coefficient: 1e-2 (Mixtral/Switch Transformer).
moe-no-aux-loss
infoAt 32 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace microsoft/Phi-tiny-MoE-instruct --plan --share

Other phimoe checkpoints

Phi-3.5-MoE-instruct
41.87B derived · -0.00% against the checkpoint