N Neurarch Architectures Models Checks Data Docs Open the app

Models / NemotronH_Nano_Omni_Reasoning_V3

Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4

Reconstructed from its own config.json with no weights read. 746K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
3.08B
3,079,759,616 parameters
In the published checkpoint
18.33B
18,326,275,008 scalars · safetensors.total, read 2026-08-24
Delta
-83.2%

quantized This checkpoint is stored quantized (modelopt). The tensor count in the file counts stored elements under a packing scheme, not logical parameters, so the two numbers below are not measuring the same thing in either direction.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
210
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$379336.61
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsfits

Structure

212 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 262144
2EmbeddingEmbedding1 × 262144 × 2688
3Positional_EmbeddingLearned Pos Embed1 × 262144 × 2688
4Attention_1Multi-Head Attention1 × 262144 × 2688
5Add_1Add1 × 262144 × 2688
6LayerNorm_1_1LayerNorm1 × 262144 × 2688
7FFN_1Feed Forward1 × 262144 × 2688
8Attention_2Multi-Head Attention1 × 262144 × 2688
9Add_2Add1 × 262144 × 2688
10LayerNorm_2_1LayerNorm1 × 262144 × 2688
11FFN_2Feed Forward1 × 262144 × 2688
12Attention_3Multi-Head Attention1 × 262144 × 2688
13Add_3Add1 × 262144 × 2688
14LayerNorm_3_1LayerNorm1 × 262144 × 2688
15FFN_3Feed Forward1 × 262144 × 2688
16Attention_4Multi-Head Attention1 × 262144 × 2688
17Add_4Add1 × 262144 × 2688
18LayerNorm_4_1LayerNorm1 × 262144 × 2688
19FFN_4Feed Forward1 × 262144 × 2688
20Attention_5Multi-Head Attention1 × 262144 × 2688
21Add_5Add1 × 262144 × 2688
22LayerNorm_5_1LayerNorm1 × 262144 × 2688
23FFN_5Feed Forward1 × 262144 × 2688
24Attention_6Multi-Head Attention1 × 262144 × 2688
25Add_6Add1 × 262144 × 2688
26LayerNorm_6_1LayerNorm1 × 262144 × 2688
27FFN_6Feed Forward1 × 262144 × 2688
28Attention_7Multi-Head Attention1 × 262144 × 2688
29Add_7Add1 × 262144 × 2688
30LayerNorm_7_1LayerNorm1 × 262144 × 2688
31FFN_7Feed Forward1 × 262144 × 2688
32Attention_8Multi-Head Attention1 × 262144 × 2688
33Add_8Add1 × 262144 × 2688
34LayerNorm_8_1LayerNorm1 × 262144 × 2688
35FFN_8Feed Forward1 × 262144 × 2688
36Attention_9Multi-Head Attention1 × 262144 × 2688
37Add_9Add1 × 262144 × 2688
38LayerNorm_9_1LayerNorm1 × 262144 × 2688
39FFN_9Feed Forward1 × 262144 × 2688
40Attention_10Multi-Head Attention1 × 262144 × 2688
41Add_10Add1 × 262144 × 2688
42LayerNorm_10_1LayerNorm1 × 262144 × 2688
43FFN_10Feed Forward1 × 262144 × 2688
44Attention_11Multi-Head Attention1 × 262144 × 2688
45Add_11Add1 × 262144 × 2688
46LayerNorm_11_1LayerNorm1 × 262144 × 2688
47FFN_11Feed Forward1 × 262144 × 2688
48Attention_12Multi-Head Attention1 × 262144 × 2688
49Add_12Add1 × 262144 × 2688
50LayerNorm_12_1LayerNorm1 × 262144 × 2688
51FFN_12Feed Forward1 × 262144 × 2688
52Attention_13Multi-Head Attention1 × 262144 × 2688
53Add_13Add1 × 262144 × 2688
54LayerNorm_13_1LayerNorm1 × 262144 × 2688
55FFN_13Feed Forward1 × 262144 × 2688
56Attention_14Multi-Head Attention1 × 262144 × 2688
57Add_14Add1 × 262144 × 2688
58LayerNorm_14_1LayerNorm1 × 262144 × 2688
59FFN_14Feed Forward1 × 262144 × 2688
60Attention_15Multi-Head Attention1 × 262144 × 2688
61Add_15Add1 × 262144 × 2688
62LayerNorm_15_1LayerNorm1 × 262144 × 2688
63FFN_15Feed Forward1 × 262144 × 2688
64Attention_16Multi-Head Attention1 × 262144 × 2688
65Add_16Add1 × 262144 × 2688
66LayerNorm_16_1LayerNorm1 × 262144 × 2688
67FFN_16Feed Forward1 × 262144 × 2688
68Attention_17Multi-Head Attention1 × 262144 × 2688
69Add_17Add1 × 262144 × 2688
70LayerNorm_17_1LayerNorm1 × 262144 × 2688
71FFN_17Feed Forward1 × 262144 × 2688
72Attention_18Multi-Head Attention1 × 262144 × 2688
73Add_18Add1 × 262144 × 2688
74LayerNorm_18_1LayerNorm1 × 262144 × 2688
75FFN_18Feed Forward1 × 262144 × 2688
76Attention_19Multi-Head Attention1 × 262144 × 2688
77Add_19Add1 × 262144 × 2688
78LayerNorm_19_1LayerNorm1 × 262144 × 2688
79FFN_19Feed Forward1 × 262144 × 2688
80Attention_20Multi-Head Attention1 × 262144 × 2688
81Add_20Add1 × 262144 × 2688
82LayerNorm_20_1LayerNorm1 × 262144 × 2688
83FFN_20Feed Forward1 × 262144 × 2688
84Attention_21Multi-Head Attention1 × 262144 × 2688
85Add_21Add1 × 262144 × 2688
86LayerNorm_21_1LayerNorm1 × 262144 × 2688
87FFN_21Feed Forward1 × 262144 × 2688
88Attention_22Multi-Head Attention1 × 262144 × 2688
89Add_22Add1 × 262144 × 2688
90LayerNorm_22_1LayerNorm1 × 262144 × 2688
91FFN_22Feed Forward1 × 262144 × 2688
92Attention_23Multi-Head Attention1 × 262144 × 2688
93Add_23Add1 × 262144 × 2688
94LayerNorm_23_1LayerNorm1 × 262144 × 2688
95FFN_23Feed Forward1 × 262144 × 2688
96Attention_24Multi-Head Attention1 × 262144 × 2688
97Add_24Add1 × 262144 × 2688
98LayerNorm_24_1LayerNorm1 × 262144 × 2688
99FFN_24Feed Forward1 × 262144 × 2688
100Attention_25Multi-Head Attention1 × 262144 × 2688
101Add_25Add1 × 262144 × 2688
102LayerNorm_25_1LayerNorm1 × 262144 × 2688
103FFN_25Feed Forward1 × 262144 × 2688
104Attention_26Multi-Head Attention1 × 262144 × 2688
105Add_26Add1 × 262144 × 2688
106LayerNorm_26_1LayerNorm1 × 262144 × 2688
107FFN_26Feed Forward1 × 262144 × 2688
108Attention_27Multi-Head Attention1 × 262144 × 2688
109Add_27Add1 × 262144 × 2688
110LayerNorm_27_1LayerNorm1 × 262144 × 2688
111FFN_27Feed Forward1 × 262144 × 2688
112Attention_28Multi-Head Attention1 × 262144 × 2688
113Add_28Add1 × 262144 × 2688
114LayerNorm_28_1LayerNorm1 × 262144 × 2688
115FFN_28Feed Forward1 × 262144 × 2688
116Attention_29Multi-Head Attention1 × 262144 × 2688
117Add_29Add1 × 262144 × 2688
118LayerNorm_29_1LayerNorm1 × 262144 × 2688
119FFN_29Feed Forward1 × 262144 × 2688
120Attention_30Multi-Head Attention1 × 262144 × 2688
121Add_30Add1 × 262144 × 2688
122LayerNorm_30_1LayerNorm1 × 262144 × 2688
123FFN_30Feed Forward1 × 262144 × 2688
124Attention_31Multi-Head Attention1 × 262144 × 2688
125Add_31Add1 × 262144 × 2688
126LayerNorm_31_1LayerNorm1 × 262144 × 2688
127FFN_31Feed Forward1 × 262144 × 2688
128Attention_32Multi-Head Attention1 × 262144 × 2688
129Add_32Add1 × 262144 × 2688
130LayerNorm_32_1LayerNorm1 × 262144 × 2688
131FFN_32Feed Forward1 × 262144 × 2688
132Attention_33Multi-Head Attention1 × 262144 × 2688
133Add_33Add1 × 262144 × 2688
134LayerNorm_33_1LayerNorm1 × 262144 × 2688
135FFN_33Feed Forward1 × 262144 × 2688
136Attention_34Multi-Head Attention1 × 262144 × 2688
137Add_34Add1 × 262144 × 2688
138LayerNorm_34_1LayerNorm1 × 262144 × 2688
139FFN_34Feed Forward1 × 262144 × 2688
140Attention_35Multi-Head Attention1 × 262144 × 2688
141Add_35Add1 × 262144 × 2688
142LayerNorm_35_1LayerNorm1 × 262144 × 2688
143FFN_35Feed Forward1 × 262144 × 2688
144Attention_36Multi-Head Attention1 × 262144 × 2688
145Add_36Add1 × 262144 × 2688
146LayerNorm_36_1LayerNorm1 × 262144 × 2688
147FFN_36Feed Forward1 × 262144 × 2688
148Attention_37Multi-Head Attention1 × 262144 × 2688
149Add_37Add1 × 262144 × 2688
150LayerNorm_37_1LayerNorm1 × 262144 × 2688
151FFN_37Feed Forward1 × 262144 × 2688
152Attention_38Multi-Head Attention1 × 262144 × 2688
153Add_38Add1 × 262144 × 2688
154LayerNorm_38_1LayerNorm1 × 262144 × 2688
155FFN_38Feed Forward1 × 262144 × 2688
156Attention_39Multi-Head Attention1 × 262144 × 2688
157Add_39Add1 × 262144 × 2688
158LayerNorm_39_1LayerNorm1 × 262144 × 2688
159FFN_39Feed Forward1 × 262144 × 2688
160Attention_40Multi-Head Attention1 × 262144 × 2688
161Add_40Add1 × 262144 × 2688
162LayerNorm_40_1LayerNorm1 × 262144 × 2688
163FFN_40Feed Forward1 × 262144 × 2688
164Attention_41Multi-Head Attention1 × 262144 × 2688
165Add_41Add1 × 262144 × 2688
166LayerNorm_41_1LayerNorm1 × 262144 × 2688
167FFN_41Feed Forward1 × 262144 × 2688
168Attention_42Multi-Head Attention1 × 262144 × 2688
169Add_42Add1 × 262144 × 2688
170LayerNorm_42_1LayerNorm1 × 262144 × 2688
171FFN_42Feed Forward1 × 262144 × 2688
172Attention_43Multi-Head Attention1 × 262144 × 2688
173Add_43Add1 × 262144 × 2688
174LayerNorm_43_1LayerNorm1 × 262144 × 2688
175FFN_43Feed Forward1 × 262144 × 2688
176Attention_44Multi-Head Attention1 × 262144 × 2688
177Add_44Add1 × 262144 × 2688
178LayerNorm_44_1LayerNorm1 × 262144 × 2688
179FFN_44Feed Forward1 × 262144 × 2688
180Attention_45Multi-Head Attention1 × 262144 × 2688
181Add_45Add1 × 262144 × 2688
182LayerNorm_45_1LayerNorm1 × 262144 × 2688
183FFN_45Feed Forward1 × 262144 × 2688
184Attention_46Multi-Head Attention1 × 262144 × 2688
185Add_46Add1 × 262144 × 2688
186LayerNorm_46_1LayerNorm1 × 262144 × 2688
187FFN_46Feed Forward1 × 262144 × 2688
188Attention_47Multi-Head Attention1 × 262144 × 2688
189Add_47Add1 × 262144 × 2688
190LayerNorm_47_1LayerNorm1 × 262144 × 2688
191FFN_47Feed Forward1 × 262144 × 2688
192Attention_48Multi-Head Attention1 × 262144 × 2688
193Add_48Add1 × 262144 × 2688
194LayerNorm_48_1LayerNorm1 × 262144 × 2688
195FFN_48Feed Forward1 × 262144 × 2688
196Attention_49Multi-Head Attention1 × 262144 × 2688
197Add_49Add1 × 262144 × 2688
198LayerNorm_49_1LayerNorm1 × 262144 × 2688
199FFN_49Feed Forward1 × 262144 × 2688
200Attention_50Multi-Head Attention1 × 262144 × 2688
201Add_50Add1 × 262144 × 2688
202LayerNorm_50_1LayerNorm1 × 262144 × 2688
203FFN_50Feed Forward1 × 262144 × 2688
204Attention_51Multi-Head Attention1 × 262144 × 2688
205Add_51Add1 × 262144 × 2688
206LayerNorm_51_1LayerNorm1 × 262144 × 2688
207FFN_51Feed Forward1 × 262144 × 2688
208Attention_52Multi-Head Attention1 × 262144 × 2688
209Add_52Add1 × 262144 × 2688
210LayerNorm_52_1LayerNorm1 × 262144 × 2688
211FFN_52Feed Forward1 × 262144 × 2688
212OutputOutput1 × 262144 × 2688

What the verifier says

info52 attention layers at embedDim 2688 cache full per-head K/V: about 546 KB per token at fp16, which dominates memory at long context. Grouped-query attention (e.g. 8:1) would cut this ~8×; multi-head latent attention (MLA) shrinks it ~10× or more. This is the move production LLMs make; it does not change the parameter count. Fix: Switch attention to groupedQueryAttention (set numKVHeads below numHeads, e.g. numHeads/4) or mla (a low-rank cached latent).
full-mha-serving-cost
infoAt 52 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init
warnAcross 52 attention layers this design caches 546 KB per token, so a single 8,192-token sequence needs ~4.6 GB of KV cache before weights or activations. That exceeds the 4 GB budget this rule assumes for serving headroom. Fix: Cut KV width: raise the GQA ratio (fewer numKVHeads), switch to MLA, reduce depth or embedDim, or accept a shorter serving context.
kv-cache-context-budget

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4 --plan --share

Other NemotronH_Nano_Omni_Reasoning_V3 checkpoints

Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
3.08B derived · -90.7% against the checkpoint
Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8
3.08B derived · -90.7% against the checkpoint