N Neurarch Architectures Models Checks Data Docs Open the app

Models / whisper

Whisper-Hindi2Hinglish-Prime

Reconstructed from its own config.json with no weights read. 254K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
1.54B
1,535,882,240 parameters
In the published checkpoint
1.54B
1,543,490,560 scalars · safetensors.total, read 2025-11-06
Delta
-0.49%

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
233
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$12.04
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsfits
H100 (80GB)weights + activationsfits

Structure

236 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 512
2MelSpectrogramMel Spectrogram1 × 80 × 1
3Enc_Conv1Conv1D1 × 1280 × 1
4Enc_Conv2Conv1D1 × 1280 × 1
5Enc_ToTokensPermute1 × 1 × 1280
6Enc_PosEncPositional Encoding1 × 1 × 1280
7Enc_Attn_1Multi-Head Attention1 × 1 × 1280
8Enc_Norm_1LayerNorm1 × 1 × 1280
9Enc_FFN_1Feed Forward1 × 1 × 1280
10Enc_Attn_2Multi-Head Attention1 × 1 × 1280
11Enc_Norm_2LayerNorm1 × 1 × 1280
12Enc_FFN_2Feed Forward1 × 1 × 1280
13Enc_Attn_3Multi-Head Attention1 × 1 × 1280
14Enc_Norm_3LayerNorm1 × 1 × 1280
15Enc_FFN_3Feed Forward1 × 1 × 1280
16Enc_Attn_4Multi-Head Attention1 × 1 × 1280
17Enc_Norm_4LayerNorm1 × 1 × 1280
18Enc_FFN_4Feed Forward1 × 1 × 1280
19Enc_Attn_5Multi-Head Attention1 × 1 × 1280
20Enc_Norm_5LayerNorm1 × 1 × 1280
21Enc_FFN_5Feed Forward1 × 1 × 1280
22Enc_Attn_6Multi-Head Attention1 × 1 × 1280
23Enc_Norm_6LayerNorm1 × 1 × 1280
24Enc_FFN_6Feed Forward1 × 1 × 1280
25Enc_Attn_7Multi-Head Attention1 × 1 × 1280
26Enc_Norm_7LayerNorm1 × 1 × 1280
27Enc_FFN_7Feed Forward1 × 1 × 1280
28Enc_Attn_8Multi-Head Attention1 × 1 × 1280
29Enc_Norm_8LayerNorm1 × 1 × 1280
30Enc_FFN_8Feed Forward1 × 1 × 1280
31Enc_Attn_9Multi-Head Attention1 × 1 × 1280
32Enc_Norm_9LayerNorm1 × 1 × 1280
33Enc_FFN_9Feed Forward1 × 1 × 1280
34Enc_Attn_10Multi-Head Attention1 × 1 × 1280
35Enc_Norm_10LayerNorm1 × 1 × 1280
36Enc_FFN_10Feed Forward1 × 1 × 1280
37Enc_Attn_11Multi-Head Attention1 × 1 × 1280
38Enc_Norm_11LayerNorm1 × 1 × 1280
39Enc_FFN_11Feed Forward1 × 1 × 1280
40Enc_Attn_12Multi-Head Attention1 × 1 × 1280
41Enc_Norm_12LayerNorm1 × 1 × 1280
42Enc_FFN_12Feed Forward1 × 1 × 1280
43Enc_Attn_13Multi-Head Attention1 × 1 × 1280
44Enc_Norm_13LayerNorm1 × 1 × 1280
45Enc_FFN_13Feed Forward1 × 1 × 1280
46Enc_Attn_14Multi-Head Attention1 × 1 × 1280
47Enc_Norm_14LayerNorm1 × 1 × 1280
48Enc_FFN_14Feed Forward1 × 1 × 1280
49Enc_Attn_15Multi-Head Attention1 × 1 × 1280
50Enc_Norm_15LayerNorm1 × 1 × 1280
51Enc_FFN_15Feed Forward1 × 1 × 1280
52Enc_Attn_16Multi-Head Attention1 × 1 × 1280
53Enc_Norm_16LayerNorm1 × 1 × 1280
54Enc_FFN_16Feed Forward1 × 1 × 1280
55Enc_Attn_17Multi-Head Attention1 × 1 × 1280
56Enc_Norm_17LayerNorm1 × 1 × 1280
57Enc_FFN_17Feed Forward1 × 1 × 1280
58Enc_Attn_18Multi-Head Attention1 × 1 × 1280
59Enc_Norm_18LayerNorm1 × 1 × 1280
60Enc_FFN_18Feed Forward1 × 1 × 1280
61Enc_Attn_19Multi-Head Attention1 × 1 × 1280
62Enc_Norm_19LayerNorm1 × 1 × 1280
63Enc_FFN_19Feed Forward1 × 1 × 1280
64Enc_Attn_20Multi-Head Attention1 × 1 × 1280
65Enc_Norm_20LayerNorm1 × 1 × 1280
66Enc_FFN_20Feed Forward1 × 1 × 1280
67Enc_Attn_21Multi-Head Attention1 × 1 × 1280
68Enc_Norm_21LayerNorm1 × 1 × 1280
69Enc_FFN_21Feed Forward1 × 1 × 1280
70Enc_Attn_22Multi-Head Attention1 × 1 × 1280
71Enc_Norm_22LayerNorm1 × 1 × 1280
72Enc_FFN_22Feed Forward1 × 1 × 1280
73Enc_Attn_23Multi-Head Attention1 × 1 × 1280
74Enc_Norm_23LayerNorm1 × 1 × 1280
75Enc_FFN_23Feed Forward1 × 1 × 1280
76Enc_Attn_24Multi-Head Attention1 × 1 × 1280
77Enc_Norm_24LayerNorm1 × 1 × 1280
78Enc_FFN_24Feed Forward1 × 1 × 1280
79Enc_Attn_25Multi-Head Attention1 × 1 × 1280
80Enc_Norm_25LayerNorm1 × 1 × 1280
81Enc_FFN_25Feed Forward1 × 1 × 1280
82Enc_Attn_26Multi-Head Attention1 × 1 × 1280
83Enc_Norm_26LayerNorm1 × 1 × 1280
84Enc_FFN_26Feed Forward1 × 1 × 1280
85Enc_Attn_27Multi-Head Attention1 × 1 × 1280
86Enc_Norm_27LayerNorm1 × 1 × 1280
87Enc_FFN_27Feed Forward1 × 1 × 1280
88Enc_Attn_28Multi-Head Attention1 × 1 × 1280
89Enc_Norm_28LayerNorm1 × 1 × 1280
90Enc_FFN_28Feed Forward1 × 1 × 1280
91Enc_Attn_29Multi-Head Attention1 × 1 × 1280
92Enc_Norm_29LayerNorm1 × 1 × 1280
93Enc_FFN_29Feed Forward1 × 1 × 1280
94Enc_Attn_30Multi-Head Attention1 × 1 × 1280
95Enc_Norm_30LayerNorm1 × 1 × 1280
96Enc_FFN_30Feed Forward1 × 1 × 1280
97Enc_Attn_31Multi-Head Attention1 × 1 × 1280
98Enc_Norm_31LayerNorm1 × 1 × 1280
99Enc_FFN_31Feed Forward1 × 1 × 1280
100Enc_Attn_32Multi-Head Attention1 × 1 × 1280
101Enc_Norm_32LayerNorm1 × 1 × 1280
102Enc_FFN_32Feed Forward1 × 1 × 1280
103decoder_tokensInput1 × 448
104Dec_Token_EmbeddingEmbedding1 × 448 × 1280
105Dec_PosEmbeddingLearned Pos Embed1 × 448 × 1280
106Dec_SelfAttn_1Causal Attention1 × 448 × 1280
107Dec_CrossAttn_1Cross-Attention1 × 448 × 1280
108Dec_Norm_1LayerNorm1 × 448 × 1280
109Dec_FFN_1Feed Forward1 × 448 × 1280
110Dec_SelfAttn_2Causal Attention1 × 448 × 1280
111Dec_CrossAttn_2Cross-Attention1 × 448 × 1280
112Dec_Norm_2LayerNorm1 × 448 × 1280
113Dec_FFN_2Feed Forward1 × 448 × 1280
114Dec_SelfAttn_3Causal Attention1 × 448 × 1280
115Dec_CrossAttn_3Cross-Attention1 × 448 × 1280
116Dec_Norm_3LayerNorm1 × 448 × 1280
117Dec_FFN_3Feed Forward1 × 448 × 1280
118Dec_SelfAttn_4Causal Attention1 × 448 × 1280
119Dec_CrossAttn_4Cross-Attention1 × 448 × 1280
120Dec_Norm_4LayerNorm1 × 448 × 1280
121Dec_FFN_4Feed Forward1 × 448 × 1280
122Dec_SelfAttn_5Causal Attention1 × 448 × 1280
123Dec_CrossAttn_5Cross-Attention1 × 448 × 1280
124Dec_Norm_5LayerNorm1 × 448 × 1280
125Dec_FFN_5Feed Forward1 × 448 × 1280
126Dec_SelfAttn_6Causal Attention1 × 448 × 1280
127Dec_CrossAttn_6Cross-Attention1 × 448 × 1280
128Dec_Norm_6LayerNorm1 × 448 × 1280
129Dec_FFN_6Feed Forward1 × 448 × 1280
130Dec_SelfAttn_7Causal Attention1 × 448 × 1280
131Dec_CrossAttn_7Cross-Attention1 × 448 × 1280
132Dec_Norm_7LayerNorm1 × 448 × 1280
133Dec_FFN_7Feed Forward1 × 448 × 1280
134Dec_SelfAttn_8Causal Attention1 × 448 × 1280
135Dec_CrossAttn_8Cross-Attention1 × 448 × 1280
136Dec_Norm_8LayerNorm1 × 448 × 1280
137Dec_FFN_8Feed Forward1 × 448 × 1280
138Dec_SelfAttn_9Causal Attention1 × 448 × 1280
139Dec_CrossAttn_9Cross-Attention1 × 448 × 1280
140Dec_Norm_9LayerNorm1 × 448 × 1280
141Dec_FFN_9Feed Forward1 × 448 × 1280
142Dec_SelfAttn_10Causal Attention1 × 448 × 1280
143Dec_CrossAttn_10Cross-Attention1 × 448 × 1280
144Dec_Norm_10LayerNorm1 × 448 × 1280
145Dec_FFN_10Feed Forward1 × 448 × 1280
146Dec_SelfAttn_11Causal Attention1 × 448 × 1280
147Dec_CrossAttn_11Cross-Attention1 × 448 × 1280
148Dec_Norm_11LayerNorm1 × 448 × 1280
149Dec_FFN_11Feed Forward1 × 448 × 1280
150Dec_SelfAttn_12Causal Attention1 × 448 × 1280
151Dec_CrossAttn_12Cross-Attention1 × 448 × 1280
152Dec_Norm_12LayerNorm1 × 448 × 1280
153Dec_FFN_12Feed Forward1 × 448 × 1280
154Dec_SelfAttn_13Causal Attention1 × 448 × 1280
155Dec_CrossAttn_13Cross-Attention1 × 448 × 1280
156Dec_Norm_13LayerNorm1 × 448 × 1280
157Dec_FFN_13Feed Forward1 × 448 × 1280
158Dec_SelfAttn_14Causal Attention1 × 448 × 1280
159Dec_CrossAttn_14Cross-Attention1 × 448 × 1280
160Dec_Norm_14LayerNorm1 × 448 × 1280
161Dec_FFN_14Feed Forward1 × 448 × 1280
162Dec_SelfAttn_15Causal Attention1 × 448 × 1280
163Dec_CrossAttn_15Cross-Attention1 × 448 × 1280
164Dec_Norm_15LayerNorm1 × 448 × 1280
165Dec_FFN_15Feed Forward1 × 448 × 1280
166Dec_SelfAttn_16Causal Attention1 × 448 × 1280
167Dec_CrossAttn_16Cross-Attention1 × 448 × 1280
168Dec_Norm_16LayerNorm1 × 448 × 1280
169Dec_FFN_16Feed Forward1 × 448 × 1280
170Dec_SelfAttn_17Causal Attention1 × 448 × 1280
171Dec_CrossAttn_17Cross-Attention1 × 448 × 1280
172Dec_Norm_17LayerNorm1 × 448 × 1280
173Dec_FFN_17Feed Forward1 × 448 × 1280
174Dec_SelfAttn_18Causal Attention1 × 448 × 1280
175Dec_CrossAttn_18Cross-Attention1 × 448 × 1280
176Dec_Norm_18LayerNorm1 × 448 × 1280
177Dec_FFN_18Feed Forward1 × 448 × 1280
178Dec_SelfAttn_19Causal Attention1 × 448 × 1280
179Dec_CrossAttn_19Cross-Attention1 × 448 × 1280
180Dec_Norm_19LayerNorm1 × 448 × 1280
181Dec_FFN_19Feed Forward1 × 448 × 1280
182Dec_SelfAttn_20Causal Attention1 × 448 × 1280
183Dec_CrossAttn_20Cross-Attention1 × 448 × 1280
184Dec_Norm_20LayerNorm1 × 448 × 1280
185Dec_FFN_20Feed Forward1 × 448 × 1280
186Dec_SelfAttn_21Causal Attention1 × 448 × 1280
187Dec_CrossAttn_21Cross-Attention1 × 448 × 1280
188Dec_Norm_21LayerNorm1 × 448 × 1280
189Dec_FFN_21Feed Forward1 × 448 × 1280
190Dec_SelfAttn_22Causal Attention1 × 448 × 1280
191Dec_CrossAttn_22Cross-Attention1 × 448 × 1280
192Dec_Norm_22LayerNorm1 × 448 × 1280
193Dec_FFN_22Feed Forward1 × 448 × 1280
194Dec_SelfAttn_23Causal Attention1 × 448 × 1280
195Dec_CrossAttn_23Cross-Attention1 × 448 × 1280
196Dec_Norm_23LayerNorm1 × 448 × 1280
197Dec_FFN_23Feed Forward1 × 448 × 1280
198Dec_SelfAttn_24Causal Attention1 × 448 × 1280
199Dec_CrossAttn_24Cross-Attention1 × 448 × 1280
200Dec_Norm_24LayerNorm1 × 448 × 1280
201Dec_FFN_24Feed Forward1 × 448 × 1280
202Dec_SelfAttn_25Causal Attention1 × 448 × 1280
203Dec_CrossAttn_25Cross-Attention1 × 448 × 1280
204Dec_Norm_25LayerNorm1 × 448 × 1280
205Dec_FFN_25Feed Forward1 × 448 × 1280
206Dec_SelfAttn_26Causal Attention1 × 448 × 1280
207Dec_CrossAttn_26Cross-Attention1 × 448 × 1280
208Dec_Norm_26LayerNorm1 × 448 × 1280
209Dec_FFN_26Feed Forward1 × 448 × 1280
210Dec_SelfAttn_27Causal Attention1 × 448 × 1280
211Dec_CrossAttn_27Cross-Attention1 × 448 × 1280
212Dec_Norm_27LayerNorm1 × 448 × 1280
213Dec_FFN_27Feed Forward1 × 448 × 1280
214Dec_SelfAttn_28Causal Attention1 × 448 × 1280
215Dec_CrossAttn_28Cross-Attention1 × 448 × 1280
216Dec_Norm_28LayerNorm1 × 448 × 1280
217Dec_FFN_28Feed Forward1 × 448 × 1280
218Dec_SelfAttn_29Causal Attention1 × 448 × 1280
219Dec_CrossAttn_29Cross-Attention1 × 448 × 1280
220Dec_Norm_29LayerNorm1 × 448 × 1280
221Dec_FFN_29Feed Forward1 × 448 × 1280
222Dec_SelfAttn_30Causal Attention1 × 448 × 1280
223Dec_CrossAttn_30Cross-Attention1 × 448 × 1280
224Dec_Norm_30LayerNorm1 × 448 × 1280
225Dec_FFN_30Feed Forward1 × 448 × 1280
226Dec_SelfAttn_31Causal Attention1 × 448 × 1280
227Dec_CrossAttn_31Cross-Attention1 × 448 × 1280
228Dec_Norm_31LayerNorm1 × 448 × 1280
229Dec_FFN_31Feed Forward1 × 448 × 1280
230Dec_SelfAttn_32Causal Attention1 × 448 × 1280
231Dec_CrossAttn_32Cross-Attention1 × 448 × 1280
232Dec_Norm_32LayerNorm1 × 448 × 1280
233Dec_FFN_32Feed Forward1 × 448 × 1280
234Dec_Norm_FinalLayerNorm1 × 448 × 1280
235LM_HeadLinear1 × 448 × 51866
236OutputOutput1 × 448 × 51866

What the verifier says

infoAt 64 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace Oriserve/Whisper-Hindi2Hinglish-Prime --plan --share

Other whisper checkpoints

Whisper-Hindi2Hinglish-Apex
801M derived · -0.91% against the checkpoint
whisper-large-v3
1.54B derived · -0.49% against the checkpoint
whisper-large-v3-turbo
801M derived · -0.91% against the checkpoint