N Neurarch Architectures Models Checks Data Docs Open the app

Models / tipsv2

tipsv2-so400m14

Reconstructed from its own config.json with no weights read. 270K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
448M
448,342,128 parameters
In the published checkpoint
862M
861,726,816 scalars · safetensors.total, read 2026-08-17
Delta
-48.0%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `configuration_tips.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
110
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$1.03
CardMemory
T4 (16GB)weights + activationsfits
A100 (40GB)weights + activationsfits
H100 (80GB)weights + activationsfits

Structure

112 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 64
2EmbeddingEmbedding1 × 64 × 1152
3Positional_EmbeddingLearned Pos Embed1 × 64 × 1152
4Attention_1Multi-Head Attention1 × 64 × 1152
5Add_1Add1 × 64 × 1152
6LayerNorm_1_1LayerNorm1 × 64 × 1152
7FFN_1Feed Forward1 × 64 × 1152
8Attention_2Multi-Head Attention1 × 64 × 1152
9Add_2Add1 × 64 × 1152
10LayerNorm_2_1LayerNorm1 × 64 × 1152
11FFN_2Feed Forward1 × 64 × 1152
12Attention_3Multi-Head Attention1 × 64 × 1152
13Add_3Add1 × 64 × 1152
14LayerNorm_3_1LayerNorm1 × 64 × 1152
15FFN_3Feed Forward1 × 64 × 1152
16Attention_4Multi-Head Attention1 × 64 × 1152
17Add_4Add1 × 64 × 1152
18LayerNorm_4_1LayerNorm1 × 64 × 1152
19FFN_4Feed Forward1 × 64 × 1152
20Attention_5Multi-Head Attention1 × 64 × 1152
21Add_5Add1 × 64 × 1152
22LayerNorm_5_1LayerNorm1 × 64 × 1152
23FFN_5Feed Forward1 × 64 × 1152
24Attention_6Multi-Head Attention1 × 64 × 1152
25Add_6Add1 × 64 × 1152
26LayerNorm_6_1LayerNorm1 × 64 × 1152
27FFN_6Feed Forward1 × 64 × 1152
28Attention_7Multi-Head Attention1 × 64 × 1152
29Add_7Add1 × 64 × 1152
30LayerNorm_7_1LayerNorm1 × 64 × 1152
31FFN_7Feed Forward1 × 64 × 1152
32Attention_8Multi-Head Attention1 × 64 × 1152
33Add_8Add1 × 64 × 1152
34LayerNorm_8_1LayerNorm1 × 64 × 1152
35FFN_8Feed Forward1 × 64 × 1152
36Attention_9Multi-Head Attention1 × 64 × 1152
37Add_9Add1 × 64 × 1152
38LayerNorm_9_1LayerNorm1 × 64 × 1152
39FFN_9Feed Forward1 × 64 × 1152
40Attention_10Multi-Head Attention1 × 64 × 1152
41Add_10Add1 × 64 × 1152
42LayerNorm_10_1LayerNorm1 × 64 × 1152
43FFN_10Feed Forward1 × 64 × 1152
44Attention_11Multi-Head Attention1 × 64 × 1152
45Add_11Add1 × 64 × 1152
46LayerNorm_11_1LayerNorm1 × 64 × 1152
47FFN_11Feed Forward1 × 64 × 1152
48Attention_12Multi-Head Attention1 × 64 × 1152
49Add_12Add1 × 64 × 1152
50LayerNorm_12_1LayerNorm1 × 64 × 1152
51FFN_12Feed Forward1 × 64 × 1152
52Attention_13Multi-Head Attention1 × 64 × 1152
53Add_13Add1 × 64 × 1152
54LayerNorm_13_1LayerNorm1 × 64 × 1152
55FFN_13Feed Forward1 × 64 × 1152
56Attention_14Multi-Head Attention1 × 64 × 1152
57Add_14Add1 × 64 × 1152
58LayerNorm_14_1LayerNorm1 × 64 × 1152
59FFN_14Feed Forward1 × 64 × 1152
60Attention_15Multi-Head Attention1 × 64 × 1152
61Add_15Add1 × 64 × 1152
62LayerNorm_15_1LayerNorm1 × 64 × 1152
63FFN_15Feed Forward1 × 64 × 1152
64Attention_16Multi-Head Attention1 × 64 × 1152
65Add_16Add1 × 64 × 1152
66LayerNorm_16_1LayerNorm1 × 64 × 1152
67FFN_16Feed Forward1 × 64 × 1152
68Attention_17Multi-Head Attention1 × 64 × 1152
69Add_17Add1 × 64 × 1152
70LayerNorm_17_1LayerNorm1 × 64 × 1152
71FFN_17Feed Forward1 × 64 × 1152
72Attention_18Multi-Head Attention1 × 64 × 1152
73Add_18Add1 × 64 × 1152
74LayerNorm_18_1LayerNorm1 × 64 × 1152
75FFN_18Feed Forward1 × 64 × 1152
76Attention_19Multi-Head Attention1 × 64 × 1152
77Add_19Add1 × 64 × 1152
78LayerNorm_19_1LayerNorm1 × 64 × 1152
79FFN_19Feed Forward1 × 64 × 1152
80Attention_20Multi-Head Attention1 × 64 × 1152
81Add_20Add1 × 64 × 1152
82LayerNorm_20_1LayerNorm1 × 64 × 1152
83FFN_20Feed Forward1 × 64 × 1152
84Attention_21Multi-Head Attention1 × 64 × 1152
85Add_21Add1 × 64 × 1152
86LayerNorm_21_1LayerNorm1 × 64 × 1152
87FFN_21Feed Forward1 × 64 × 1152
88Attention_22Multi-Head Attention1 × 64 × 1152
89Add_22Add1 × 64 × 1152
90LayerNorm_22_1LayerNorm1 × 64 × 1152
91FFN_22Feed Forward1 × 64 × 1152
92Attention_23Multi-Head Attention1 × 64 × 1152
93Add_23Add1 × 64 × 1152
94LayerNorm_23_1LayerNorm1 × 64 × 1152
95FFN_23Feed Forward1 × 64 × 1152
96Attention_24Multi-Head Attention1 × 64 × 1152
97Add_24Add1 × 64 × 1152
98LayerNorm_24_1LayerNorm1 × 64 × 1152
99FFN_24Feed Forward1 × 64 × 1152
100Attention_25Multi-Head Attention1 × 64 × 1152
101Add_25Add1 × 64 × 1152
102LayerNorm_25_1LayerNorm1 × 64 × 1152
103FFN_25Feed Forward1 × 64 × 1152
104Attention_26Multi-Head Attention1 × 64 × 1152
105Add_26Add1 × 64 × 1152
106LayerNorm_26_1LayerNorm1 × 64 × 1152
107FFN_26Feed Forward1 × 64 × 1152
108Attention_27Multi-Head Attention1 × 64 × 1152
109Add_27Add1 × 64 × 1152
110LayerNorm_27_1LayerNorm1 × 64 × 1152
111FFN_27Feed Forward1 × 64 × 1152
112OutputOutput1 × 64 × 1152

What the verifier says

infoAt 27 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace google/tipsv2-so400m14 --plan --share

Other tipsv2 checkpoints

tipsv2-b14
110M derived · -44.0% against the checkpoint