N Neurarch Architectures Models Checks Data Docs Open the app

Models / llada2_moe

LLaDA2.1-mini

Reconstructed from its own config.json with no weights read. 110K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
1.14B
1,144,385,536 parameters
In the published checkpoint
16.26B
16,255,643,392 scalars · safetensors.total, read 2026-04-13
Delta
-93.0%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `configuration_llada2_moe.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
82
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$2572.10
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsfits
H100 (80GB)weights + activationsfits

Structure

84 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 32768
2EmbeddingEmbedding1 × 32768 × 2048
3Positional_EmbeddingLearned Pos Embed1 × 32768 × 2048
4Attention_1Multi-Head Attention1 × 32768 × 2048
5Add_1Add1 × 32768 × 2048
6LayerNorm_1_1LayerNorm1 × 32768 × 2048
7FFN_1Feed Forward1 × 32768 × 2048
8Attention_2Multi-Head Attention1 × 32768 × 2048
9Add_2Add1 × 32768 × 2048
10LayerNorm_2_1LayerNorm1 × 32768 × 2048
11FFN_2Feed Forward1 × 32768 × 2048
12Attention_3Multi-Head Attention1 × 32768 × 2048
13Add_3Add1 × 32768 × 2048
14LayerNorm_3_1LayerNorm1 × 32768 × 2048
15FFN_3Feed Forward1 × 32768 × 2048
16Attention_4Multi-Head Attention1 × 32768 × 2048
17Add_4Add1 × 32768 × 2048
18LayerNorm_4_1LayerNorm1 × 32768 × 2048
19FFN_4Feed Forward1 × 32768 × 2048
20Attention_5Multi-Head Attention1 × 32768 × 2048
21Add_5Add1 × 32768 × 2048
22LayerNorm_5_1LayerNorm1 × 32768 × 2048
23FFN_5Feed Forward1 × 32768 × 2048
24Attention_6Multi-Head Attention1 × 32768 × 2048
25Add_6Add1 × 32768 × 2048
26LayerNorm_6_1LayerNorm1 × 32768 × 2048
27FFN_6Feed Forward1 × 32768 × 2048
28Attention_7Multi-Head Attention1 × 32768 × 2048
29Add_7Add1 × 32768 × 2048
30LayerNorm_7_1LayerNorm1 × 32768 × 2048
31FFN_7Feed Forward1 × 32768 × 2048
32Attention_8Multi-Head Attention1 × 32768 × 2048
33Add_8Add1 × 32768 × 2048
34LayerNorm_8_1LayerNorm1 × 32768 × 2048
35FFN_8Feed Forward1 × 32768 × 2048
36Attention_9Multi-Head Attention1 × 32768 × 2048
37Add_9Add1 × 32768 × 2048
38LayerNorm_9_1LayerNorm1 × 32768 × 2048
39FFN_9Feed Forward1 × 32768 × 2048
40Attention_10Multi-Head Attention1 × 32768 × 2048
41Add_10Add1 × 32768 × 2048
42LayerNorm_10_1LayerNorm1 × 32768 × 2048
43FFN_10Feed Forward1 × 32768 × 2048
44Attention_11Multi-Head Attention1 × 32768 × 2048
45Add_11Add1 × 32768 × 2048
46LayerNorm_11_1LayerNorm1 × 32768 × 2048
47FFN_11Feed Forward1 × 32768 × 2048
48Attention_12Multi-Head Attention1 × 32768 × 2048
49Add_12Add1 × 32768 × 2048
50LayerNorm_12_1LayerNorm1 × 32768 × 2048
51FFN_12Feed Forward1 × 32768 × 2048
52Attention_13Multi-Head Attention1 × 32768 × 2048
53Add_13Add1 × 32768 × 2048
54LayerNorm_13_1LayerNorm1 × 32768 × 2048
55FFN_13Feed Forward1 × 32768 × 2048
56Attention_14Multi-Head Attention1 × 32768 × 2048
57Add_14Add1 × 32768 × 2048
58LayerNorm_14_1LayerNorm1 × 32768 × 2048
59FFN_14Feed Forward1 × 32768 × 2048
60Attention_15Multi-Head Attention1 × 32768 × 2048
61Add_15Add1 × 32768 × 2048
62LayerNorm_15_1LayerNorm1 × 32768 × 2048
63FFN_15Feed Forward1 × 32768 × 2048
64Attention_16Multi-Head Attention1 × 32768 × 2048
65Add_16Add1 × 32768 × 2048
66LayerNorm_16_1LayerNorm1 × 32768 × 2048
67FFN_16Feed Forward1 × 32768 × 2048
68Attention_17Multi-Head Attention1 × 32768 × 2048
69Add_17Add1 × 32768 × 2048
70LayerNorm_17_1LayerNorm1 × 32768 × 2048
71FFN_17Feed Forward1 × 32768 × 2048
72Attention_18Multi-Head Attention1 × 32768 × 2048
73Add_18Add1 × 32768 × 2048
74LayerNorm_18_1LayerNorm1 × 32768 × 2048
75FFN_18Feed Forward1 × 32768 × 2048
76Attention_19Multi-Head Attention1 × 32768 × 2048
77Add_19Add1 × 32768 × 2048
78LayerNorm_19_1LayerNorm1 × 32768 × 2048
79FFN_19Feed Forward1 × 32768 × 2048
80Attention_20Multi-Head Attention1 × 32768 × 2048
81Add_20Add1 × 32768 × 2048
82LayerNorm_20_1LayerNorm1 × 32768 × 2048
83FFN_20Feed Forward1 × 32768 × 2048
84OutputOutput1 × 32768 × 2048

What the verifier says

info20 attention layers at embedDim 2048 cache full per-head K/V: about 160 KB per token at fp16, which dominates memory at long context. Grouped-query attention (e.g. 8:1) would cut this ~8×; multi-head latent attention (MLA) shrinks it ~10× or more. This is the move production LLMs make; it does not change the parameter count. Fix: Switch attention to groupedQueryAttention (set numKVHeads below numHeads, e.g. numHeads/4) or mla (a low-rank cached latent).
full-mha-serving-cost
infoAt 20 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace inclusionAI/LLaDA2.1-mini --plan --share

Other llada2_moe checkpoints

LLaDA2.0-mini
1.14B derived · -93.0% against the checkpoint