N Neurarch Architectures Models Checks Data Docs Open the app

Models / deepseek_vl_v2

DeepSeek-OCR-2

Reconstructed from its own config.json with no weights read. 963K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
2.78B
2,777,122,560 parameters
In the published checkpoint
3.39B
3,389,119,360 scalars · safetensors.total, read 2026-02-03
Delta
-18.1%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `modeling_deepseekocr2.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
79
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$274.99
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsfits

Structure

82 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 8192
2EmbeddingEmbedding1 × 8192 × 1280
3RoPERoPE1 × 8192 × 1280
4Vision inputInput3 × 1024 × 1024
5PatchEmbedPatch Embed4096 × 1280
6Patch_Position_EmbeddingLearned Pos Embed4096 × 1280
7Vision encoder (internals not in config)Projection4096 × 1280
8Vision tokensReshape1 × 4096 × 1280
9Multimodal fusion (concat tokens)Concatenate1 × 12288 × 1280
10RMSNorm_1_1RMSNorm1 × 12288 × 1280
11Attention_1Grouped Query Attn1 × 12288 × 1280
12Add_1_attnAdd1 × 12288 × 1280
13RMSNorm_1_2RMSNorm1 × 12288 × 1280
14FFN_1SwiGLU1 × 12288 × 1280
15Add_1_ffnAdd1 × 12288 × 1280
16RMSNorm_2_1RMSNorm1 × 12288 × 1280
17Attention_2Grouped Query Attn1 × 12288 × 1280
18Add_2_attnAdd1 × 12288 × 1280
19RMSNorm_2_2RMSNorm1 × 12288 × 1280
20MoE_2Shared-Expert MoE1 × 12288 × 1280
21Add_2_ffnAdd1 × 12288 × 1280
22RMSNorm_3_1RMSNorm1 × 12288 × 1280
23Attention_3Grouped Query Attn1 × 12288 × 1280
24Add_3_attnAdd1 × 12288 × 1280
25RMSNorm_3_2RMSNorm1 × 12288 × 1280
26MoE_3Shared-Expert MoE1 × 12288 × 1280
27Add_3_ffnAdd1 × 12288 × 1280
28RMSNorm_4_1RMSNorm1 × 12288 × 1280
29Attention_4Grouped Query Attn1 × 12288 × 1280
30Add_4_attnAdd1 × 12288 × 1280
31RMSNorm_4_2RMSNorm1 × 12288 × 1280
32MoE_4Shared-Expert MoE1 × 12288 × 1280
33Add_4_ffnAdd1 × 12288 × 1280
34RMSNorm_5_1RMSNorm1 × 12288 × 1280
35Attention_5Grouped Query Attn1 × 12288 × 1280
36Add_5_attnAdd1 × 12288 × 1280
37RMSNorm_5_2RMSNorm1 × 12288 × 1280
38MoE_5Shared-Expert MoE1 × 12288 × 1280
39Add_5_ffnAdd1 × 12288 × 1280
40RMSNorm_6_1RMSNorm1 × 12288 × 1280
41Attention_6Grouped Query Attn1 × 12288 × 1280
42Add_6_attnAdd1 × 12288 × 1280
43RMSNorm_6_2RMSNorm1 × 12288 × 1280
44MoE_6Shared-Expert MoE1 × 12288 × 1280
45Add_6_ffnAdd1 × 12288 × 1280
46RMSNorm_7_1RMSNorm1 × 12288 × 1280
47Attention_7Grouped Query Attn1 × 12288 × 1280
48Add_7_attnAdd1 × 12288 × 1280
49RMSNorm_7_2RMSNorm1 × 12288 × 1280
50MoE_7Shared-Expert MoE1 × 12288 × 1280
51Add_7_ffnAdd1 × 12288 × 1280
52RMSNorm_8_1RMSNorm1 × 12288 × 1280
53Attention_8Grouped Query Attn1 × 12288 × 1280
54Add_8_attnAdd1 × 12288 × 1280
55RMSNorm_8_2RMSNorm1 × 12288 × 1280
56MoE_8Shared-Expert MoE1 × 12288 × 1280
57Add_8_ffnAdd1 × 12288 × 1280
58RMSNorm_9_1RMSNorm1 × 12288 × 1280
59Attention_9Grouped Query Attn1 × 12288 × 1280
60Add_9_attnAdd1 × 12288 × 1280
61RMSNorm_9_2RMSNorm1 × 12288 × 1280
62MoE_9Shared-Expert MoE1 × 12288 × 1280
63Add_9_ffnAdd1 × 12288 × 1280
64RMSNorm_10_1RMSNorm1 × 12288 × 1280
65Attention_10Grouped Query Attn1 × 12288 × 1280
66Add_10_attnAdd1 × 12288 × 1280
67RMSNorm_10_2RMSNorm1 × 12288 × 1280
68MoE_10Shared-Expert MoE1 × 12288 × 1280
69Add_10_ffnAdd1 × 12288 × 1280
70RMSNorm_11_1RMSNorm1 × 12288 × 1280
71Attention_11Grouped Query Attn1 × 12288 × 1280
72Add_11_attnAdd1 × 12288 × 1280
73RMSNorm_11_2RMSNorm1 × 12288 × 1280
74MoE_11Shared-Expert MoE1 × 12288 × 1280
75Add_11_ffnAdd1 × 12288 × 1280
76RMSNorm_12_1RMSNorm1 × 12288 × 1280
77Attention_12Grouped Query Attn1 × 12288 × 1280
78Add_12_attnAdd1 × 12288 × 1280
79RMSNorm_12_2RMSNorm1 × 12288 × 1280
80MoE_12Shared-Expert MoE1 × 12288 × 1280
81Add_12_ffnAdd1 × 12288 × 1280
82OutputOutput1 × 12288 × 1280

What the verifier says

warn"RoPE" receives input but its output is not connected. This layer will be unreachable in the forward pass. Fix: Connect the output forward, or add an Output node if this is the final layer.
dead-end
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 6848 (5.35× embedDim). Expected: ~3328. Fix: Set intermediateSize to 3328 for embedDim=1280.
swiglu-dim-convention
infoAt 12 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace deepseek-ai/DeepSeek-OCR-2 --plan --share

Other deepseek_vl_v2 checkpoints

DeepSeek-OCR
2.78B derived · -16.8% against the checkpoint