Models / unlimited-ocr
Unlimited-OCR-AWQ
Reconstructed from its own config.json
with no weights read. 1.3M downloads on Hugging Face.
Our count against the checkpoint
The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.
quantized This checkpoint is stored quantized (compressed-tensors). The tensor count in the file counts stored elements under a packing scheme, not logical parameters, so the two numbers below are not measuring the same thing in either direction.
What it costs to run
Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.
| Card | Memory | |
|---|---|---|
| T4 (16GB) | weights + activations | does not fit |
| A100 (40GB) | weights + activations | does not fit |
| H100 (80GB) | weights + activations | fits |
Structure
84 nodes. Output shapes are propagated from the input shape, batch dimension excluded.
| Layer | Type | Output shape | |
|---|---|---|---|
| 1 | Input | Input | 1 × 32768 |
| 2 | Embedding | Embedding | 1 × 32768 × 1280 |
| 3 | RoPE | RoPE | 1 × 32768 × 1280 |
| 4 | Vision input | Input | 3 × 1024 × 1024 |
| 5 | PatchEmbed | Patch Embed | 4096 × 1280 |
| 6 | Patch_Position_Embedding | Learned Pos Embed | 4096 × 1280 |
| 7 | Vision encoder (internals not in config) | Projection | 4096 × 1280 |
| 8 | Vision tokens | Reshape | 1 × 4096 × 1280 |
| 9 | Multimodal fusion (concat tokens) | Concatenate | 1 × 36864 × 1280 |
| 10 | RMSNorm_1_1 | RMSNorm | 1 × 36864 × 1280 |
| 11 | Attention_1 | Grouped Query Attn | 1 × 36864 × 1280 |
| 12 | Add_1_attn | Add | 1 × 36864 × 1280 |
| 13 | RMSNorm_1_2 | RMSNorm | 1 × 36864 × 1280 |
| 14 | FFN_1 | SwiGLU | 1 × 36864 × 1280 |
| 15 | Add_1_ffn | Add | 1 × 36864 × 1280 |
| 16 | RMSNorm_2_1 | RMSNorm | 1 × 36864 × 1280 |
| 17 | Attention_2 | Grouped Query Attn | 1 × 36864 × 1280 |
| 18 | Add_2_attn | Add | 1 × 36864 × 1280 |
| 19 | RMSNorm_2_2 | RMSNorm | 1 × 36864 × 1280 |
| 20 | MoE_2 | Shared-Expert MoE | 1 × 36864 × 1280 |
| 21 | Add_2_ffn | Add | 1 × 36864 × 1280 |
| 22 | RMSNorm_3_1 | RMSNorm | 1 × 36864 × 1280 |
| 23 | Attention_3 | Grouped Query Attn | 1 × 36864 × 1280 |
| 24 | Add_3_attn | Add | 1 × 36864 × 1280 |
| 25 | RMSNorm_3_2 | RMSNorm | 1 × 36864 × 1280 |
| 26 | MoE_3 | Shared-Expert MoE | 1 × 36864 × 1280 |
| 27 | Add_3_ffn | Add | 1 × 36864 × 1280 |
| 28 | RMSNorm_4_1 | RMSNorm | 1 × 36864 × 1280 |
| 29 | Attention_4 | Grouped Query Attn | 1 × 36864 × 1280 |
| 30 | Add_4_attn | Add | 1 × 36864 × 1280 |
| 31 | RMSNorm_4_2 | RMSNorm | 1 × 36864 × 1280 |
| 32 | MoE_4 | Shared-Expert MoE | 1 × 36864 × 1280 |
| 33 | Add_4_ffn | Add | 1 × 36864 × 1280 |
| 34 | RMSNorm_5_1 | RMSNorm | 1 × 36864 × 1280 |
| 35 | Attention_5 | Grouped Query Attn | 1 × 36864 × 1280 |
| 36 | Add_5_attn | Add | 1 × 36864 × 1280 |
| 37 | RMSNorm_5_2 | RMSNorm | 1 × 36864 × 1280 |
| 38 | MoE_5 | Shared-Expert MoE | 1 × 36864 × 1280 |
| 39 | Add_5_ffn | Add | 1 × 36864 × 1280 |
| 40 | RMSNorm_6_1 | RMSNorm | 1 × 36864 × 1280 |
| 41 | Attention_6 | Grouped Query Attn | 1 × 36864 × 1280 |
| 42 | Add_6_attn | Add | 1 × 36864 × 1280 |
| 43 | RMSNorm_6_2 | RMSNorm | 1 × 36864 × 1280 |
| 44 | MoE_6 | Shared-Expert MoE | 1 × 36864 × 1280 |
| 45 | Add_6_ffn | Add | 1 × 36864 × 1280 |
| 46 | RMSNorm_7_1 | RMSNorm | 1 × 36864 × 1280 |
| 47 | Attention_7 | Grouped Query Attn | 1 × 36864 × 1280 |
| 48 | Add_7_attn | Add | 1 × 36864 × 1280 |
| 49 | RMSNorm_7_2 | RMSNorm | 1 × 36864 × 1280 |
| 50 | MoE_7 | Shared-Expert MoE | 1 × 36864 × 1280 |
| 51 | Add_7_ffn | Add | 1 × 36864 × 1280 |
| 52 | RMSNorm_8_1 | RMSNorm | 1 × 36864 × 1280 |
| 53 | Attention_8 | Grouped Query Attn | 1 × 36864 × 1280 |
| 54 | Add_8_attn | Add | 1 × 36864 × 1280 |
| 55 | RMSNorm_8_2 | RMSNorm | 1 × 36864 × 1280 |
| 56 | MoE_8 | Shared-Expert MoE | 1 × 36864 × 1280 |
| 57 | Add_8_ffn | Add | 1 × 36864 × 1280 |
| 58 | RMSNorm_9_1 | RMSNorm | 1 × 36864 × 1280 |
| 59 | Attention_9 | Grouped Query Attn | 1 × 36864 × 1280 |
| 60 | Add_9_attn | Add | 1 × 36864 × 1280 |
| 61 | RMSNorm_9_2 | RMSNorm | 1 × 36864 × 1280 |
| 62 | MoE_9 | Shared-Expert MoE | 1 × 36864 × 1280 |
| 63 | Add_9_ffn | Add | 1 × 36864 × 1280 |
| 64 | RMSNorm_10_1 | RMSNorm | 1 × 36864 × 1280 |
| 65 | Attention_10 | Grouped Query Attn | 1 × 36864 × 1280 |
| 66 | Add_10_attn | Add | 1 × 36864 × 1280 |
| 67 | RMSNorm_10_2 | RMSNorm | 1 × 36864 × 1280 |
| 68 | MoE_10 | Shared-Expert MoE | 1 × 36864 × 1280 |
| 69 | Add_10_ffn | Add | 1 × 36864 × 1280 |
| 70 | RMSNorm_11_1 | RMSNorm | 1 × 36864 × 1280 |
| 71 | Attention_11 | Grouped Query Attn | 1 × 36864 × 1280 |
| 72 | Add_11_attn | Add | 1 × 36864 × 1280 |
| 73 | RMSNorm_11_2 | RMSNorm | 1 × 36864 × 1280 |
| 74 | MoE_11 | Shared-Expert MoE | 1 × 36864 × 1280 |
| 75 | Add_11_ffn | Add | 1 × 36864 × 1280 |
| 76 | RMSNorm_12_1 | RMSNorm | 1 × 36864 × 1280 |
| 77 | Attention_12 | Grouped Query Attn | 1 × 36864 × 1280 |
| 78 | Add_12_attn | Add | 1 × 36864 × 1280 |
| 79 | RMSNorm_12_2 | RMSNorm | 1 × 36864 × 1280 |
| 80 | MoE_12 | Shared-Expert MoE | 1 × 36864 × 1280 |
| 81 | Add_12_ffn | Add | 1 × 36864 × 1280 |
| 82 | Final_RMSNorm | RMSNorm | 1 × 36864 × 1280 |
| 83 | LM_Head | Linear | 1 × 36864 × 129280 |
| 84 | Output | Output | 1 × 36864 × 129280 |
What the verifier says
dead-end
swiglu-dim-convention
deep-attention-default-init
Do this to your own model
Same numbers, on a model in your repo, in one command. No account.
pip install neurarch-trace
neurarch-trace sahilchachra/Unlimited-OCR-AWQ --plan --share