# Mamba SSM Block vs Jamba

Pure state space against a hybrid that keeps some attention.

**Jamba has 13B more parameters than Mamba SSM Block: 42 layers added, 7 removed, 8 changed.**

Source: https://neurarch.com/diff/mamba-block-vs-jamba.html

## Sides

| | Mamba SSM Block | Jamba |
|---|---|---|
| Layers | 16 | 51 |
| Parameters | 168M | 13B |
| Input | 1 × 1024 | 1 × 4096 |
| Output | 1 × 1024 × 50280 | 1 × 4096 × 65536 |
| Forward-passes | yes | yes |
| Est. train cost | $2.58 | $338.33 |
| T4 16GB | fits | no |
| A100 40GB | fits | no |
| H100 80GB | fits | no |

## Deltas (Jamba relative to Mamba SSM Block)

- Parameters: +13B (77× the size)
- Layers: +35
- Added 42, removed 7, changed 8, unchanged 3

## Layer by layer

| # | Status | Mamba SSM Block | Params | Output | Jamba | Params | Output |
|---|---|---|---|---|---|---|---|
| 1 | changed (shape) | tokens (Input) |  | 1 × 1024 | tokens (Input) |  | 1 × 4096 |
| 2 | changed (numEmbeddings, embeddingDim) | embed (Embedding) | 51M | 1 × 1024 × 1024 | token_embed (Embedding) | 268M | 1 × 4096 × 4096 |
| 3 | added |  | |  | mixer_norm_1 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 4 | added |  | |  | mamba_1 (Mamba) | 101M | 1 × 4096 × 4096 |
| 5 | added |  | |  | mixer_residual_1 (Add) |  | 1 × 4096 × 4096 |
| 6 | changed (normalizedShape) | norm_ssm (Rms Norm) | 1.0K | 1 × 1024 × 1024 | ffn_norm_1 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 7 | changed (type, outFeatures, embedDim, ffDim) | in_proj (Linear) |  | 1 × 1024 × 4096 | mlp_1 (Feed Forward) | 117M | 1 × 4096 × 4096 |
| 8 | removed | to_channels (Permute) |  | 1 × 4096 × 1024 |  | |  |
| 9 | removed | z_gate (Swish) |  | 1 × 1024 × 4096 |  | |  |
| 10 | removed | causal_conv (Conv1d) | 16K | 1 × 4096 × 1024 |  | |  |
| 11 | removed | to_tokens (Permute) |  | 1 × 1024 × 4096 |  | |  |
| 12 | removed | silu_x (Swish) |  | 1 × 1024 × 4096 |  | |  |
| 13 | added |  | |  | ffn_residual_1 (Add) |  | 1 × 4096 × 4096 |
| 14 | added |  | |  | mixer_norm_2 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 15 | changed (dConv, expand) | ssm_scan (Mamba) | 51M | 1 × 1024 × 4096 | mamba_2 (Mamba) | 101M | 1 × 4096 × 4096 |
| 16 | removed | gate_out (Multiply) |  | 1 × 1024 × 4096 |  | |  |
| 17 | added |  | |  | mixer_residual_2 (Add) |  | 1 × 4096 × 4096 |
| 18 | added |  | |  | ffn_norm_2 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 19 | added |  | |  | moe_2 (Moe Layer) | 2.8B | 1 × 4096 × 4096 |
| 20 | added |  | |  | ffn_residual_2 (Add) |  | 1 × 4096 × 4096 |
| 21 | added |  | |  | mixer_norm_3 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 22 | added |  | |  | mamba_3 (Mamba) | 101M | 1 × 4096 × 4096 |
| 23 | added |  | |  | mixer_residual_3 (Add) |  | 1 × 4096 × 4096 |
| 24 | added |  | |  | ffn_norm_3 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 25 | changed (type, outFeatures, embedDim, ffDim) | out_proj (Linear) |  | 1 × 1024 × 1024 | mlp_3 (Feed Forward) | 117M | 1 × 4096 × 4096 |
| 26 | same | residual_1 (Add) |  | 1 × 1024 × 1024 | ffn_residual_3 (Add) |  | 1 × 4096 × 4096 |
| 27 | changed (normalizedShape) | norm_ffn (Rms Norm) | 1.0K | 1 × 1024 × 1024 | mixer_norm_4 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 28 | removed | ffn (Swiglu) | 6.3M | 1 × 1024 × 1024 |  | |  |
| 29 | added |  | |  | mamba_4 (Mamba) | 101M | 1 × 4096 × 4096 |
| 30 | added |  | |  | mixer_residual_4 (Add) |  | 1 × 4096 × 4096 |
| 31 | added |  | |  | ffn_norm_4 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 32 | added |  | |  | moe_4 (Moe Layer) | 2.8B | 1 × 4096 × 4096 |
| 33 | added |  | |  | ffn_residual_4 (Add) |  | 1 × 4096 × 4096 |
| 34 | added |  | |  | mixer_norm_5 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 35 | added |  | |  | attn_5 (Grouped Query Attention) | 42M | 1 × 4096 × 4096 |
| 36 | added |  | |  | mixer_residual_5 (Add) |  | 1 × 4096 × 4096 |
| 37 | added |  | |  | ffn_norm_5 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 38 | added |  | |  | mlp_5 (Feed Forward) | 117M | 1 × 4096 × 4096 |
| 39 | added |  | |  | ffn_residual_5 (Add) |  | 1 × 4096 × 4096 |
| 40 | added |  | |  | mixer_norm_6 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 41 | added |  | |  | mamba_6 (Mamba) | 101M | 1 × 4096 × 4096 |
| 42 | added |  | |  | mixer_residual_6 (Add) |  | 1 × 4096 × 4096 |
| 43 | added |  | |  | ffn_norm_6 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 44 | added |  | |  | moe_6 (Moe Layer) | 2.8B | 1 × 4096 × 4096 |
| 45 | added |  | |  | ffn_residual_6 (Add) |  | 1 × 4096 × 4096 |
| 46 | added |  | |  | mixer_norm_7 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 47 | added |  | |  | mamba_7 (Mamba) | 101M | 1 × 4096 × 4096 |
| 48 | added |  | |  | mixer_residual_7 (Add) |  | 1 × 4096 × 4096 |
| 49 | added |  | |  | ffn_norm_7 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 50 | added |  | |  | mlp_7 (Feed Forward) | 117M | 1 × 4096 × 4096 |
| 51 | added |  | |  | ffn_residual_7 (Add) |  | 1 × 4096 × 4096 |
| 52 | added |  | |  | mixer_norm_8 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 53 | added |  | |  | mamba_8 (Mamba) | 101M | 1 × 4096 × 4096 |
| 54 | added |  | |  | mixer_residual_8 (Add) |  | 1 × 4096 × 4096 |
| 55 | added |  | |  | ffn_norm_8 (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 56 | added |  | |  | moe_8 (Moe Layer) | 2.8B | 1 × 4096 × 4096 |
| 57 | same | residual_2 (Add) |  | 1 × 1024 × 1024 | ffn_residual_8 (Add) |  | 1 × 4096 × 4096 |
| 58 | added |  | |  | final_norm (Rms Norm) | 4.1K | 1 × 4096 × 4096 |
| 59 | changed (outFeatures, inFeatures) | lm_head (Linear) |  | 1 × 1024 × 50280 | lm_head (Linear) | 269M | 1 × 4096 × 65536 |
| 60 | same | output (Output) |  | 1 × 1024 × 50280 | logits (Output) |  | 1 × 4096 × 65536 |

## What this is not

- The two are priced at different declared inputs (1 × 1024 against 1 × 4096), so memory, cost and GPU fit are each right about their own model and are not a comparison between them. The layer and parameter deltas are unaffected.
- Parameter counts are derived from the graph, not read from a checkpoint. They are exact for a graph that is fully specified and approximate for one that is not.
- Cost and GPU fit are estimates from the graph under one set of assumptions, not measurements of a run.

## Graphs

- Mamba SSM Block: https://neurarch.com/templates/mamba-block/model.json
- Jamba: https://neurarch.com/templates/jamba/model.json
- Check a graph of your own: `POST https://www.neurarch.com/api/v1/plan`
