# Transformer Block vs Mamba SSM Block

Attention against a state-space layer for the same job.

**Mamba SSM Block has 161M more parameters than Transformer Block: 12 layers added, 2 removed, 4 changed.**

Source: https://neurarch.com/diff/transformer-block-vs-mamba-block.html

## Sides

| | Transformer Block | Mamba SSM Block |
|---|---|---|
| Layers | 6 | 16 |
| Parameters | 7.1M | 168M |
| Input | 512 × 768 | 1 × 1024 |
| Output | 512 × 768 | 1 × 1024 × 50280 |
| Forward-passes | yes | yes |
| Est. train cost | $0.185 | $2.58 |
| T4 16GB | fits | fits |
| A100 40GB | fits | fits |
| H100 80GB | fits | fits |

## Deltas (Mamba SSM Block relative to Transformer Block)

- Parameters: +161M (24× the size)
- Layers: +10
- Added 12, removed 2, changed 4, unchanged 2

## Layer by layer

| # | Status | Transformer Block | Params | Output | Mamba SSM Block | Params | Output |
|---|---|---|---|---|---|---|---|
| 1 | changed (shape) | Input (Input) |  | 512 × 768 | tokens (Input) |  | 1 × 1024 |
| 2 | removed | MultiHeadAttention (Multi Head Attention) | 2.4M | 512 × 768 |  | |  |
| 3 | removed | Add_1 (Add) |  | 512 × 768 |  | |  |
| 4 | added |  | |  | embed (Embedding) | 51M | 1 × 1024 × 1024 |
| 5 | changed (type, normalizedShape) | LayerNorm_1 (Layer Norm) |  | 512 × 768 | norm_ssm (Rms Norm) | 1.0K | 1 × 1024 × 1024 |
| 6 | added |  | |  | in_proj (Linear) |  | 1 × 1024 × 4096 |
| 7 | added |  | |  | to_channels (Permute) |  | 1 × 4096 × 1024 |
| 8 | added |  | |  | z_gate (Swish) |  | 1 × 1024 × 4096 |
| 9 | added |  | |  | causal_conv (Conv1d) | 16K | 1 × 4096 × 1024 |
| 10 | added |  | |  | to_tokens (Permute) |  | 1 × 1024 × 4096 |
| 11 | added |  | |  | silu_x (Swish) |  | 1 × 1024 × 4096 |
| 12 | added |  | |  | ssm_scan (Mamba) | 51M | 1 × 1024 × 4096 |
| 13 | added |  | |  | gate_out (Multiply) |  | 1 × 1024 × 4096 |
| 14 | changed (type, hiddenDim, ffDim, outFeatures) | FeedForward (Feed Forward) | 4.7M | 512 × 768 | out_proj (Linear) |  | 1 × 1024 × 1024 |
| 15 | same | Add_2 (Add) |  | 512 × 768 | residual_1 (Add) |  | 1 × 1024 × 1024 |
| 16 | changed (type, normalizedShape) | LayerNorm_2 (Layer Norm) |  | 512 × 768 | norm_ffn (Rms Norm) | 1.0K | 1 × 1024 × 1024 |
| 17 | added |  | |  | ffn (Swiglu) | 6.3M | 1 × 1024 × 1024 |
| 18 | added |  | |  | residual_2 (Add) |  | 1 × 1024 × 1024 |
| 19 | added |  | |  | lm_head (Linear) |  | 1 × 1024 × 50280 |
| 20 | same | Output (Output) |  | 512 × 768 | output (Output) |  | 1 × 1024 × 50280 |

## What this is not

- The two are priced at different declared inputs (512 × 768 against 1 × 1024), so memory, cost and GPU fit are each right about their own model and are not a comparison between them. The layer and parameter deltas are unaffected.
- Parameter counts are derived from the graph, not read from a checkpoint. They are exact for a graph that is fully specified and approximate for one that is not.
- Cost and GPU fit are estimates from the graph under one set of assumptions, not measurements of a run.

## Graphs

- Transformer Block: https://neurarch.com/templates/transformer-block/model.json
- Mamba SSM Block: https://neurarch.com/templates/mamba-block/model.json
- Check a graph of your own: `POST https://www.neurarch.com/api/v1/plan`
