# Simple RNN vs Mamba SSM Block

The recurrent layer everyone started with against its modern replacement.

**Mamba SSM Block has 167M more parameters than Simple RNN: 15 layers added, 1 removed, 2 changed.**

Source: https://neurarch.com/diff/simple-rnn-vs-mamba-block.html

## Sides

| | Simple RNN | Mamba SSM Block |
|---|---|---|
| Layers | 2 | 16 |
| Parameters | 1.1M | 168M |
| Input | 128 × 300 | 1 × 1024 |
| Output | 128 × 10 | 1 × 1024 × 50280 |
| Forward-passes | yes | yes |
| Est. train cost | $0.047 | $2.58 |
| T4 16GB | fits | fits |
| A100 40GB | fits | fits |
| H100 80GB | fits | fits |

## Deltas (Mamba SSM Block relative to Simple RNN)

- Parameters: +167M (153× the size)
- Layers: +14
- Added 15, removed 1, changed 2, unchanged 1

## Layer by layer

| # | Status | Simple RNN | Params | Output | Mamba SSM Block | Params | Output |
|---|---|---|---|---|---|---|---|
| 1 | changed (shape) | Input (Input) |  | 128 × 300 | tokens (Input) |  | 1 × 1024 |
| 2 | removed | LSTM (Lstm) | 791K | 128 × 256 |  | |  |
| 3 | added |  | |  | embed (Embedding) | 51M | 1 × 1024 × 1024 |
| 4 | added |  | |  | norm_ssm (Rms Norm) | 1.0K | 1 × 1024 × 1024 |
| 5 | added |  | |  | in_proj (Linear) |  | 1 × 1024 × 4096 |
| 6 | added |  | |  | to_channels (Permute) |  | 1 × 4096 × 1024 |
| 7 | added |  | |  | z_gate (Swish) |  | 1 × 1024 × 4096 |
| 8 | added |  | |  | causal_conv (Conv1d) | 16K | 1 × 4096 × 1024 |
| 9 | added |  | |  | to_tokens (Permute) |  | 1 × 1024 × 4096 |
| 10 | added |  | |  | silu_x (Swish) |  | 1 × 1024 × 4096 |
| 11 | added |  | |  | ssm_scan (Mamba) | 51M | 1 × 1024 × 4096 |
| 12 | added |  | |  | gate_out (Multiply) |  | 1 × 1024 × 4096 |
| 13 | added |  | |  | out_proj (Linear) |  | 1 × 1024 × 1024 |
| 14 | added |  | |  | residual_1 (Add) |  | 1 × 1024 × 1024 |
| 15 | added |  | |  | norm_ffn (Rms Norm) | 1.0K | 1 × 1024 × 1024 |
| 16 | added |  | |  | ffn (Swiglu) | 6.3M | 1 × 1024 × 1024 |
| 17 | added |  | |  | residual_2 (Add) |  | 1 × 1024 × 1024 |
| 18 | changed (outFeatures) | Linear (Linear) |  | 128 × 10 | lm_head (Linear) |  | 1 × 1024 × 50280 |
| 19 | same | Output (Output) |  | 128 × 10 | output (Output) |  | 1 × 1024 × 50280 |

## What this is not

- The two are priced at different declared inputs (128 × 300 against 1 × 1024), so memory, cost and GPU fit are each right about their own model and are not a comparison between them. The layer and parameter deltas are unaffected.
- Parameter counts are derived from the graph, not read from a checkpoint. They are exact for a graph that is fully specified and approximate for one that is not.
- Cost and GPU fit are estimates from the graph under one set of assumptions, not measurements of a run.

## Graphs

- Simple RNN: https://neurarch.com/templates/simple-rnn/model.json
- Mamba SSM Block: https://neurarch.com/templates/mamba-block/model.json
- Check a graph of your own: `POST https://www.neurarch.com/api/v1/plan`
