# GPT-2 vs BERT Base

Decoder-only against encoder-only, same era, same size class.

**BERT Base has 53M fewer parameters than GPT-2: 2 layers added, 3 removed, 5 changed.**

Source: https://neurarch.com/diff/gpt2-vs-bert-base.html

## Sides

| | GPT-2 | BERT Base |
|---|---|---|
| Layers | 10 | 9 |
| Parameters | 84M | 31M |
| Input | 1 × 1024 | 1 × 512 |
| Output | 1 × 1024 × 50257 | 1 × 512 × 768 |
| Forward-passes | yes | yes |
| Est. train cost | $1.82 | $0.197 |
| T4 16GB | fits | fits |
| A100 40GB | fits | fits |
| H100 80GB | fits | fits |

## Deltas (BERT Base relative to GPT-2)

- Parameters: -53M (-63.1%)
- Layers: -1
- Added 2, removed 3, changed 5, unchanged 4

## Layer by layer

| # | Status | GPT-2 | Params | Output | BERT Base | Params | Output |
|---|---|---|---|---|---|---|---|
| 1 | changed (shape) | tokens (Input) |  | 1 × 1024 | input_ids (Input) |  | 1 × 512 |
| 2 | changed (numEmbeddings) | token_embed (Embedding) | 39M | 1 × 1024 × 768 | word_embed (Embedding) | 23M | 1 × 512 × 768 |
| 3 | changed (maxLen) | pos_embed (Positional Encoding) |  | 1 × 1024 × 768 | pos_embed (Positional Encoding) |  | 1 × 512 × 768 |
| 4 | same | ln_1 (Layer Norm) | 1.5K | 1 × 1024 × 768 | embed_norm (Layer Norm) | 1.5K | 1 × 512 × 768 |
| 5 | removed | attn (Causal Attention) | 2.4M | 1 × 1024 × 768 |  | |  |
| 6 | removed | residual_1 (Add) |  | 1 × 1024 × 768 |  | |  |
| 7 | added |  | |  | embed_drop (Dropout) |  | 1 × 512 × 768 |
| 8 | added |  | |  | self_attn (Multi Head Attention) | 2.4M | 1 × 512 × 768 |
| 9 | same | ln_2 (Layer Norm) | 1.5K | 1 × 1024 × 768 | norm (Layer Norm) | 1.5K | 1 × 512 × 768 |
| 10 | changed (hiddenDim, embedDim) | mlp (Feed Forward) | 4.7M | 1 × 1024 × 768 | dense (Feed Forward) | 4.7M | 1 × 512 × 768 |
| 11 | removed | residual_2 (Add) |  | 1 × 1024 × 768 |  | |  |
| 12 | same | ln_f (Layer Norm) | 1.5K | 1 × 1024 × 768 | norm (Layer Norm) | 1.5K | 1 × 512 × 768 |
| 13 | changed (outFeatures, inFeatures) | lm_head (Linear) |  | 1 × 1024 × 50257 | dense (Linear) | 591K | 1 × 512 × 768 |
| 14 | same | logits (Output) |  | 1 × 1024 × 50257 | cls_embedding (Output) |  | 1 × 512 × 768 |

## What this is not

- The two are priced at different declared inputs (1 × 1024 against 1 × 512), so memory, cost and GPU fit are each right about their own model and are not a comparison between them. The layer and parameter deltas are unaffected.
- Parameter counts are derived from the graph, not read from a checkpoint. They are exact for a graph that is fully specified and approximate for one that is not.
- Cost and GPU fit are estimates from the graph under one set of assumptions, not measurements of a run.

## Graphs

- GPT-2: https://neurarch.com/templates/gpt2/model.json
- BERT Base: https://neurarch.com/templates/bert-base/model.json
- Check a graph of your own: `POST https://www.neurarch.com/api/v1/plan`
