N Neurarch Architectures Checks Data Docs Open the app

Resources / Datasets

Arch-Bench task set

The task definitions behind the benchmark: design-from-spec and repair-and-extend instances for agents that build neural network architectures. Each task pairs a natural-language design brief with a starting graph the agent edits and a set of programmatic pass criteria (structural blockers, parameter budgets and bands, required layer families on an input-to-output path, KV cache and decode-latency ceilings). Twelve curated tasks with eight starting fixtures, plus a deterministic generator that mints a larger split from a seed.

12 curated tasks, 8 fixtures MIT Free to use The environment Source repo On Hugging Face

Get it

Task definitionsraw.githubusercontent.com/neurarch-ai/neurarch-arch-bench/main/tasks.json
application/json
Hugging Face mirror (dataset viewer, load_dataset)huggingface.co/datasets/neurarch-ai/arch-bench-tasks
text/html
curl -s https://raw.githubusercontent.com/neurarch-ai/neurarch-arch-bench/main/tasks.json | jq '.tasks[] | {id, spec, constraints}'

What is in a row

task id design spec starting graph pass constraints difficulty

What this dataset is not

Twelve curated tasks is a seed, not a set you can rank models on with confidence. The generated split exists because of that, and the numbers worth comparing are the generated-split ones. The curated tasks are best read as the worked examples that show what a task is.

Licence and citation

Released under MIT License. Cite it as:

Neurarch. Arch-Bench task set. https://neurarch.com/d/bench-tasks.html

The rest of the set

Neurarch architecture corpus
36 neural network architectures kept as typed graphs rather than diagrams.
36 architectures · CC0-1.0
Neurarch structural check catalogue
The 41 structural checks Neurarch runs on a model graph, as data: id, severity, category, the condition that triggers it, why it costs something, and the fix.
41 checks · CC-BY-4.0
Verifier grounding study (264 graphs)
Clean reference architectures plus systematically corrupted variants (broken attention head divisibility, linear width mismatches, severed connections), each built as a real PyTorch model and run on a GPU.
264 graphs · MIT
Arch-Bench arena results
Frontier language models scored on design-from-spec tasks by a deterministic verifier rather than a human or an LLM judge.
18 model-split results · CC-BY-4.0
arch-design-sft: verified architecture-design SFT data
Supervised fine-tuning data for neural architecture design treated as structured graph editing.
3,010 verified examples · MIT
Structure, verdict and trained outcome triples
The corpus that pairs what a design looks like with what it did.
80 trained graphs · MIT
Verified architecture-design reasoning traces (Claude)
Spec to reasoning to design triples where the design is re-graded by the same deterministic verifier the benchmark uses, and only passing traces are kept.
306 verified traces · MIT
Verified architecture-design reasoning traces (Grok)
The same verified spec to reasoning to design triples as the Claude split, rejection-sampled from a different frontier model, so the two can be pooled for volume or held apart to see how much of the reasoning style is model-specific.
376 verified traces · MIT