The system of record
for model structure.
Every change: planned, approved, kept.

Models are now changed by agents as often as by people. Git records the text and the experiment tracker records the run; neither knows what the model is, or can enforce a policy before a GPU is dispatched. From the moment a weight file or a repository is brought in, Neurarch plans every change to it, takes the approval, and keeps what actually happened when it ran: what it is, whether it runs, what it costs on which card, whether it is allowed here, what this structure did the last time it trained here. The same record is served to a person in the app, to an agent over MCP and the API, and to CI on every model pull request.

40 popular checkpoints reconciled against their weights: 15 of 28 comparable within 2%, from config alone · Live since early May 2026 · neurarch.com
Pre-seed · Pitch Deck 2026

Team

I've already shipped this pattern in production,
at AWS, for cloud architectures.

XG Xin Gao
Xin Gao
Founder
Neurarch is the second time I'm shipping this thesis

At AWS I led an agentic LLM that converted natural-language cloud requirements into validated AWS architecture diagrams via typed-schema function-calling. 5× faster Solutions Architect workflow.

Neurarch applies the same pattern to neural networks: the model is a typed graph, every change is a schema-validated edit, and the approval is computed rather than argued.

Career arc

Meta · Research Scientist, LLM foundation models for recommendation (current)
AI startup · Founding Lead Scientist · solo-shipped LLM vuln triage in 10 wks, led 5-person team
AWS AI · Applied Scientist · 5 yrs · Bedrock GenAI, agentic LLM NL→AWS-architecture system

Supporting record

2
US patents
granted
308
Scholar
citations
ICML
2023
workshop
PhD
CS
NJIT
21d
solo
MVP
195
ML layer
types
43
design-time
structural checks
7
framework
export

Solo by design, to ship at a team's velocity (full-stack MVP in 21 days, then shipped solo nights and weekends).

2

The Problem

A model change is written in seconds.
Approving one still takes a senior engineer's afternoon.

Git records the text. The tracker records the run. Nothing in the stack knows what the model is, whether this change is allowed here, or what this structure did the last time it ran.

2020 2026 cost Cost to approve and record a model change a senior engineer's afternoon, or a GPU run to find out. Unchanged. Cost to produce a model change an agent writes one in seconds the gap we sell into widens with every change an agent makes
What it is

A weight file or a repository comes in and nothing derives what is inside it, or whether it matches what it claims. We derive it from the file: on 40 of the most-downloaded checkpoints, 15 of the 28 comparable ones reproduce their published parameter count within 2% from config alone.

Whether it is allowed here

The budget ceiling, the card whitelist and the size limit live in a wiki. Nothing enforces them at the moment a GPU is about to be dispatched. The plan carries the organisation's policy, and a refusal names the rule it came from.

What happened last time

Nothing joins a structure to its verdict, its spend and its outcome. The ledger does, per organisation: runs dispatched through Neurarch land in it, and a run on your own cluster joins it when the exported script reports under your key, one environment variable at run time, never a credential in the file.

"Whether it runs" is the cheap half, and it is real: in an Alibaba study of 12,289 failed training jobs, tensor-shape errors were the second most common framework-specific failure. One forward pass finds it, so it is the floor of the record, not the reason to keep one.

3

The Product

The agent lives in the pull request.
The gate is deterministic. The ledger is ours.

model-ci-example · pull request #1 · neurarch bot commented
neurarch-bot · plan for models/tiny_gpt.py:TinyGPT
params +178.4K (+295%) · cost +<$0.01 · blast radius 4 layers
Found 2 structural issues in this PR:
⛔ head-dim-divisibility (block)
models/tiny_gpt.py:21
MultiheadAttention has embed_dim=384, num_heads=5. head_dim would be 76.80 (must be an integer).
⚠ softmax-cross-entropy (warn)
models/tiny_gpt.py:56
nn.CrossEntropyLoss applies LogSoftmax internally; an explicit Softmax before it double-applies.
✕ neurarch-bot failed · 1 blocking · the graph does not forward-pass
# next, in progress: a fix branch opened by the bot, and /neurarch train on request
  • ›
    The bot, in the pull request
    neurarch-ai/neurarch-bot: traces base and head in your CI, posts one plan per model, turns the check red on a blocker. Two files to install.
  • ›
    One command, one card
    neurarch-trace … --plan --share: the same plan in the terminal, with a link you can paste into Slack.
  • ›
    The same verifier for other agents
    neurarch-mcp for Claude Code and Cursor, POST /api/v1/check and /plan for anything else, the Arch-Bench environment for RL.
  • ›
    The ledger, per organisation
    Every plan and every run outcome is a row keyed to the structure. The next change is checked against what happened last time here. This is the line no model call can write.

The editor at neurarch.com is the reference client: it proves the verifier is real and generates corpus rows. It is not the product.

4

In CI

The gate sits in the pull request,
in front of the GPU.

pull request #212 · checks
neurarch-lint
Found 3 structural issues in this PR:
⛔ head-dim-divisibility (block)
models/encoder.py:18
MultiheadAttention has embed_dim=384, num_heads=5. head_dim would be 76.80 (must be an integer).
⛔ groupnorm-channel-divisibility (block)
models/encoder.py:24
GroupNorm has num_channels=16, num_groups=3. num_channels must be divisible by num_groups.
⚠ softmax-no-dim (warn)
models/head.py:9
Softmax called without an explicit dim. Pass dim= (usually dim=-1).
✕ neurarch-lint failed · exit code 1 · 2 blocking
ruff, flake8 and mypy all pass this file

It imports cleanly. It crashes once the module is built on a GPU. The fix in the README example is one character, and the Action finds it from source, with no Python install and no model loaded.

Same text locally
$ npx neurarch-lint --markdown --dir=.

A block fails the check. A warn is informational unless fail-on-warn is set. SARIF output lands findings in the Security tab and inline on the diff.

Why CI and not the editor

ML engineers live in pip and GitHub. A gate that runs where the code already goes needs no one to open a new tool, and every catch is a row in the corpus behind the verifier.

5

Why Now

Generation got a hundred times cheaper.
Verification did not move at all.

⚡
The customer becomes a machine

Agents write model code now, and an agent asks for a verdict a thousand times a day. Demand scales with token volume, not with the number of ML engineers.

💸
The mistake costs a run

The unit of a wrong architecture is not a bug fix, it is $10K to $10M of GPU time. Nothing else in the stack checks it before the spend.

🧭
Why it opens now, not in 2019

Code became an agent domain because it already had a compiler and tests to close the loop against. Architecture design had neither, so agents cannot self-correct there. Building the verifier is what makes the domain agent-addressable at all.

Every discipline that spends real money per artifact built a verification layer first. Software has the compiler, silicon has formal verification, a hundred-billion-dollar industry. Neural architecture spends the most per mistake and has nothing.
Our thesis
6

Market Size

Seats are the floor, not the ceiling.
A seat is priced against a software budget. A verdict is priced against the run it prevents.

3.5M
ML / AI engineers worldwide, LinkedIn + GitHub ML authors + HuggingFace, 2026
1.2M
Actively design architectures, not just fine-tuning off-the-shelf
600K
At SaaS-paying companies, addressable buyers
$2.2M
per 1% of the 600K seats at $30/mo blended ARPU. Single-digit penetration = $10 to $20M ARR.
Proven interest

33K GitHub stars on Netron, a read-only ML inspection tool. Stars are softer than dollars, but 33,000 engineers wanted half of what we ship badly enough to star it.

Comparable: Copilot

~$400M ARR (2024) on a similar per-IC-developer seat motion. Younger, faster-growing, higher-ACV category than ours.

5-year expansion

+400K academic ML researchers, +500K Fortune-5000 data-science seats. $1.5B+ TAM by year 5.

Second, non-seat path: licensing the environment to frontier labs training design agents. Environment deals in this market run $1-10M/yr per lab against roughly 10-20 credible buyers today.

Third, and the one that sets the ceiling: a verdict metered in front of a training job is drawn from the compute budget, which is two orders of magnitude larger than the software budget and growing faster than any other line in the industry.

7

Business Model

Freemium SaaS with a
self-served Pro tier.

Free
$0 /mo
  • Typed-graph editor + 28 SOTA templates
  • PyTorch / Keras export
  • Bring-your-own LLM key
  • Public sharing
Pro
$19 /mo
  • Everything in Free
  • Hosted AI agent, DeepSeek V3 (500/mo)
  • HuggingFace direct import
  • FastAPI & cross-framework export
  • Private models + version history
Team / Enterprise
Contact
  • SSO, audit log, RBAC
  • Real-time collaboration
  • On-prem / VPC deploy
  • Custom layer plugin SDK
  • SLA + dedicated support
End of Year 1
$30K MRR

~1.3K Pro · 50 paying teams · Netron outreach + waitlist conversion

Year 2
$1M ARR

~3K Pro/Plus · 10 Team customers · seed round

Year 3
$4M ARR

~10K Pro · 30 Team / Enterprise

Year 5
$20M+ ARR

~50K Pro · 100+ Enterprise · metered machine calls and environment licences

Stripe live (off pending pricing study) · Modal.com GPU backend integrated · Supabase auth shipped.

8

Traction · since launch

Live since May. 22 accounts, all organic.
The MCP server is the first external pull.

22
registered accounts
zero paid acquisition
~2,300
npm downloads of neurarch-mcp
last month
15/28
comparable popular checkpoints
reproduced within 2% from config
22 registered accounts since 2026-04-26, all organic. 9 ML engineers recruited one per company, sending feedback on a <24h feedback-to-ship cycle (7 fixes in one day on a researcher's V4 neuron paper).
Arch-Bench, our live arena, runs frontier models on architecture tasks with anti-gaming graders: Claude Opus 5 currently leads the board. The harness and its GRPO training loop are open source and runnable; the held-out split is not published.
9 engineers I recruited, one per company
Google DeepMindResearcher
AppleSoftware Engineer
AmazonApplied Scientist
AdobeEngineer
ZooxEngineer
AbbVieBiostat Manager
NJIT / GA TechCS PhD

"If the full paper-to-runnable-code path lands end-to-end, this tool is unbeatable."

CS PhD researcher, NJIT · after I shipped 7 fixes in 24h on his V4 neuron paper
Stated plainly, because we would rather you hear it from us

Of the machine-facing surfaces (developer API, MCP server, CI Action, GitHub App, self-hosted container), only the MCP server shows measurable external pull, and we cannot yet tell a recruited caller from one who found us. External Arch-Bench submissions: zero. Paying customers: zero; Stripe is wired and tested end to end, switched off pending pricing. The argument here is that the record demonstrably accumulates, not that distribution is solved. It is not, and it is the first thing this round buys.

9

Go-to-market

Open-source wedge → self-serve Pro
→ team expansion.

The free tool is the funnel. Pull before push: 20 signups, 9 recruited engineers across 7 companies in active feedback, $0 marketing.

🪝
Free wedge, zero CAC

neurarch-lint (source-available CLI) + MCP server + open SOTA templates run in any PyTorch repo, CI, or Claude. Each bug they catch is a reason to open the hosted app.

💳
Self-serve conversion

Free web app → Pro $19 when they want the hosted agent, private models, and cross-framework export. Paywalls already live behind Stripe.

🌱
Land & expand

Individual seat → team. Engineers at Apple, AbbVie and others already use it individually; we expand seat-by-seat from the accounts we are already inside.

🔁
Compounding loop

Every model designed = a labeled architecture. Exported figures and public templates carry a "Made with Neurarch" mark back to the top of funnel.

What this round lets us nail
Instrument

CLI → web → Pro conversion: the one metric that sets CAC

Ship the hook

In-CLI "open in Neurarch" CTA on every caught bug

Convert teams

Design-partner accounts → first paid seats

Amplify

OSS template drops on HF · X · r/MachineLearning

10

Competition

No one else ships a typed-graph contract
with an agent that respects it.

Typed Graph Live FLOPs AI Agent Code Export Cross-FW Convert Cloud Launch
Netronview-only, , , , ,
TensorBoardpost-hoc✓, , , ,
Excalidraw / Miro, , , , , ,
HuggingFace Hub, , , ✓PT↔TF,
Cursor / Copilot, , ✓text¹, ,
MMdnn (MS, archived), , , ✓7 FW,
SkyPilot, , , , , ✓
Neurarch✓✓✓✓7 FW✓

Cursor writes code. Neurarch designs architecture. Complementary, but the typed-graph + agent + ML-domain combo is uncontested.

¹ Cursor edits source files; it does not export model architectures or propagate tensor shapes through a typed graph.

11

Positioning

The empty quadrant: design-time intelligence.

Structural / Graph-native
ML-Domain Aware
High
Low
Low
High
Design-time
intelligence
Mermaid
Excalidraw / Miro
AWS Step Functions
Cursor / Copilot
SkyPilot
MMdnn
Netron (read-only)
TensorBoard
HuggingFace Hub
Neurarch

Tools that show your model are read-only. Tools you can edit don't know it's a model. Neurarch is a typed graph you can edit.

12

Objection: "Doesn't HuggingFace already show model structure?"

Seeing a model isn't designing one.

HuggingFace renders a model that already exists, read-only. Neurarch sits one step upstream, where the model is still being built.

HuggingFace
Store · Distribute · Display

✓ Find a pre-built model
✓ Read-only structure view
✓ Hosting & inference compute
✕ Can't edit the graph
✕ No pre-training bug checks
✕ No runnable glue code

Neurarch
Design · Catch · Generate

✓ An editable typed graph
✓ Catch bugs before the GPU run
dim mismatch · param blow-up · cycles
✓ Emit runnable code
train · eval · deploy
→ then push to HuggingFace to host

Before a model exists · Neurarch ───────► HuggingFace · after a model exists

We're HuggingFace's on-ramp and complement, not its replacement.

13

Objection: "The labs will build this themselves."

Yes, for their own models.
The buyer is everyone else running a research agent.

What a lab builds
One organisation · one model family · one corpus

A preference model inside its own agent, trained on its own runs. Foster et al. (33 authors, 2026) built exactly that inside AIRA-dojo, for AIRA-dojo: 69.35% pairwise, reading code with three frontier models.
It sees that lab's designs and that lab's outcomes. Nobody else's.

What we sell
The long tail · every organisation · one corpus

Teams running AIDE or AIRA-dojo style research agents without a lab behind them will not build a verifier. They call one.
Every call pairs a structure with a verified outcome, across organisations. No single lab sees that breadth.
Measured 2026-09-01: a frontier model reading our exported code ranks 15 held-out pairs at 73 to 80%; always picking the larger design scores 73% on the same pairs. Our static score abstains on 11. The corpus is the only route past those numbers, and it is the one nobody else is collecting.

Weights & Biases

System of record for experiments across about 1,400 organisations. Sold to CoreWeave for about $1.7B. Not for an algorithm: for the record.

Braintrust

Every lab evaluates its own LLM apps. Evals still became the bottleneck for everyone else, and a company.

Snyk, Sonar

Verification gets big when it sits on the flow of every artifact, not when it is cleverer than the in-house check.

Against a frontier lab, no startup has a data moat on that lab's designs. Across labs, nobody has one at all. That is the seat.

14

Moat

Copy the rules in a week.
The other three take the loop.

The 43 checks are the part a funded team clones fastest, and we say so. What they would still not have: the substrate, a typed graph the params, the memory, the GPU fit, the cost and the blast radius are computed from rather than judged; the round trip, a verifier and a training path in one loop, which is the only way a static score can ever be measured against a trained outcome; and the ledger, one organisation's structures and what they each trained to, which compounds and which no model has in its weights.

24 designs. All passed the verifier. All trained on identical configs. The static score cannot tell them apart.

1 · Structure paired with outcome

These pairs exist only where one party owns the verifier and the training round trip. Labs have code and papers. Nobody is collecting this.

2 · It mines rules from outcomes

One structural feature mined from completed runs moved the in-sample correlation from 0.17 to 0.41 at zero GPU cost. Out of sample (2026-09-01, 15 held-out designs) it moved 0.015, and the score declined to separate 11 of 15 pairs. Two frontier models reading the same code decided all 15 at 73 to 80%; so did a rule with no model in it, always pick the larger design, at 73%. The corpus is the lever. The first rule is not yet it. All of it is published.

3 · Integrated head start (6 to 12 mo)

195 layer types · 43 structural checks · agent on a typed graph · export to 7 frameworks · free GPU round trip. A copy starts at zero on all of it.

15

The next 90 days

Three results, each dated and falsifiable.
No new features.

Days 1 to 30 · calibration, out of sample · ran 2026-09-01

Five held-out tasks, two held-out designers, 15 designs trained under identical configs with graphs persisted. The mined rule moved the correlation by 0.015. The score declined to separate 11 of 15 pairs of legal designs it had never seen.

Exit, met: the number is published, including the part that did not lift, and the baselines it has to beat: Opus 5 reading the same code, 12 of 15, and "pick the larger design", 11 of 15. The next campaign fits a ranking signal on more rows instead of shipping more rules.

Days 30 to 60 · machine traffic

Distribute and meter the API, the MCP server and the CI Action. Publish the share of verifier calls originating from agents rather than browsers, weekly.

Exit: a live ratio, and the first outside repo running our gate in front of its own GPU spend.

Days 60 to 90 · the RL result

Train a small open model inside our environment with GRPO and measure it on the held-out Arch-Bench split against frontier models. Opus 5 leads the board today, so the baseline is already standing.

Exit: a curve showing a trained small model beating a frontier model at architecture design.

Why the third one is the whole thesis

One result proves three things at once: the environment's reward is real signal rather than noise, the signal can be trained into a model, and the signal exists nowhere else. Every part is already built: the environment, the GRPO loop, the leaderboard, and approved managed-GPU capacity.

Soundness is settled and is not the headline: across 264 graphs, 96 of 96 designs the verifier blocked crashed in PyTorch forward and 80 of 80 it passed ran clean; 24/24 verifier-passed designs trained end to end on managed GPUs. A forward pass gives an agent the same answer in seconds. The unsolved half, and the one above, is which legal design deserves the run.

16

The ladder

Five rungs. We are on the second.
Each one has a precondition we can name.

L1 · humans L2 · machines we are here L3 · ranking L4 · the designer L5 · the record
L1 · verified design for humans

Sells seats.
True today. Live, and the verdicts are measured against 264 real runs.

L2 · verified design for machines

Sells calls and CI gates.
Built, never distributed. This is what the round buys.

L3 · ranking, not just legality

Sells which design deserves the GPU.
Needs a score that decides pairs. The first out-of-sample campaign abstained on 11 of 15.

L4 · the model that designs models

Sells the designer, or the environment that trains one.
Needs the RL result and corpus scale.

L5 · the record every model carries

Sells structural provenance at the point of spend.
Needs L2 adoption wide enough to be the default path.

We are raising on rung two, with rung three inside the round. Four and five are stated so you can hold us to the preconditions, not so you can price them today.

17

The verification layer for machine-designed models.

A typed-graph contract for ML architectures, verified in milliseconds, called by humans and increasingly by machines.

The long-term bet: pairs of architecture structure and verified training outcome exist only where someone owns both the verifier and the training round trip. Once that corpus ranks reliably, the verdict stops merely legalising a design and starts pricing it, and a verdict in front of a training job becomes the thing a compute budget is released against.

Built solo, nights and weekends, while working full-time · 20 signups, all organic.

Raising pre-seed
Funding the distribution of the machine-facing surfaces, the out-of-sample calibration campaign, and the RL result. Happy to share the details and discuss what fits.

neurarch.com / pitch

Live product demo · investor data room available on request.
Xin Gao · xin.gao.njit@gmail.com

18