Models are now changed by agents as often as by people. Git records the text and the experiment tracker records the run; neither knows what the model is, or can enforce a policy before a GPU is dispatched. From the moment a weight file or a repository is brought in, Neurarch plans every change to it, takes the approval, and keeps what actually happened when it ran: what it is, whether it runs, what it costs on which card, whether it is allowed here, what this structure did the last time it trained here. The same record is served to a person in the app, to an agent over MCP and the API, and to CI on every model pull request.
At AWS I led an agentic LLM that converted natural-language cloud requirements into validated AWS architecture diagrams via typed-schema function-calling. 5× faster Solutions Architect workflow.
Neurarch applies the same pattern to neural networks: the model is a typed graph, every change is a schema-validated edit, and the approval is computed rather than argued.
Meta · Research Scientist, LLM foundation models for recommendation (current)
AI startup · Founding Lead Scientist · solo-shipped LLM vuln triage in 10 wks, led 5-person team
AWS AI · Applied Scientist · 5 yrs · Bedrock GenAI, agentic LLM NL→AWS-architecture system
Solo by design, to ship at a team's velocity (full-stack MVP in 21 days, then shipped solo nights and weekends).
Git records the text. The tracker records the run. Nothing in the stack knows what the model is, whether this change is allowed here, or what this structure did the last time it ran.
A weight file or a repository comes in and nothing derives what is inside it, or whether it matches what it claims. We derive it from the file: on 40 of the most-downloaded checkpoints, 15 of the 28 comparable ones reproduce their published parameter count within 2% from config alone.
The budget ceiling, the card whitelist and the size limit live in a wiki. Nothing enforces them at the moment a GPU is about to be dispatched. The plan carries the organisation's policy, and a refusal names the rule it came from.
Nothing joins a structure to its verdict, its spend and its outcome. The ledger does, per organisation: runs dispatched through Neurarch land in it, and a run on your own cluster joins it when the exported script reports under your key, one environment variable at run time, never a credential in the file.
"Whether it runs" is the cheap half, and it is real: in an Alibaba study of 12,289 failed training jobs, tensor-shape errors were the second most common framework-specific failure. One forward pass finds it, so it is the floor of the record, not the reason to keep one.
neurarch-ai/neurarch-bot: traces base and head in your CI, posts one plan per model, turns the check red on a blocker. Two files to install.neurarch-trace … --plan --share: the same plan in the terminal, with a link you can paste into Slack.neurarch-mcp for Claude Code and Cursor, POST /api/v1/check and /plan for anything else, the Arch-Bench environment for RL.The editor at neurarch.com is the reference client: it proves the verifier is real and generates corpus rows. It is not the product.
It imports cleanly. It crashes once the module is built on a GPU. The fix in the README example is one character, and the Action finds it from source, with no Python install and no model loaded.
A block fails the check. A warn is informational unless fail-on-warn is set. SARIF output lands findings in the Security tab and inline on the diff.
ML engineers live in pip and GitHub. A gate that runs where the code already goes needs no one to open a new tool, and every catch is a row in the corpus behind the verifier.
Agents write model code now, and an agent asks for a verdict a thousand times a day. Demand scales with token volume, not with the number of ML engineers.
The unit of a wrong architecture is not a bug fix, it is $10K to $10M of GPU time. Nothing else in the stack checks it before the spend.
Code became an agent domain because it already had a compiler and tests to close the loop against. Architecture design had neither, so agents cannot self-correct there. Building the verifier is what makes the domain agent-addressable at all.
33K GitHub stars on Netron, a read-only ML inspection tool. Stars are softer than dollars, but 33,000 engineers wanted half of what we ship badly enough to star it.
~$400M ARR (2024) on a similar per-IC-developer seat motion. Younger, faster-growing, higher-ACV category than ours.
+400K academic ML researchers, +500K Fortune-5000 data-science seats. $1.5B+ TAM by year 5.
Second, non-seat path: licensing the environment to frontier labs training design agents. Environment deals in this market run $1-10M/yr per lab against roughly 10-20 credible buyers today.
Third, and the one that sets the ceiling: a verdict metered in front of a training job is drawn from the compute budget, which is two orders of magnitude larger than the software budget and growing faster than any other line in the industry.
~1.3K Pro · 50 paying teams · Netron outreach + waitlist conversion
~3K Pro/Plus · 10 Team customers · seed round
~10K Pro · 30 Team / Enterprise
~50K Pro · 100+ Enterprise · metered machine calls and environment licences
Stripe live (off pending pricing study) · Modal.com GPU backend integrated · Supabase auth shipped.
"If the full paper-to-runnable-code path lands end-to-end, this tool is unbeatable."
Of the machine-facing surfaces (developer API, MCP server, CI Action, GitHub App, self-hosted container), only the MCP server shows measurable external pull, and we cannot yet tell a recruited caller from one who found us. External Arch-Bench submissions: zero. Paying customers: zero; Stripe is wired and tested end to end, switched off pending pricing. The argument here is that the record demonstrably accumulates, not that distribution is solved. It is not, and it is the first thing this round buys.
The free tool is the funnel. Pull before push: 20 signups, 9 recruited engineers across 7 companies in active feedback, $0 marketing.
neurarch-lint (source-available CLI) + MCP server + open SOTA templates run in any PyTorch repo, CI, or Claude. Each bug they catch is a reason to open the hosted app.
Free web app → Pro $19 when they want the hosted agent, private models, and cross-framework export. Paywalls already live behind Stripe.
Individual seat → team. Engineers at Apple, AbbVie and others already use it individually; we expand seat-by-seat from the accounts we are already inside.
Every model designed = a labeled architecture. Exported figures and public templates carry a "Made with Neurarch" mark back to the top of funnel.
CLI → web → Pro conversion: the one metric that sets CAC
In-CLI "open in Neurarch" CTA on every caught bug
Design-partner accounts → first paid seats
OSS template drops on HF · X · r/MachineLearning
| Typed Graph | Live FLOPs | AI Agent | Code Export | Cross-FW Convert | Cloud Launch | |
|---|---|---|---|---|---|---|
| Netron | view-only | , | , | , | , | , |
| TensorBoard | post-hoc | ✓ | , | , | , | , |
| Excalidraw / Miro | , | , | , | , | , | , |
| HuggingFace Hub | , | , | , | ✓ | PT↔TF | , |
| Cursor / Copilot | , | , | ✓ | text¹ | , | , |
| MMdnn (MS, archived) | , | , | , | ✓ | 7 FW | , |
| SkyPilot | , | , | , | , | , | ✓ |
| Neurarch | ✓ | ✓ | ✓ | ✓ | 7 FW | ✓ |
Cursor writes code. Neurarch designs architecture. Complementary, but the typed-graph + agent + ML-domain combo is uncontested.
¹ Cursor edits source files; it does not export model architectures or propagate tensor shapes through a typed graph.
Tools that show your model are read-only. Tools you can edit don't know it's a model. Neurarch is a typed graph you can edit.
HuggingFace renders a model that already exists, read-only. Neurarch sits one step upstream, where the model is still being built.
✓ Find a pre-built model
✓ Read-only structure view
✓ Hosting & inference compute
✕ Can't edit the graph
✕ No pre-training bug checks
✕ No runnable glue code
✓ An editable typed graph
✓ Catch bugs before the GPU run
dim mismatch · param blow-up · cycles
✓ Emit runnable code
train · eval · deploy
→ then push to HuggingFace to host
We're HuggingFace's on-ramp and complement, not its replacement.
A preference model inside its own agent, trained on its own runs. Foster et al. (33 authors, 2026) built exactly that inside AIRA-dojo, for AIRA-dojo: 69.35% pairwise, reading code with three frontier models.
It sees that lab's designs and that lab's outcomes. Nobody else's.
Teams running AIDE or AIRA-dojo style research agents without a lab behind them will not build a verifier. They call one.
Every call pairs a structure with a verified outcome, across organisations. No single lab sees that breadth.
Measured 2026-09-01: a frontier model reading our exported code ranks 15 held-out pairs at 73 to 80%; always picking the larger design scores 73% on the same pairs. Our static score abstains on 11. The corpus is the only route past those numbers, and it is the one nobody else is collecting.
System of record for experiments across about 1,400 organisations. Sold to CoreWeave for about $1.7B. Not for an algorithm: for the record.
Every lab evaluates its own LLM apps. Evals still became the bottleneck for everyone else, and a company.
Verification gets big when it sits on the flow of every artifact, not when it is cleverer than the in-house check.
Against a frontier lab, no startup has a data moat on that lab's designs. Across labs, nobody has one at all. That is the seat.
The 43 checks are the part a funded team clones fastest, and we say so. What they would still not have: the substrate, a typed graph the params, the memory, the GPU fit, the cost and the blast radius are computed from rather than judged; the round trip, a verifier and a training path in one loop, which is the only way a static score can ever be measured against a trained outcome; and the ledger, one organisation's structures and what they each trained to, which compounds and which no model has in its weights.
24 designs. All passed the verifier. All trained on identical configs. The static score cannot tell them apart.
These pairs exist only where one party owns the verifier and the training round trip. Labs have code and papers. Nobody is collecting this.
One structural feature mined from completed runs moved the in-sample correlation from 0.17 to 0.41 at zero GPU cost. Out of sample (2026-09-01, 15 held-out designs) it moved 0.015, and the score declined to separate 11 of 15 pairs. Two frontier models reading the same code decided all 15 at 73 to 80%; so did a rule with no model in it, always pick the larger design, at 73%. The corpus is the lever. The first rule is not yet it. All of it is published.
195 layer types · 43 structural checks · agent on a typed graph · export to 7 frameworks · free GPU round trip. A copy starts at zero on all of it.
Five held-out tasks, two held-out designers, 15 designs trained under identical configs with graphs persisted. The mined rule moved the correlation by 0.015. The score declined to separate 11 of 15 pairs of legal designs it had never seen.
Exit, met: the number is published, including the part that did not lift, and the baselines it has to beat: Opus 5 reading the same code, 12 of 15, and "pick the larger design", 11 of 15. The next campaign fits a ranking signal on more rows instead of shipping more rules.
Distribute and meter the API, the MCP server and the CI Action. Publish the share of verifier calls originating from agents rather than browsers, weekly.
Exit: a live ratio, and the first outside repo running our gate in front of its own GPU spend.
Train a small open model inside our environment with GRPO and measure it on the held-out Arch-Bench split against frontier models. Opus 5 leads the board today, so the baseline is already standing.
Exit: a curve showing a trained small model beating a frontier model at architecture design.
One result proves three things at once: the environment's reward is real signal rather than noise, the signal can be trained into a model, and the signal exists nowhere else. Every part is already built: the environment, the GRPO loop, the leaderboard, and approved managed-GPU capacity.
Soundness is settled and is not the headline: across 264 graphs, 96 of 96 designs the verifier blocked crashed in PyTorch forward and 80 of 80 it passed ran clean; 24/24 verifier-passed designs trained end to end on managed GPUs. A forward pass gives an agent the same answer in seconds. The unsolved half, and the one above, is which legal design deserves the run.
Sells seats.
True today. Live, and the verdicts are measured against 264 real runs.
Sells calls and CI gates.
Built, never distributed. This is what the round buys.
Sells which design deserves the GPU.
Needs a score that decides pairs. The first out-of-sample campaign abstained on 11 of 15.
Sells the designer, or the environment that trains one.
Needs the RL result and corpus scale.
Sells structural provenance at the point of spend.
Needs L2 adoption wide enough to be the default path.
We are raising on rung two, with rung three inside the round. Four and five are stated so you can hold us to the preconditions, not so you can price them today.
A typed-graph contract for ML architectures, verified in milliseconds, called by humans and increasingly by machines.
The long-term bet: pairs of architecture structure and verified training outcome exist only where someone owns both the verifier and the training round trip. Once that corpus ranks reliably, the verdict stops merely legalising a design and starts pricing it, and a verdict in front of a training job becomes the thing a compute budget is released against.
Built solo, nights and weekends, while working full-time · 20 signups, all organic.
neurarch.com / pitch
Live product demo · investor data room available on request.
Xin Gao · xin.gao.njit@gmail.com