A recording, not a live demo

One task, start to finish.
Nothing here was typed by hand.

Loading the recording.

What this shows. That the loop closes: designs get proposed, the ones that cannot run are deleted before anything is booked, exactly two questions are asked because they are the only two nothing can derive, one run is paid for, and the result is kept and handed back.

What this does not show. That we know which design trains better. Between two designs that both run, we do not. Our own selection benchmark expresses a preference on 8.3% of pairs and lands at 51.4% pairwise accuracy, which is a coin flip, and two 2026 papers beat that with a frontier model reading the source. The ordering below is legality first, then cost, and cost is our money, not a judgement about the design.