Code-repair training data where every example proves itself by running.

Most training data is graded by another model's opinion. Mine is graded by execution: the broken program has to fail a hidden checker, the fixed program has to pass the same checker, both runs verified before an example is allowed to exist. No AI judges anywhere in the correctness path — and every result on this site ships with the receipts to back it.

Where frontier models miss Get in touch
17% → 23–40%
hard-tier fix-rate, tuned 7B — four runs, full spread published
0
AI judges in the correctness path
every run
published with its full spread, not just the headline

Support the research

This is one person paying for GPUs and API calls out of pocket. If you want more of this kind of work — verified data, published spreads, results you can read in full — you can fund it directly. Donations aren't earmarked: they fund whatever experiment is next on the blog, and every result is published whichever way it goes, wins and losses alike.

What your support buys is more published results — another training run, another frontier comparison, another batch of execution-verified data — each one shipped with the receipts to check it yourself.

Support on Ko-fi

The proof, not the pitch

How it's made (the short honest version)

Faults are constructed backwards from working code, so the fix is known and provable. Every checker gets attacked with wrong-but-plausible code before its task can ship — if a bad fix can pass, the task is cut. On multi-file faults, fixing only one file provably still fails, so partial credit is impossible by construction. The generation pipeline itself is the one thing I don't publish; the proofs are the product, and you can run those.

Code repair is where this starts. The method — construct the fault backwards from a known-good answer, then prove the fix by execution — isn't specific to code; it works anywhere correctness can be run rather than judged. That's the direction. Code repair is just the first domain with the receipts to show it.

Share this: X Hacker News Reddit LinkedIn Email