Building this in public
One person, a data factory, and a rule: every claim ships with the receipt that proves it. Posts happen when something real happens.
2026-07-20
I ran the same recipe on a 32B. It made the model worse.
Same frozen data, same script, model swapped 7B→32B: overall fix-rate DOWN 7.4 points, medium tier significantly damaged. Every number published — negative results are results.
2026-07-19
I retrained my own headline number four times — here is every result
The published run said 17.1% → 40.0%. Three re-runs of the same frozen recipe said 34.3%, 34.3%, 22.9%. Why I'm publishing the spread nobody else publishes.