Building this in public

One person, a data factory, and a rule: every claim ships with the receipt that proves it. Posts happen when something real happens.

2026-07-20

I ran the same recipe on a 32B. It made the model worse.

Same frozen data, same script, model swapped 7B→32B: overall fix-rate DOWN 7.4 points, medium tier significantly damaged. Every number published — negative results are results.

2026-07-19

I retrained my own headline number four times — here is every result

The published run said 17.1% → 40.0%. Three re-runs of the same frozen recipe said 34.3%, 34.3%, 22.9%. Why I'm publishing the spread nobody else publishes.