The run log
69 experiments. 23 of them failed, and they are the point.
Every row is a real experiment from the archive: a hypothesis, the config it ran at, the metric that moved, and a verdict. Failed runs are shown by default because the archive kept them — one of its files has a section titled “the graveyard — what failed and why”.
42
shipped or proven
23
killed
7
projects
A number without its configuration is not quotable. A deck once cited one model at F1 77.0 while the reports cited it at 78.5 — same model, same ground truth, same scorer, different inference resolution. Every figure has been stamped since.