Frozen gold suites
Benchmarks
Reproducible recipe-generation results scored against frozen references. The headline is only the start: inspect structure-specific performance and the bounded evidence behind every member outcome.
No benchmark runs are published yet
This page will compare reviewed model runs after a maintainer executes
the frozen gold suite locally, validates the canonical result, and
deliberately commits that small result document. No placeholder scores
are shown while benchmarks/results/ is empty.
Read the benchmark methodology for the frozen suite, metrics, revision policy, contamination caveat, and reproduction command.