Frozen gold suites

Benchmarks

Reproducible recipe-generation results scored against frozen references. The headline is only the start: inspect structure-specific performance and the bounded evidence behind every member outcome.

Read the methodology and caveats

No benchmark runs are published yet

This page will compare reviewed model runs after a maintainer executes the frozen gold suite locally, validates the canonical result, and deliberately commits that small result document. No placeholder scores are shown while benchmarks/results/ is empty.

Read the benchmark methodology for the frozen suite, metrics, revision policy, contamination caveat, and reproduction command.