Changelog
Changelog
A dated record of what changed on this site: new benchmarks, new results, methodology updates and corrections. Version histories of individual benchmarks are on the benchmark pages.
- Benchmark
Creative Reimplementation Benchmark v1.0
First publication of the benchmark: three models, one prompt, single shot, ranked by the original authors of the reference demo.
- Results
The original authors of the reference demo ranked the three reimplementations: 1. Claude Fable 5.1, 2. GPT-6 Astra, 3. Claude Opus 4.7. Their notes are on the benchmark page.
- Results
Recordings added for three models
Side-by-side comparison recordings published for Claude Fable 5.1, Claude Opus 4.7 and GPT-6 Astra.
- Site
Site published at ai.jumalauta.org