Benchmark 01 / Creative intelligence
Expert-rankedActiveCreative Reimplementation Benchmark
One reference. One prompt. A complete audiovisual work. Each model receives a video capture of a 2001 demoscene production and rebuilds it for the web in a single shot. The original authors of the demo rank the results.
Standings
- 1Claude Fable 5.1Anthropic
- 2GPT-6 AstraOpenAI
- 3Claude Opus 4.7Anthropic
Ranked by the original authors of the reference demo (members of Jumalauta). A ranking only: no scores.