01 / Leaderboard
Leaderboard
| Rank | Model | Organization | Configuration | Recording |
|---|---|---|---|---|
| 1 | Claude Fable 5.1Anthropic | Anthropic | Single shot · High | Watch the Claude Fable 5.1 recording on YouTube |
| 2 | GPT-6 AstraOpenAI | OpenAI | Single shot · High | Watch the GPT-6 Astra recording on YouTube |
| 3 | Claude Opus 4.7Anthropic | Anthropic | Single shot · High | Watch the Claude Opus 4.7 recording on YouTube |
- Ranking: an order of preference set by the original authors of the reference demo (members of Jumalauta). This benchmark produces a ranking only. There are no scores, by design.
02 / Key takeaways
Key takeaways
- The original authors of the demo ranked the three reimplementations: 1. Claude Fable 5.1, 2. GPT-6 Astra, 3. Claude Opus 4.7.
- The order did not follow visual polish. According to the original authors, the GPT-6 Astra version looks better at first glance, but Claude Fable 5.1 understood the essence of the demo, and they weighted that above surface quality.
- The benchmark measures interpretation and an understanding of the spirit of a work, not visual quality alone. It produces a ranking by design: there are no scores.
- All three models produced a full-length, working audiovisual work from a single prompt.
- Every evaluated output is published as a side-by-side recording against the original, so readers can check the judgment for themselves.
03 / Recordings
Recordings
04 / Evaluators' notes
Evaluators' notes
The original authors gave the following reasons for their order. They are reported here as the authors' views, not as findings of this site. No other evaluation comments have been published.
- Rank 1
Claude Fable 5.1
According to the original authors, Claude Fable 5.1 understood the essence of the demo. Their example is the credits, where the model credited itself under an invented handle in the style of the group's own credits. The output is fully procedural, as its own credits state.
gfx: taas yks tekoaely joka luulee osaavansa piirtaeae
gfx: yet another AI that thinks it can drawRecording, about 1:28 gfx: ei yhtaeaen kuvaa, kaikki proseduraalista
gfx: not a single image, everything proceduralRecording, about 1:40 - Rank 2
GPT-6 Astra
According to the original authors, the GPT-6 Astra version looks better at first glance: the model used image generation models for its assets, and the output contains four generated images. The authors weighted an understanding of the spirit of the work above visual polish.
- Rank 3
Claude Opus 4.7
According to the original authors, the Claude Opus 4.7 version looks noticeably rougher than the two newer entries.
05 / The task
The task
Models receive the same video capture and rebuild the work for the web. The prompt deliberately leaves room for interpretation: composition, effects, assets and creative choices belong to the model.
Kiinalaisen saatanan palvontademo
- Released
- 2001-06-16, Proxy 20016th place, combined demo competition
- Reference capture
- 169.62 s (2 min 50 s)
- Target platform
- HTML5 Canvas / WebGL
- Original soundtrack
- Permitted
- Code, assets and effects
- AI-generated
06 / Prompt
Prompt
So here's an old demoscene demo. what i would like is that you'd recreate this demo in demo_capture.mp4 rules: - use the capture as reference but it can be inspired-by/your interpretation partially and not frame-by-frame accurate - music can be used as is but code and assets and effects need to be generated by AI - technology: Web (HTML5 Canvas/WebGL) - demo's length must match the original's length
demo_capture.mp4Same reference capture for every modelThis is the complete prompt, reproduced verbatim. Every model receives exactly this text together with the input file, and nothing else.
To test another model, copy the prompt and supply the same reference capture.
A benchmark kit containing the prompt, the reference capture, the protocol and the tooling used to produce the comparison recordings exists. It is not yet public; publication is planned.
07 / Evaluation protocol
Evaluation protocol
Every model receives the same two inputs: the prompt, reproduced verbatim on this page, and the reference
capture demo_capture.mp4. The model produces its reimplementation in a single shot at High reasoning
effort, in a vanilla harness setup: the default configuration, with no customisation. There is no human
follow-up: nobody corrects, steers or re-prompts the model after the initial prompt. A single run per model
is used. No retries, no best-of-N selection.
The evaluators are the original authors of the reference demo: members of Jumalauta, the group that made it in 2001. This is the core of the benchmark and a deliberate design decision. The authors know the intention and the cultural context of the work better than anyone else, and the benchmark measures how well a model understands exactly that.
The authors place the reimplementations in order of preference, judging each as a whole work: interpretation, creativity and implementation together. A frame-by-frame reconstruction is not required. Interpretation is explicitly permitted by the prompt: original credits or invented creator names, for example, can be intentional creative decisions within the work and are treated as such.
The result is an ordinal ranking. By design there are no scores, no sub-scores and no confidence intervals.
Each output is documented as a side-by-side comparison recording: the original on the left, the model's version on the right, with the audio taken from the original. The recordings are 1920x1080, open with a 5-second title card and run about 2 min 55 s.
Publication status
- Prompt
- Published verbatim
- Reference work
- Published
- Recordings
- 3 of 3 models
- Ranking
- 3 of 3 models placed
- Evaluators' notes
- 3 of 3 models
- Scores
- None, by design
- Evaluators
- The original authors of the reference demo (members of Jumalauta)
- Number, names and roles of evaluators
- Not published
- Harness names, run dates, cost and tokens
- Not yet published
- Benchmark kit
- Not yet public
Shared conditions
- Generation
- Single shotOne initial prompt. No subsequent human direction.
- Reasoning effort
- HighExtra-high and max settings are excluded.
- Harness setup
- VanillaDefault configuration of the harness, no customisation. Harness names and versions have not been published.
- Human follow-up
- NoneThe model's first complete output is the evaluated output.
- Runs per model
- OneA single run per model is used. No retries, no best-of-N selection.
- Implementation
- HTML5 Canvas / WebGLThe work must run on the web.
- Code, visual assets and effects
- AI-generatedEverything except the soundtrack must be produced by the model.
- Original soundtrack
- PermittedThe music of the reference may be reused as is.
- Duration
- Must match the referenceThe reference capture file is 169.62 seconds long.
- Time, cost and token budgets
- Not specifiedVersion 1.0 of the protocol sets no budget limits.
08 / Limitations
Limitations
- An author-centred view, by design
- The evaluators are the original authors of the reference demo. Their ranking reflects how well each output captures the work as they understand it. It is a subjective view and not a general verdict on quality.
- A small group of evaluators
- The evaluators are a small group. Their number, names and roles have not been published, and the evaluation is not claimed to be blind.
- An ordinal ranking only
- The result is an order of preference. It does not say how large the differences between the works are, and there is no score to compare across versions or benchmarks.
- One task, one run per model
- The benchmark consists of a single audiovisual task built on a single reference work, and each model was run once. Run-to-run variance has not been measured, so a different sample from the same model could look different.
- Three models
- Only three models have been evaluated. The ranking says nothing about models that were not run.
- Possible training data contamination
- The reference demo has been publicly available online for years. It may be present in the training data of the evaluated models.
- Not a measure of general capability
- The results describe performance on this benchmark only. They say nothing about how the models compare elsewhere.
09 / FAQ
FAQ
- Isn't this just taste?
- Yes, and deliberately so. The outputs are creative works, and the benchmark asks how well a model understood the work it was given. That is a judgment, not a measurement. The benchmark makes the judgment inspectable instead: the prompt, the conditions and every recording are on this page.
- Why do the original authors judge? Isn't that a conflict of interest?
- It is the central design decision of this benchmark. The authors know the intention and the cultural context of the original better than anyone else, and understanding exactly that is what the benchmark measures. The price is that the result is an author-centred, subjective view and not a general verdict on quality. According to the authors, their order did not follow visual polish, which is the kind of distinction this choice of evaluators is meant to capture.
- Why is there no score?
- The evaluators put the works in order of preference as whole works. There is no rubric, no sub-scores and no confidence interval, and none are planned for this benchmark. A ranking says which work the authors preferred. It does not say by how much.
- Is the original demo in the training data?
- Possibly. The demo was released in 2001 and has been publicly available online since. This cannot be ruled out for any of the evaluated models, and it is listed under Limitations.
- Why only three models?
- These are the models evaluated so far. Any model can be run with the same prompt, the same reference capture and the same conditions.
- Was each model run more than once?
- No. A single run per model is used. No retries, no best-of-N selection.
10 / Version history
Version history
- v1.0Active
- Initial publication with three evaluated models.
- Side-by-side comparison recordings published for every model.
- Full ranking by the original authors published, with the authors' notes.
11 / How to cite
How to cite
Cite the benchmark together with its version. Results from different major versions are not comparable.
@misc{jumalaita2026creativereimplementation,
author = {{JumalAIta Intelligence}},
title = {Creative Reimplementation Benchmark},
year = {2026},
url = {https://ai.jumalauta.org/benchmarks/creative-reimplementation/},
note = {Version 1.0}
}The principles that apply to every JumalAIta Intelligence benchmark are described in the general methodology. Changes to this site are recorded in the changelog.



