JumalAIta IntelligenceIndependent evaluation

Benchmark 01 / Creative intelligence

Creative Reimplementation Benchmark

v1.0ActiveUpdated · 3 models · Expert-ranked

One reference. One prompt. A complete audiovisual work. Each model receives a video capture of a 2001 demoscene production and rebuilds it for the web in a single shot. The original authors of the demo rank the results.

Measures How a model interprets, creates and implements a complete audiovisual work from one reference and one prompt.

Generation
Single shot
Reasoning effort
High
Harness setup
Vanilla
Human follow-up
None

01 / Leaderboard

Leaderboard

v1.0 · Expert-ranked
Creative Reimplementation Benchmark v1.0 standings, active.Ranking: an order of preference set by the original authors of the reference demo (members of Jumalauta). This benchmark produces a ranking only. There are no scores, by design.
RankModelOrganizationConfigurationRecording
1Claude Fable 5.1AnthropicAnthropicSingle shot · HighWatch the Claude Fable 5.1 recording on YouTube
2GPT-6 AstraOpenAIOpenAISingle shot · HighWatch the GPT-6 Astra recording on YouTube
3Claude Opus 4.7AnthropicAnthropicSingle shot · HighWatch the Claude Opus 4.7 recording on YouTube
  • Ranking: an order of preference set by the original authors of the reference demo (members of Jumalauta). This benchmark produces a ranking only. There are no scores, by design.

02 / Key takeaways

Key takeaways

  • The original authors of the demo ranked the three reimplementations: 1. Claude Fable 5.1, 2. GPT-6 Astra, 3. Claude Opus 4.7.
  • The order did not follow visual polish. According to the original authors, the GPT-6 Astra version looks better at first glance, but Claude Fable 5.1 understood the essence of the demo, and they weighted that above surface quality.
  • The benchmark measures interpretation and an understanding of the spirit of a work, not visual quality alone. It produces a ranking by design: there are no scores.
  • All three models produced a full-length, working audiovisual work from a single prompt.
  • Every evaluated output is published as a side-by-side recording against the original, so readers can check the judgment for themselves.

03 / Recordings

Recordings

In rank order. Side by side: the original on the left, the model's version on the right.

Anthropic · Single shot · High reasoning

04 / Evaluators' notes

Evaluators' notes

The original authors gave the following reasons for their order. They are reported here as the authors' views, not as findings of this site. No other evaluation comments have been published.

  1. According to the original authors, Claude Fable 5.1 understood the essence of the demo. Their example is the credits, where the model credited itself under an invented handle in the style of the group's own credits. The output is fully procedural, as its own credits state.

    gfx: taas yks tekoaely joka luulee osaavansa piirtaeae
    gfx: yet another AI that thinks it can drawRecording, about 1:28
    gfx: ei yhtaeaen kuvaa, kaikki proseduraalista
    gfx: not a single image, everything proceduralRecording, about 1:40
  2. According to the original authors, the GPT-6 Astra version looks better at first glance: the model used image generation models for its assets, and the output contains four generated images. The authors weighted an understanding of the spirit of the work above visual polish.

  3. According to the original authors, the Claude Opus 4.7 version looks noticeably rougher than the two newer entries.

05 / The task

The task

Original referenceWatch on YouTube

Models receive the same video capture and rebuild the work for the web. The prompt deliberately leaves room for interpretation: composition, effects, assets and creative choices belong to the model.

Kiinalaisen saatanan palvontademo

Jumalauta · Proxy 2001 · Windows

Released
2001-06-16, Proxy 20016th place, combined demo competition
Reference capture
169.62 s (2 min 50 s)
Target platform
HTML5 Canvas / WebGL
Original soundtrack
Permitted
Code, assets and effects
AI-generated

Release archive on Demozoo

06 / Prompt

Prompt

The benchmark prompt, verbatim
So here's an old demoscene demo. what i would like is that you'd recreate this demo in demo_capture.mp4
rules:
- use the capture as reference but it can be inspired-by/your interpretation partially and not frame-by-frame accurate
- music can be used as is but code and assets and effects need to be generated by AI
- technology: Web (HTML5 Canvas/WebGL)
- demo's length must match the original's length
Inputdemo_capture.mp4Same reference capture for every model

This is the complete prompt, reproduced verbatim. Every model receives exactly this text together with the input file, and nothing else.

To test another model, copy the prompt and supply the same reference capture.

A benchmark kit containing the prompt, the reference capture, the protocol and the tooling used to produce the comparison recordings exists. It is not yet public; publication is planned.

07 / Evaluation protocol

Evaluation protocol

Every model receives the same two inputs: the prompt, reproduced verbatim on this page, and the reference capture demo_capture.mp4. The model produces its reimplementation in a single shot at High reasoning effort, in a vanilla harness setup: the default configuration, with no customisation. There is no human follow-up: nobody corrects, steers or re-prompts the model after the initial prompt. A single run per model is used. No retries, no best-of-N selection.

The evaluators are the original authors of the reference demo: members of Jumalauta, the group that made it in 2001. This is the core of the benchmark and a deliberate design decision. The authors know the intention and the cultural context of the work better than anyone else, and the benchmark measures how well a model understands exactly that.

The authors place the reimplementations in order of preference, judging each as a whole work: interpretation, creativity and implementation together. A frame-by-frame reconstruction is not required. Interpretation is explicitly permitted by the prompt: original credits or invented creator names, for example, can be intentional creative decisions within the work and are treated as such.

The result is an ordinal ranking. By design there are no scores, no sub-scores and no confidence intervals.

Each output is documented as a side-by-side comparison recording: the original on the left, the model's version on the right, with the audio taken from the original. The recordings are 1920x1080, open with a 5-second title card and run about 2 min 55 s.

Publication status

Prompt
Published verbatim
Reference work
Published
Recordings
3 of 3 models
Ranking
3 of 3 models placed
Evaluators' notes
3 of 3 models
Scores
None, by design
Evaluators
The original authors of the reference demo (members of Jumalauta)
Number, names and roles of evaluators
Not published
Harness names, run dates, cost and tokens
Not yet published
Benchmark kit
Not yet public

Shared conditions

Generation
Single shotOne initial prompt. No subsequent human direction.
Reasoning effort
HighExtra-high and max settings are excluded.
Harness setup
VanillaDefault configuration of the harness, no customisation. Harness names and versions have not been published.
Human follow-up
NoneThe model's first complete output is the evaluated output.
Runs per model
OneA single run per model is used. No retries, no best-of-N selection.
Implementation
HTML5 Canvas / WebGLThe work must run on the web.
Code, visual assets and effects
AI-generatedEverything except the soundtrack must be produced by the model.
Original soundtrack
PermittedThe music of the reference may be reused as is.
Duration
Must match the referenceThe reference capture file is 169.62 seconds long.
Time, cost and token budgets
Not specifiedVersion 1.0 of the protocol sets no budget limits.

08 / Limitations

Limitations

An author-centred view, by design
The evaluators are the original authors of the reference demo. Their ranking reflects how well each output captures the work as they understand it. It is a subjective view and not a general verdict on quality.
A small group of evaluators
The evaluators are a small group. Their number, names and roles have not been published, and the evaluation is not claimed to be blind.
An ordinal ranking only
The result is an order of preference. It does not say how large the differences between the works are, and there is no score to compare across versions or benchmarks.
One task, one run per model
The benchmark consists of a single audiovisual task built on a single reference work, and each model was run once. Run-to-run variance has not been measured, so a different sample from the same model could look different.
Three models
Only three models have been evaluated. The ranking says nothing about models that were not run.
Possible training data contamination
The reference demo has been publicly available online for years. It may be present in the training data of the evaluated models.
Not a measure of general capability
The results describe performance on this benchmark only. They say nothing about how the models compare elsewhere.

09 / FAQ

FAQ

Isn't this just taste?
Yes, and deliberately so. The outputs are creative works, and the benchmark asks how well a model understood the work it was given. That is a judgment, not a measurement. The benchmark makes the judgment inspectable instead: the prompt, the conditions and every recording are on this page.
Why do the original authors judge? Isn't that a conflict of interest?
It is the central design decision of this benchmark. The authors know the intention and the cultural context of the original better than anyone else, and understanding exactly that is what the benchmark measures. The price is that the result is an author-centred, subjective view and not a general verdict on quality. According to the authors, their order did not follow visual polish, which is the kind of distinction this choice of evaluators is meant to capture.
Why is there no score?
The evaluators put the works in order of preference as whole works. There is no rubric, no sub-scores and no confidence interval, and none are planned for this benchmark. A ranking says which work the authors preferred. It does not say by how much.
Is the original demo in the training data?
Possibly. The demo was released in 2001 and has been publicly available online since. This cannot be ruled out for any of the evaluated models, and it is listed under Limitations.
Why only three models?
These are the models evaluated so far. Any model can be run with the same prompt, the same reference capture and the same conditions.
Was each model run more than once?
No. A single run per model is used. No retries, no best-of-N selection.

10 / Version history

Version history

  1. v1.0Active
    • Initial publication with three evaluated models.
    • Side-by-side comparison recordings published for every model.
    • Full ranking by the original authors published, with the authors' notes.

11 / How to cite

How to cite

Cite the benchmark together with its version. Results from different major versions are not comparable.

BibTeX
@misc{jumalaita2026creativereimplementation,
  author = {{JumalAIta Intelligence}},
  title  = {Creative Reimplementation Benchmark},
  year   = {2026},
  url    = {https://ai.jumalauta.org/benchmarks/creative-reimplementation/},
  note   = {Version 1.0}
}

The principles that apply to every JumalAIta Intelligence benchmark are described in the general methodology. Changes to this site are recorded in the changelog.