JumalAIta IntelligenceIndependent evaluation

Benchmarks

Benchmarks

Each benchmark has its own page with the leaderboard, the recorded outputs, the verbatim prompt, the evaluation protocol and its limitations. Benchmarks are versioned, and every status is labeled.

Benchmark 01 / Creative intelligence

Expert-rankedActive

Creative Reimplementation Benchmark

One reference. One prompt. A complete audiovisual work. Each model receives a video capture of a 2001 demoscene production and rebuilds it for the web in a single shot. The original authors of the demo rank the results.

Standings

  1. 1Claude Fable 5.1Anthropic
  2. 2GPT-6 AstraOpenAI
  3. 3Claude Opus 4.7Anthropic

Ranked by the original authors of the reference demo (members of Jumalauta). A ranking only: no scores.

In development

More benchmarks in development.

New benchmarks are announced in the changelog when they are published.