JumalAIta IntelligenceIndependent evaluation

Changelog

Changelog

A dated record of what changed on this site: new benchmarks, new results, methodology updates and corrections. Version histories of individual benchmarks are on the benchmark pages.

  1. Benchmark

    Creative Reimplementation Benchmark v1.0

    First publication of the benchmark: three models, one prompt, single shot, ranked by the original authors of the reference demo.

  2. Results

    Full ranking published

    The original authors of the reference demo ranked the three reimplementations: 1. Claude Fable 5.1, 2. GPT-6 Astra, 3. Claude Opus 4.7. Their notes are on the benchmark page.

  3. Results

    Recordings added for three models

    Side-by-side comparison recordings published for Claude Fable 5.1, Claude Opus 4.7 and GPT-6 Astra.

  4. Site

    Site published at ai.jumalauta.org