Methodology
Methodology
JumalAIta Intelligence is an independent evaluator. This page describes what every benchmark on this site has in common. The task, the prompt, the evaluation protocol and the limitations of an individual benchmark are documented on its own page.
Principles
- Identical conditions
- Every model in a benchmark receives the same prompt and the same inputs, and runs under the same stated settings: how many attempts it gets, which reasoning effort is used and whether a human may follow up. The conditions are listed on the benchmark page, and a model is never reported without its configuration.
- Published verbatim
- The exact prompt is printed on the benchmark page, word for word, including its typos and formatting. Anyone can copy it and run it against a new model.
- Evidence on the page
- Model outputs are recorded and published next to the standings. A reader who disagrees with a judgment should be able to see exactly what was judged.
- Gaps stay visible
- What has not been published is labeled "Not yet published", "TBA" or "Preliminary". Nothing is estimated, rounded up or filled in. A benchmark that produces a ranking is shown as a ranking, not dressed up with scores.
What a benchmark page contains
Every benchmark is documented on a single page, results first, in the same order:
- Leaderboard with the benchmark version, the status and the date of the last update.
- Key takeaways: what the current results do and do not show.
- Recordings of every evaluated output, and the evaluators' notes where they exist.
- The task and the prompt, with the prompt reproduced verbatim.
- Evaluation protocol: who evaluates and how, the conditions shared by all models and a list of what has and has not been published.
- Limitations and a FAQ that takes the uncomfortable questions first.
- Version history and a citation entry.
Published benchmarks: Creative Reimplementation Benchmark.
Models and configurations
A model name alone does not identify what was evaluated. Results are always reported together with the configuration of the run: how the output was generated (for example single shot) and the reasoning effort that was used. The same model evaluated under a different configuration is a separate result.
Run details such as the harness, the run date, the cost and the token usage have a place on every model page. Where they have not been published, the page says so.
Who evaluates, and how
Each benchmark names its evaluators and its evaluation method on its own page. The method decides what the leaderboard shows, and the site does not convert one kind of result into another.
- Expert-ranked
- Named experts place the outputs in order of preference as whole works. The result is an ordinal ranking only: no scores, no sub-scores and no confidence intervals. A ranking says which output the evaluators preferred. It does not say by how much.
- Scored
- Outputs receive numerical scores. Scores are reported with the number of evaluators and, where available, a 95% confidence interval next to the value. No scored benchmark has been published yet.
The Creative Reimplementation Benchmark is expert-ranked, and its evaluators are the original authors of the reference work. Evaluation guidelines commonly exclude the authors of a work from judging it. Here the choice is deliberate and stated openly: the authors know the intention and the cultural context of the work better than anyone else, and the benchmark measures how well a model understands exactly that. According to the authors, their order did not follow visual polish, which is the kind of distinction this choice is meant to capture. The price is an author-centred, subjective view, and the benchmark page says so under Limitations.
Every benchmark states how many runs each model gets. Where a single run per model is used, there are no retries and no best-of-N selection.
Status labels
- Preliminary
- Some part of the result has not yet been published, for example part of the ranking. Preliminary standings may change.
- Active
- The result for the current version is published in full, and more models may be added in later versions.
- Frozen
- The benchmark is no longer updated. Its results stay online as published.
In tables, TBA marks a placement that has not been announced, and "Not yet published" marks a value that exists or is planned but is not public. Models without an announced placement are listed alphabetically, which implies no order.
Versioning
Every benchmark carries a version number, and every result records the version it was produced under.
- Major (2.0)
- The prompt, the inputs or the evaluation method change. Results are not comparable across major versions.
- Minor (1.1)
- New models or a new evaluation round are added under an unchanged task.
- Patch (1.0.1)
- Corrections that do not change the task: fixed errors, clarified wording, corrected metadata.
Version histories are kept on the benchmark pages. Changes to the site as a whole are recorded in the changelog.
Scope
A benchmark result describes how a model performed on that benchmark, under that configuration, at that time. It is not a measure of general capability, and standings in one benchmark say nothing about another.
Reference material that is publicly available online may be present in a model's training data. Where that applies, the benchmark page says so under Limitations.