article methodology

Benchmark methodology and reproducibility limits

Benchmark claims on this site are attributed to Prime Intellect unless an independent artifact supports stronger wording. Public sources lack complete prompts, model snapshots, run counts, variance, hardware, and raw outputs for most evaluations. This page documents what exists, what is unresolved, and how the site updates claims.

v v0.7.1reviewed 2026-08-09evidence self-reportedsources S001, S037, S038cutoff 2026-08-09
Benchmarks

Benchmark or case-study claim made by Prime Intellect without independent reproduction in this corpus.

Reproducibility standard

Reproducibility here means enough information to rerun the benchmark under the same conditions and compare the outputs honestly. Most launch claims do not yet meet that threshold.

What is disclosed

The corpus does include references to benchmark names, some scorecards, and linked articles. That is still not the same thing as a fully disclosed evaluation pipeline.

What is missing

The missing fields are the ones that matter most for trust: prompts, seed control, model snapshots, hardware, raw outputs, and variance.

Scorecards and artifacts

Scorecards help with citation, but not necessarily with re-running the experiment. Treat them as evidence of a reported claim, not a full verification bundle.

Update policy

If Prime Intellect later publishes independent artifacts, this site should upgrade the claim class only when the new evidence is actually inspectable and version-stamped.