article methodology
Benchmark methodology and reproducibility limits
Benchmark claims on this site are attributed to Prime Intellect unless an independent artifact supports stronger wording. Public sources lack complete prompts, model snapshots, run counts, variance, hardware, and raw outputs for most evaluations. This page documents what exists, what is unresolved, and how the site updates claims.
Benchmark or case-study claim made by Prime Intellect without independent reproduction in this corpus.
Reproducibility standard
Reproducibility here means enough information to rerun the benchmark under the same conditions and compare the outputs honestly. Most launch claims do not yet meet that threshold.
What is disclosed
The corpus does include references to benchmark names, some scorecards, and linked articles. That is still not the same thing as a fully disclosed evaluation pipeline.
What is missing
The missing fields are the ones that matter most for trust: prompts, seed control, model snapshots, hardware, raw outputs, and variance.
Scorecards and artifacts
Scorecards help with citation, but not necessarily with re-running the experiment. Treat them as evidence of a reported claim, not a full verification bundle.
Update policy
If Prime Intellect later publishes independent artifacts, this site should upgrade the claim class only when the new evidence is actually inspectable and version-stamped.