benchmark

Prime Agent on ARC-AGI-3: reported score and caveats

Prime Intellect reports Opus 5 in Prime Agent at 95.5% RHAE Best@1 on public ARC-AGI-3, with three runs of 95.0, 95.2, and 95.5, and a linked median scorecard. The result is self-reported, not independently reproduced here, and public/private eval rule compatibility remains unresolved.

v v0.7.1reviewed 2026-08-09evidence self-reportedsources S001, S037, S036, S035cutoff 2026-08-09
Benchmarks

Benchmark or case-study claim made by Prime Intellect without independent reproduction in this corpus.

Published claim

Prime Intellect reports a 95.5% RHAE Best@1 score for Opus 5 in Prime Agent on public ARC-AGI-3. This page repeats that claim as a self-reported figure, not as independently reproduced evidence.

Linked scorecard

The linked scorecard shows three runs at 95.0, 95.2, and 95.5. Use the scorecard to verify the citation, but not to infer independent reproduction.

Reported runs

The published run bundle does not include all prompts, hardware details, or variance analysis. That makes the result useful as a claim, but incomplete as a reproducibility artifact.

Cross-harness comparison limits

Public and private ARC-AGI-3 rule sets may not transfer cleanly, so the benchmark is not a universal cross-harness proof. Use it as a bounded claim with source attribution only.

Missing evidence

  • raw prompts,
  • run configuration,
  • hardware,
  • and variance data.