benchmark
Prime Agent on ARC-AGI-3: reported score and caveats
Prime Intellect reports Opus 5 in Prime Agent at 95.5% RHAE Best@1 on public ARC-AGI-3, with three runs of 95.0, 95.2, and 95.5, and a linked median scorecard. The result is self-reported, not independently reproduced here, and public/private eval rule compatibility remains unresolved.
Benchmark or case-study claim made by Prime Intellect without independent reproduction in this corpus.
Published claim
Prime Intellect reports a 95.5% RHAE Best@1 score for Opus 5 in Prime Agent on public ARC-AGI-3. This page repeats that claim as a self-reported figure, not as independently reproduced evidence.
Linked scorecard
The linked scorecard shows three runs at 95.0, 95.2, and 95.5. Use the scorecard to verify the citation, but not to infer independent reproduction.
Reported runs
The published run bundle does not include all prompts, hardware details, or variance analysis. That makes the result useful as a claim, but incomplete as a reproducibility artifact.
Cross-harness comparison limits
Public and private ARC-AGI-3 rule sets may not transfer cleanly, so the benchmark is not a universal cross-harness proof. Use it as a bounded claim with source attribution only.
Missing evidence
- raw prompts,
- run configuration,
- hardware,
- and variance data.