hub

Prime Agent benchmarks: self-reported result ledger

Prime Intellect reported ARC-AGI-3, long-context, EmulatorBench, PMPP-Hard, Factorio, and MazeBench results in its launch article. Most numbers are self-reported and lack publicly available prompts, raw runs, variance, or hardware details. This hub lists what is inspectable, what is missing, and links to methodology.

v v0.7.1reviewed 2026-08-09evidence self-reportedsources S001, S037cutoff 2026-08-09
Home

Benchmark or case-study claim made by Prime Intellect without independent reproduction in this corpus.

Self-reported results

SuiteClaimEvidence classGap
ARC-AGI-395.5% RHAE Best@1Self-reportedNo raw runs or variance published
Long-context suiteMixed win/lose rowsSelf-reportedSome competitor figures are reused
MethodologyN/AMethodology pageMissing prompts and hardware on most rows

Available results

SuiteClaimEvidence classGap
ARC-AGI-395.5% RHAE Best@1Self-reportedNo raw runs or variance published
Long-context suiteMixed win/lose rowsSelf-reportedSome competitor figures are reused
MethodologyN/AMethodology pageMissing prompts and hardware on most rows

Missing evidence

SuiteClaimEvidence classGap
ARC-AGI-395.5% RHAE Best@1Self-reportedNo raw runs or variance published
Long-context suiteMixed win/lose rowsSelf-reportedSome competitor figures are reused
MethodologyN/AMethodology pageMissing prompts and hardware on most rows

Methodology and next step

Use the methodology page before you compare results. If you need a public citation, start from the source browser and follow the self-reported claim chain rather than quoting a single number in isolation.