benchmark
Prime Agent on the long-context benchmark suite
The launch article compares Prime Agent against native harnesses across OOLONG, OBLIQ-Bench, LongBench, ManyIH, LongCoT-Mini, and EmulatorBench. Prime Agent wins many cells but loses or ties others. The comparison offloaded context to files and used official competitor figures where local reproductions underperformed; raw run bundles are not published.
Benchmark or case-study claim made by Prime Intellect without independent reproduction in this corpus.
Published table
The launch table compares Prime Agent against other harnesses on OOLONG, OBLIQ-Bench, LongBench, ManyIH, LongCoT-Mini, and EmulatorBench. It is a useful snapshot, not a complete scientific reproduction package.
Where Prime Agent loses
Prime Agent does not win every row. Where it loses or ties, the published table is still valuable because it marks the boundary of the harness advantage instead of hiding the misses.
Methodology disclosed
The article at least states that some competitor figures were reused and that context was offloaded to files where appropriate. That is better than opaque marketing, but still not enough for full reproducibility.
Methodology missing
Missing prompts, budgets, model snapshots, and hardware details make apples-to-apples comparison impossible from the public record alone.