OpenAI’s announcement of the GPT-6 Astra achieving a 99.9% AGI benchmark score has been complicated by analysis showing the same model scores 62.7% under a standard harness. The supporting software structure around the AI, not just the model itself, substantially affects the results.
- 99.9% AGI score derived from OpenAI’s specialized harness, not the core model alone.
- Standard ARC benchmark harness scores Astra at 62.7%, still far above previous models.
- ARC Prize and experts caution this observed performance does not yet prove true AGI.
What happened
OpenAI launched GPT-6 Astra boasting a headline 99.9% score on the ARC Prize benchmark, a test designed to evaluate reasoning capabilities expected of artificial general intelligence. However, when ARC Prize ran Astra under its own standard testing harness—the software and environment that control how the model interacts with the test—the score was only 62.7%.
This significant difference is due to the harness, which manages factors such as memory between questions, access to tools, and context handling. OpenAI’s specialized Provider Adapter harness keeps and compresses reasoning state between questions, enabling scores and efficiency well beyond the minimal interface used in the standard harness. The latter is what ARC Prize and many outside researchers consider the baseline for meaningful comparison.
Why it matters
The dramatic discrepancy highlights how benchmark results can be influenced not only by the AI model but also by the surrounding software framework. This calls for cautious interpretation of headline AGI claims based on benchmark scores, emphasizing the need for transparency in evaluation methods.
Despite the score gap, even the standard harness result of 62.7% is a large leap over previous models, such as Anthropic’s Sol at 7.8%, marking real progress in reasoning ability. Yet ARC Prize co-founder and independent experts stress there is insufficient evidence to definitively declare this model as achieving true AGI at this stage.
What to watch next
The ARC Prize foundation plans to continue publishing results from both OpenAI’s specialized harness and the standard harness side by side to improve clarity around future model evaluations. This dual-reporting aims to prevent confusion and allow more nuanced benchmarking discussions.
Industry observers and AI governance stakeholders will be monitoring how OpenAI and competitors handle transparency in performance reporting, especially as new models push the boundaries of reasoning and generalization. The debate over what constitutes AGI benchmarks and responsible disclosure practices will remain central to the field’s progression.