OpenAI altered several benchmark figures for its GPT-6 Astra model after publishing its launch announcement, with some revisions making Astra appear stronger and temporarily reducing the reported scores of rival Anthropic models.
The changes emerged during an unusually delayed release on Thursday, September 3. OpenAI had intended to publish the announcement at 2pm Eastern Time, but the page was withdrawn and did not become consistently accessible until almost two hours later.
The company’s X account shared the post at 3.32pm, although many users received an error message. OpenAI chief executive Sam Altman later wrote that the company had “hit a little snag getting the blog post deployed”, while the firm gave varying explanations for the disruption, including a content-management problem and an internet outage.
OpenAI said it could not disclose why the post had been retracted, but insisted the reason was unrelated to the benchmark results. When the announcement reappeared, several figures had changed, and some continued to move afterwards.
GPT-6 Astra benchmark figures changed after launch
One of the most significant revisions concerned Astra’s reported hallucination rate. An early version listed the figure at 4.2 per cent, before a later version reduced it to 2 per cent. The corresponding figure for GPT-5.6 Sol, Astra’s predecessor, fell from 12.2 per cent to 9.4 per cent.
Both figures subsequently returned to their original levels. OpenAI’s official Astra materials now list a range of evaluation results across coding, mathematics, cybersecurity and other tasks, with performance dependent on the model configuration and testing conditions used.
The company also increased GPT-5.6 Sol’s result on its internal ExploitBench cybersecurity test from 5.5 per cent to 11.5 per cent. OpenAI said it was investigating whether to reverse that change because the higher score reflected a reasoning setting that is not commercially available for Sol.
Astra’s mathematics score on the FrontierMath Tier 4 evaluation remained at 97.6 per cent. However, the published scores for Anthropic’s Fable 5.1 and GPT-5.6 Sol changed during the same period, briefly making Astra’s lead appear wider.
Fable 5.1’s score moved from 87.8 per cent to 78 per cent before returning to 83 per cent. Sol’s figure shifted from 83 per cent to 80.5 per cent and then back to 83 per cent.
The figures had already changed between an embargoed draft and the first public version. Astra’s score on ARC-AGI-3 rose from 98.6 per cent in the draft to 99.99 per cent in the live announcement, although the benchmark’s own published results show performance varies sharply according to the “harness” — the set of tools and interface through which the model operates.
Under ARC Prize’s standard harness, Astra’s best observed ARC-AGI-3 score was 62.7 per cent, while its provider-adapted harness produced a result of 99.9 per cent. The organisation said the two tests answer different questions about model capability rather than providing a straightforward like-for-like comparison.
OpenAI says revisions reflect testing conditions
OpenAI said benchmark results can vary by several percentage points depending on the model checkpoint, the tools provided, the amount of reasoning allowed and the precise evaluation run.
“We care deeply about getting evaluations right,” a spokesperson said. “For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons.”
The company’s announcement also warns that its scores represent the maximum result achieved at any level of effort and may not exactly match the version of the model available in the production ChatGPT service.
That explanation has not ended concerns among independent researchers about what some in the industry call “benchmaxxing” — repeatedly running evaluations under different conditions in order to obtain the most favourable score.
Anka Reuel and Mike Hardy, researchers at Stanford, said benchmark results could be rerun quickly and that the practice could serve marketing purposes. They also argued that Astra’s system card gave too little information about some tests, including the internal hallucination evaluation.
Vincent Sunn Chen, an AI engineer at Snorkel AI, said changes shortly before a model launch were not unusual because the checkpoint, configuration, harness and grading system were often still being finalised. He said companies should explain what had changed whenever they revised a published score.
Not every revision benefited OpenAI. Scores for Anthropic’s Claude Fable 5.1 and Opus 5 on the HealthBench Professional evaluation increased in later versions of the material, while Astra’s coding result rose only marginally from 57.7 per cent to 57.9 per cent.
The episode highlights the difficulty of comparing rapidly evolving AI models through a single set of headline figures. Benchmark scores are used by companies to demonstrate progress and attract customers, but differences in tools, computing time and evaluation methods can make apparently precise rankings difficult to interpret.
