GPT-6 Astra Benchmark Results Show Discrepancies, Efficiency Gains

2026-09-04

OpenAI's GPT-6 Astra is receiving conflicting evaluations from AI benchmark platforms. While one assessment places it ahead of competitors, another rates it similarly to its predecessor, lagging behind another model.

VERA Brief

AI-generated. Grounded in the article and its cited sources.

OpenAI's GPT-6 Astra has produced conflicting benchmark results, with one assessment placing it ahead of competitors and another rating it similarly to its predecessor. A notable development is Astra's increased efficiency on the ARC-AGI-3 benchmark, surpassing average human performance.

Key facts

  • OpenAI's GPT-6 Astra has received divergent benchmark results from different AI evaluation platforms.
  • Epoch AI reported Astra achieving 169 points, positioning it as a leading model.
  • Artificial Analysis rated Astra no better than its predecessor and behind Claude Fable 5.1.
  • On the ARC-AGI-3 benchmark, Astra demonstrated greater efficiency than average human performance for the first time.
  • The varying benchmark performances raise questions about consistent measurement and verification of AI system capabilities.

Source: The Decoder

Reported by VERA Newswire.

More from September 2026 in The Record.