AIThe Decoder1h ago

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency

TL;DRGPT-6 Astra shows mixed benchmark results but excels at ARC-AGI-3, sparking AGI timeline debates.

Why it matters: Conflicting performance data complicates AI capability assessment and shapes when experts think AGI arrives.

OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. The biggest surprise comes from ARC-AGI-3, where Astra works more…

Read full article

Source: The Decoder · Opens in new tab