AIFortune3h ago
OpenAI quietly updates its evaluation metrics for GPT-6 Astra, making
OpenAI quietly updates its evaluation metrics for GPT-6 Astra, making changes that appear to favor Astra and continuing to revise other metrics after launch

TL;DROpenAI adjusted GPT-6 benchmarks after launch, raising questions about metric reliability.
Why it matters: Benchmark changes post-launch undermine confidence in model comparisons and competitive claims.
OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement mid-afternoon on Sept. 3.
Read full articleSource: Fortune · Opens in new tab