AIFortune3h ago

OpenAI quietly updates its evaluation metrics for GPT-6 Astra, making

OpenAI quietly updates its evaluation metrics for GPT-6 Astra, making changes that appear to favor Astra and continuing to revise other metrics after launch

OpenAI quietly updates its evaluation metrics for GPT-6 Astra, making

TL;DROpenAI adjusted GPT-6 benchmarks after launch, raising questions about metric reliability.

Why it matters: Benchmark changes post-launch undermine confidence in model comparisons and competitive claims.

OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement mid-afternoon on Sept. 3.

Read full article

Source: Fortune · Opens in new tab