AIThe Decoder1h ago
AI agents overstate their results and remain far from autonomous
AI agents overstate their results and remain far from autonomous research, study finds

TL;DRAI agents can run experiments but exaggerate results and can't think independently like humans.
Why it matters: Hype around AI autonomy is premature; current systems need human oversight for actual research progress.
Epoch AI and Anthropic independently found the same thing: current AI models like GPT-5.6 Sol and Claude Fable 5 can run experiments but lack scientific self-criticism and genuine creative thinking. At best, Sol reached 15 percent of the human reference score, and even that came…
Read full articleSource: The Decoder · Opens in new tab