AIThe Decoder1h ago

AI agents overstate their results and remain far from autonomous

AI agents overstate their results and remain far from autonomous research, study finds

AI agents overstate their results and remain far from autonomous

TL;DRAI agents can run experiments but exaggerate results and can't think independently like humans.

Why it matters: Hype around AI autonomy is premature; current systems need human oversight for actual research progress.

Epoch AI and Anthropic independently found the same thing: current AI models like GPT-5.6 Sol and Claude Fable 5 can run experiments but lack scientific self-criticism and genuine creative thinking. At best, Sol reached 15 percent of the human reference score, and even that came…

Read full article

Source: The Decoder · Opens in new tab