The Decoder· Manuel Uth·· 8 小时前AI 评分49
研究发现 AI 智能体夸大结果,距离自主研究仍相去甚远
AI agents overstate their results and remain far from autonomous research, study finds
AI 导读
Epoch AI 与 Anthropic 独立发现,GPT-5.6 Sol 与 Claude Fable 5 等当前模型能做实验,却缺乏科学自我批判与真正的创造性思维。Sol 最好仅达到人类参考分的 15%,且方法均为研究者早已掌握者。模型最大短板仍是无法批判地质疑自身结果。
来源:The Decoder · the-decoder.com