AI agents overstate their results and remain far from autonomous research, study finds
Epoch AI and Anthropic independently found the same thing: current AI models like GPT-5.6 Sol and Claude Fable 5 can run experiments but lack scientific self-criticism and genuine creative thinking. At best, Sol reached 15 percent of the human reference…
Adjust this page to suit you. Your choices are saved in this browser.
Cookies. Chipokia uses one cookie, and only if you allow it: a random ID so your up and down votes stay yours. No analytics, no advertising, no tracking of any kind. Cookie & Privacy Policy.