A study conducted by Princeton University and the UK AI Security Institute questions Anthropic's and OpenAI's claims about the ability of artificial intelligence agents to conduct scientific research autonomously. In the evaluation, published in August 2026, two frontier models, Claude Opus 4.8 and GPT-5.6 Sol, were given research tasks based on unpublished papers submitted to the NeurIPS conference — and both generated works were rejected by the original authors, who served as reviewers.

The methodology, called Shadow Evaluation, subjected agents to central research questions from two not-yet-published papers: one on how personality traits of language models can be steered by weights, and another on a method called TabPFN, which detects when a tabular prediction model faces data very different from its training data.

Evaluation results

The agents had six days, US$ 3,000 in API credits, a GPU budget, and access to a virtual machine and the open web. The researchers used the OpenClaw scaffold, an open-source framework. Despite completing engineering tasks and compiling full LaTeX papers, both works were rejected — one received “Strong Reject.” The reviewers pointed to weakly motivated data, poorly justified experiments, unreadable prose, and a lack of new contributions.

Analysis of the logs revealed systematic failures: a lack of judgment about what is publishable, difficulty solving problems creatively, inability to reverse strategies when initial hypotheses failed, and poor resource management. Both agents spent less than half of the API budget and abandoned the most ambitious goals within the first ten hours.

The agents also violated page limits — the papers would have been rejected by desk review at NeurIPS — and exhibited “behavioral state decay,” a phenomenon similar to the one described by Meta AI, in which the model recognizes a requirement but later disregards it while fixing a bug.

Replication with another model

The team repeated an experiment with GPT-5.6 Sol and OpenAI's Codex scaffold. Almost all failure modes repeated, with the model consuming the US$ 3,000 budget in just over two days and producing experiments of insufficient size.

The results contrast with statements from the companies themselves. In June, Anthropic published the text “When AI Builds Itself,” with internal data on research acceleration and the idea of a coordinated pause in global development. OpenAI said GPT-5.6 Sol helped in the post-training of a smaller model and saved researchers weeks of work — a contribution that, according to the study, does not even appear in the model's 81-page system card.

Study limitations

The authors acknowledge that the study covers only two papers and that the reviewers were not blind, since they knew the research question and knew they were evaluating AI-generated works. Even so, they argue that the weakness of the results is so clear that these limitations would hardly change the overall picture. The professional reviewers, logs, and agent repositories were made available online.

The study also questions the use of workshop acceptance as a success metric. According to the researchers, Sakana AI submitted three papers from The AI Scientist-v2 to an ICLR workshop in 2025; one was accepted with an average score of 6.33 and later withdrawn after citation errors were discovered. Workshop acceptance rates reach 60–70%, compared with 20–30% at main conferences.

More from Radar