Frontier AI agents can write LaTeX, debug GPU clusters and run hundreds of experiments — yet they fall short when asked to perform open-ended scientific discovery, according to a new Princeton-led shadow evaluation. The team provided top-tier agents with six days, about $3,000 in API compute per run and real, unpublished research questions. The resulting manuscripts did the engineering work but were unambiguously rejected by the scientists who posed the problems.
What the shadow evaluation tested
The study was designed to probe a capability that many benchmark suites do not: producing genuinely new, expert-credible research when the answers are not known in advance. Standard benchmarks such as SWE-Bench, MLE-Bench and EXP-Bench typically evaluate agents on tasks with known solutions — for example, fixing a software bug, reaching a target metric in a competition, or reproducing an experiment where success can be measured automatically. That structure can allow models to optimize toward the test metric rather than demonstrate open-ended research reasoning.
Princeton's team introduced a different protocol called shadow evaluations, in which frontier agents were presented with original research questions that had no ground-truth answers. The objective was to see whether an agent could:
- identify which experiments mattered,
- detect when a line of inquiry was failing, and
- produce work that subject-matter experts would accept as legitimate research.
"automated AI research intern"
The timing of the paper is notable. OpenAI CEO Sam Altman publicly set a goal for an "automated AI research intern" by September 2026, a target that the Princeton study directly interrogates. With that deadline less than a month away at the time of the study's reporting, the shadow evaluation provides the first empirical test designed to measure whether current agents meet the practical demands of independent research assistance.
Findings and what they mean
Across runs where agents were given compute and time resources representative of a real-world research assistant, the teams observed the following pattern:
- Agents performed substantial technical and engineering tasks: drafting LaTeX, orchestrating large-scale experiments, and managing infrastructure.
- Despite completing experiment pipelines, the final papers did not convince the domain experts who supplied the questions and were rejected on scientific grounds.
The paper also highlights a limitation in current benchmark-driven assessments. EXP-Bench, which assembled 461 research tasks from 51 top-tier AI papers and tests complete, executable experiments, reported a success rate of only 0.5% for frontier models when answers were known and measurable. In contrast, the Princeton shadow evaluation exposed agents to unknowns; in that setting, the agents produced work that failed expert scrutiny.
| Evaluation feature | Parameter |
|---|---|
| Compute budget per run | $3,000 |
| Time allotted | Six days |
| EXP-Bench success (known answers) | 0.5% |
Implications for AI research and policy
The findings carry several practical consequences. First, progress on leaderboard-style benchmarks does not necessarily imply agents can replace human judgment in open research. Benchmarks that have measurable, automatable targets can be overfit during training and evaluation; they may reward optimization toward a metric rather than robust scientific reasoning.
Second, as organizations and funders discuss deploying AI as research assistants or automating parts of the scientific workflow, shadow evaluations offer a more stringent yardstick: do agents generate work that domain experts would accept? For now, the Princeton study suggests the answer is no.
Finally, the study provides a reality check for timelines that envision near-term, fully autonomous AI researchers. Even highly capable agents that can handle engineering and experimental throughput still struggle with the judgment calls, hypothesis framing and interpretive nuance required for credible, novel science.
In short, today’s top agents can do much of the heavy lifting of experimental work, but they do not yet meet the bar for independent scientific discovery when evaluated against the standards of practicing researchers.