New studies find AI agents handle the grunt work of research but lack the judgment and originality to run projects end-to-end — a reality check on self-improving AI.
For the past year, the loudest story in AI has been about acceleration: models that write their own code, agents that run their own experiments, and — the ultimate prize — systems that improve themselves, kicking off a feedback loop that leaves human researchers in the dust. It's a compelling narrative. It's also, according to a fresh batch of empirical work, running well ahead of the evidence.
Three separate research efforts released this month landed on the same uncomfortable conclusion. AI systems are genuinely good at the mechanical parts of research. They are still bad at the parts that actually make research research.
What the studies actually found
A multi-institution group led by Peter Kirgis and Sayash Kapoor at Princeton put AI agents to the test on real machine-learning research tasks. The agents could reliably solve the engineering problems — writing code, running experiments, not getting stuck on setup. But when it came to producing original work at the caliber of a paper accepted at a top ML conference, they fell short. They lacked the judgment and creativity to generate genuinely novel contributions.
In a related effort, Kapoor's team (he co-wrote the book AI Snake Oil) asked an AI tool to conduct research and write up its findings based on questions drawn from two published papers — then had the original human authors grade the output. The results were mixed in an instructive way. The AI ran experiments without getting stuck, a real capability. But the human experts judged that it wandered down fruitless paths and failed to exercise the discipline a competent researcher brings to knowing which threads are worth pulling.
Then there's ASI-Bench, a new benchmark built by more than 40 experts over 31,000-plus hours of human labor. Its headline finding: current systems "remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research." Not "getting close." Far.

Why this matters more than it sounds
It would be easy to file this under "AI is overhyped, film at eleven." But the specific shape of the gap is what's useful.
The tasks AI is good at — writing code, running experiments, searching literature, executing a defined procedure — are exactly the tasks that are checkable. There's a right answer, or at least a runnable result. The tasks it's bad at — deciding which questions are worth asking, recognizing when a promising direction is actually a dead end, knowing what counts as a novel contribution — are the ones that require taste, context, and judgment. Those don't come with an answer key.
This distinction isn't academic. It maps almost perfectly onto how AI is showing up inside real companies right now.
The parts of knowledge work AI is genuinely transforming are the bounded, verifiable ones: drafting, summarizing, transforming data, generating first-pass code. The parts it struggles with are the open-ended, judgment-heavy ones: deciding strategy, prioritizing between imperfect options, knowing when the output is subtly wrong. If you've deployed AI agents in your organization and felt both impressed and quietly nervous, this is why. The capability and the gap are two sides of the same coin.
The self-improvement question, deflated
The most consequential implication is for the recursive self-improvement thesis — the idea that AI will soon get good enough at AI research to accelerate its own development in a runaway loop. That story depends on agents being able to do original, project-level research autonomously. The evidence this month says they can't, not yet, and that the missing ingredient isn't more compute or a bigger context window. It's judgment.
Notably, one of the research agents in the open-ended study was tested and correctly explained why a task it was given couldn't be done — a genuine sign of grasping constraints. But recognizing a limit and generating a breakthrough are very different skills, and the gap between them is where the hype currently lives.
What to do with this
For leaders making AI decisions, the takeaway isn't "AI can't do research, so relax." It's more precise and more useful than that.
- Deploy AI where the work is checkable. The clearest wins are tasks with verifiable outputs and a human who can catch errors. That's not a limitation to apologize for — it's where the real productivity is right now.
- Keep humans on the judgment. Anywhere the value comes from taste, prioritization, or originality, the agent is an assistant, not an owner. Design workflows that assume this, rather than hoping the model grows into the role.
- Measure maturity honestly. The organizations that get the most out of AI aren't the ones with the boldest self-improvement roadmap. They're the ones who know exactly which of their AI systems are doing real work, what those systems can and can't be trusted with, and where the gaps are. That clear-eyed inventory beats a hype cycle every time.
The machines are getting remarkably capable at the middle of the research process. The beginning and the end — the parts that require knowing what matters — still belong to us.
The lesson for anyone adopting AI: match the tool to the task's verifiability, not to the marketing. The gap between "can execute" and "can decide" is the single most important thing to understand about where AI is today.
