Evum AI logo
AI Is Writing 90% of the Papers Now

AI Is Writing 90% of the Papers Now

← Back to blog

Nine in ten biomedical papers now show signs of AI help, and detection tools just got scary-good — a preview of what's coming for every knowledge-work team.

Here's a number that should reframe how you think about AI adoption in your own organization: an estimated 90% of recent biomedical papers now show signs of AI assistance. That figure, reported by Springer Nature, isn't about a niche group of early adopters or a splashy lab demo. It's the ordinary, unremarkable state of a rigorous, high-stakes professional field — one where careers, funding, and reputations ride on the words being published.

If AI has quietly become the default co-author in one of the most credibility-obsessed corners of knowledge work, it has almost certainly done the same inside your company. You just haven't measured it.

Two trends collided this month

Two developments landed at once, and they're more interesting together than apart.

The first is that adoption number. Signs of AI involvement — telltale phrasing, statistical fingerprints in word choice, structural patterns — now appear in the overwhelming majority of biomedical manuscripts. Note the careful language: "signs of AI help." This isn't a claim that 90% of papers were written by a chatbot start to finish. It spans everything from a researcher polishing a clumsy sentence to wholesale drafting. But the range itself is the point. AI has become woven into the writing process so thoroughly that the clean line between "human-written" and "machine-written" has effectively dissolved.

The second development is that the tools for spotting AI text have made a genuine leap. A detector called Pangram was reported by research firm Epoch AI to have produced zero false positives across 495 human-written texts. Pangram's own technical paper claims a 0.0041% false-positive rate on English, with comparable performance across more than 100 languages — and, importantly, without the bias against non-native English writers that plagued earlier detectors.

That last part matters. First-generation AI detectors were quietly a disaster: they flagged plenty of honest human writing as machine-made, and they disproportionately punished people writing in a second language. Institutions that leaned on them made real, damaging mistakes. The new generation is a different class of tool.

AI Is Writing 90% of the Papers Now — infographic

Why "detection got better" is a double-edged result

At first glance, better detection sounds like the answer to the adoption surge. Ninety percent of papers touched by AI? Fine — we'll just detect and label them.

It's not that simple, for two reasons.

First, detection accuracy and detection usefulness are different things. A tool can be 99.99% accurate at spotting machine-generated prose and still tell you almost nothing you can act on. If most of the flagged text is a human researcher using AI to fix grammar — an entirely legitimate use — then a positive result isn't misconduct. It's Tuesday. The tool answers "was a model involved?" when the question you actually care about is "was this work done honestly and competently?" Those aren't the same question, and no detector answers the second one.

Second, we're now in an arms race with a moving finish line. Every improvement in detection creates pressure to make AI text harder to detect — through paraphrasing, "humanizing" tools, or simply better models. Nature's own coverage frames it plainly as an arms race, and arms races don't end with one side winning permanently. They end with everyone spending more and trusting less.

The real lesson isn't about scientists

It's tempting to file this under "academic publishing has a problem." But the more useful reading is that biomedical research is a leading indicator for everyone else.

Research writing is about the hardest possible environment for undisclosed AI use to take hold: it's peer-reviewed, adversarially scrutinized, and career-defining. If AI assistance saturated that field to 90%, then your marketing copy, your analyst reports, your legal memos, your customer emails, and your internal strategy docs are already running at least as high. Probably higher. Nobody's peer-reviewing the quarterly deck.

Which surfaces the uncomfortable truth: most organizations have no idea how much of their own output is now AI-assisted, where, or by whom. They're in exactly the position of a journal editor before the detection tools existed — sensing that something has shifted, unable to quantify it, and reaching for policies that assume a bright line that no longer exists.

What to actually do about it

The instinct to buy a detector and start policing is the wrong first move. Detection is a downstream control that fails when you don't understand the upstream reality. A better sequence:

  • Inventory before you police. You can't govern what you can't see. Know which teams use which AI tools, for what kinds of work, and where the output goes. The 90% figure is only scary because it's usually invisible until someone measures it.
  • Write disclosure norms, not bans. Blanket prohibitions push AI use underground, which is the worst outcome — you get all the risk and none of the visibility. Clear, low-friction disclosure ("say when and how you used AI") keeps behavior in the open. Nature's research even found that transparently disclosed, human-edited work is often trusted more than work with murky origins.
  • Judge the output, not the tool. A detector tells you a model was involved. Your quality process should tell you whether the result is accurate, original, and defensible. Keep the second question firmly in human hands.
  • Assume the arms race, plan for it. Any control that depends on detection being permanently better than generation is a control with an expiry date. Build governance that survives the day detection stops working.

The 90% number isn't a scandal. It's a status report — a snapshot of how fast, how quietly, and how completely AI has embedded itself in serious work. The organizations that come out ahead won't be the ones with the best detector. They'll be the ones who measured their own reality first, and governed the work instead of chasing the tool.