AI Is Now Writing the Papers and Reviewing Them Too. Science Has a Problem.

Submissions to academic journals have surged 42% since ChatGPT launched. Now 21% of peer reviews are fully AI-generated. The system built to validate science is breaking.

August 24, 2026Updated August 24, 20267 min read
AI Is Now Writing the Papers and Reviewing Them Too. Science Has a Problem.

Academic peer review is one of the oldest quality-control systems in human knowledge production. It's slow, imperfect, and frequently gamed. But it has, broadly, held. What's happening to it right now is different.

AI-generated papers are flooding journals at a scale the system was never designed to absorb. The numbers aren't speculative anymore. One major management journal saw submissions rise 42% after ChatGPT launched in late 2022, roughly double the increase it recorded during the entire COVID-19 pandemic. The editors attribute nearly all of that jump to AI involvement. Submissions scored as human-only actually fell. By early 2026, the majority of manuscripts arriving at that journal show some degree of AI involvement, and the fastest-growing category is papers with 70% or higher AI-generated content.

That's the writing side. The reviewing side is worse.

The Peer Review System Is Now Reviewing Itself With AI

An analysis of submissions to ICLR 2026, one of the most competitive machine learning conferences in the world, found that approximately 21% of reviews were entirely AI-generated. More than half showed some degree of AI involvement. The program chairs acknowledged the problem publicly. A separate survey of 1,600 researchers found that more than 50% have used AI tools during peer review, often in direct conflict with journal guidance.

This isn't a niche concern limited to computer science conferences. A detection model applied to Nature Communications reviews classified roughly 12% of 2025 reviews as AI-generated, with the sharpest growth concentrated in late 2024. An audit of published biomedical literature found detectable traces of LLM processing in at least 13.5% of PubMed abstracts from 2024.

The system built to validate what counts as scientific knowledge is now, in meaningful part, automated on both ends.

What AI-Generated Papers Actually Look Like

There's a common assumption that AI papers are obviously bad. That's not quite right, and it's the more dangerous misconception.

A formally passable machine learning paper can be produced by an AI system in roughly 15 hours at a cost of around $140. A graduate student might spend an entire semester producing their first accepted workshop paper. The speed and cost differential isn't a marginal gap. It's a structural problem. AI can generate plausible-looking research infinitely faster than humans can evaluate it.

The quality gap is real, but it's narrowing. Detection models catch stylistic signatures left by earlier models like GPT-3.5 and GPT-4, including bloated nominal prose, characteristic vocabulary patterns, and inflated reading complexity scores. Newer models are better at hiding those tells. The detection tools that journals rely on today are essentially calibrated to yesterday's AI output.

One specific problem has attracted particular attention: fake citations. Hallucinated references are appearing in peer-reviewed medical papers at scale. An AI-assisted audit identified nearly 3,000 peer-reviewed medical papers containing fake citations. These aren't papers that were rejected. They passed review and got published.

This connects to a broader pattern in AI output that we've covered before. AI Hallucinations Hit the Courtroom: When Your AI Tool Gets You Sanctioned is about legal sanctions, but the underlying mechanism is identical: AI systems producing confident, plausible, wrong citations that humans fail to catch before publication.

Why the Review System Can't Keep Up

Peer review was designed around human-scale production. A journal's editorial board could realistically evaluate the volume of serious research humans could produce. That assumption is gone.

The submission surge is overwhelming reviewers. Journals that already struggled to find qualified reviewers for legitimate submissions are now drowning in a much larger pool. Editors who try to use AI detection tools to manage the volume face a real problem: false positive rates are high enough that those tools can't support editorial decisions on their own. Flag too aggressively and you reject legitimate work. Flag too loosely and you publish slop.

Reviewers facing their own overload are, predictably, turning to the same tools. The pattern of AI writing a paper, submitting it for review, and having that review conducted by AI is no longer hypothetical. It's measurable, at ICLR scale, right now.

ArXiv, the preprint server that sits upstream of formal peer review, already moved to ban authors for a year if AI writes their papers. That's a significant policy shift that acknowledges the problem directly. But preprint enforcement doesn't solve the downstream journal problem, and journals are moving much more slowly.

The Hallucinated Citation Problem Is Worse in Medicine

The fake citation issue deserves its own attention, particularly in medical research. An audit found nearly 3,000 peer-reviewed medical papers with fabricated references. In a field where clinical practice depends on published evidence, fake citations aren't an academic integrity problem. They're a patient safety problem.

AI Radiology Is Now Infrastructure in Rural Hospitals. The Safety Debate Is Just Getting Started. covers what happens when AI tools reach clinical deployment before the validation framework catches up. The citation problem hits earlier in the pipeline. Clinicians making treatment decisions based on published research can't verify that the papers they're reading cited real studies, because journal review didn't catch the hallucinations.

The FDA's push toward real-time clinical trial frameworks assumes the underlying research literature is trustworthy. If AI-generated fake citations are contaminating the evidence base, that assumption has a hole in it.

What the Research Community Is Actually Doing About It

Responses are fragmentary. ICLR's program chairs acknowledged the problem but noted detection tools alone can't support decisions. Some journals are requiring authors to declare AI use explicitly. Others are experimenting with post-publication audits. A few have raised submission fees to create friction.

None of these are solutions. They're friction. The incentive structure that drives AI paper submission hasn't changed: publish-or-perish pressure is real, AI lowers the marginal cost of a submission to near zero, and detection is unreliable. Until the incentive structure changes, volume will keep climbing.

The optimistic framing is that detection models and AI-assisted review tools will eventually match the output tools in sophistication. That may be true. But the gap between "eventually" and "right now" is where the damage accumulates. Published papers with fake citations don't unpublish themselves. The scientific literature is a cumulative record, and it's being contaminated in real time.

One useful comparison: think about what happened to web content when AI writing became cheap. The web adapted, slowly and imperfectly, through better filtering, reader skepticism, and platform-level interventions. Academic publishing is less equipped for that kind of adaptation. It moves slower, its quality signals are harder to reconstruct once compromised, and the consequences of bad information are more serious than a low-quality blog post.

What Researchers and Institutions Should Do Right Now

The practical reality is that researchers, reviewers, and editors are operating in this environment whether they're ready for it or not. A few things actually help:

If you're submitting research: Document your AI use explicitly, even if the journal doesn't require it yet. Check every citation your AI tools produce against the original source. Don't assume the reference list is real. Citation hallucination is the most dangerous failure mode because it's the hardest for reviewers to catch.

If you're reviewing: Treat any paper with citation-dense claims as requiring spot-check verification. Look up the three or four most critical references. Fake citations cluster around supporting evidence for the paper's main argument. That's where the hallucinations appear most often.

If you're running an editorial board or conference program: The detection tool question matters less than the policy question. Your authors and reviewers need explicit, enforced guidance, not a terms-of-service paragraph nobody reads. ICLR acknowledged the problem after it was already measurable at scale. Getting ahead of it requires deciding what you'll actually enforce before the data makes the decision for you.

The deeper issue isn't that AI is writing papers. It's that the system built to validate science has no baseline for the current volume, no reliable detection at the frontier of model capability, and no consensus on what AI involvement actually disqualifies. Those three gaps, together, are what makes this a structural problem rather than a moderation problem.

Related News