The most-read paper in AI drug discovery this summer is not a victory lap. It is an audit, and the audit does not balance.
On 7 August, Nature Reviews Drug Discovery published "Artificial intelligence in drug discovery — what it is, where we stand and the path forward," a Perspective running to 221 references and more than a dozen authors. It has drawn roughly 26,000 accesses in under a month, which for a review journal is a small stampede. The reason is the sentence sitting in its abstract, delivered without a hedge: "Although a wide variety of AI methods have been developed, applied and benchmarked, evidence of their clinically relevant impact is, so far, disappointingly limited."
This is not a gathering of temperamental skeptics. The byline includes lead author Andreas Bender, Jack W. Scannell — whose "Eroom's law" analysis defined pharma's productivity decline for a generation of executives — Francesca Grisoni, and Isidro Cortés-Ciriano of the European Bioinformatics Institute. These are people who build the models. Their complaint is not that AI does nothing, but that the field has spent a decade measuring the wrong thing. Benchmarking, they argue, "need[s] to move on from model validation and instead focus on their ability to improve decision making."
The credit column: structure and sequence
Where AI has unambiguously won is upstream, in the physics-adjacent problems where data is dense and ground truth is cheap. Protein structure prediction is the canonical case. What was once a multi-year crystallography project is now a query you run before lunch, and the downstream field — docking, binder design, target triage — has been rebuilt around that assumption.
The second win is maturing fast. A survey published 3 July in npj Drug Discovery by Yiheng Zhu and colleagues, "Generative AI for controllable protein sequence design," catalogues a field that went from curiosity to infrastructure in about four years. "Fueled by advances in generative AI, the field of protein design is undergoing a paradigm shift," the authors write. The scale numbers make the point: ProtGPT2 at 738 million parameters, ProGen2 at 6.4 billion, xTrimoPGLM at 100 billion. ESM-IF was trained using 12 million AlphaFold2-predicted structures and gained close to ten percentage points in sequence recovery on held-out native backbones.
Read the fine print, though, and the metrics are almost entirely computational — perplexity, sequence recovery on CATH benchmarks, self-consistency scores. They measure whether a model reproduces known biology, not whether a designed protein helps a patient.
The debit column: what happened in the clinic
The best available industry-wide numbers still come from a 2024 analysis in Drug Discovery Today by Madura Jayatunga and colleagues at Boston Consulting Group — a paper the Nature Reviews authors themselves cite. AI-discovered molecules cleared Phase 1 at an 80 to 90 percent rate, against a historic industry range of roughly 40 to 65 percent. That is a genuine result, and a modest one: Phase 1 mostly tests whether a molecule behaves like a drug, which is precisely what generative chemistry is good at. In Phase 2, where compounds must show they actually treat disease, AI-originated drugs succeeded at about 40 percent — indistinguishable from everything else, on a small sample. End to end, the analysis projected AI might lift probability of success from 5–10 percent to 9–18 percent. Useful. Not a revolution.
The field's single most concrete clinical readout arrived in June 2025, when Nature Medicine published Phase 2a results for rentosertib, Insilico Medicine's TNIK inhibitor for idiopathic pulmonary fibrosis — a molecule whose target and structure were both generated by AI. Across 71 patients at more than 20 sites in China over 12 weeks, the 60 mg once-daily arm showed a mean forced vital capacity gain of +98.4 mL versus a −20.3 mL decline on placebo. In patients not taking standard antifibrotics, the improvement reached 187.8 mL.
Insilico founder and CEO Alex Zhavoronkov called the data proof that AI has "transformative potential in drug discovery and development." The trial's lead investigator was notably more restrained. "The sample size in each patient group was relatively limited, and these findings will need to be validated in larger cohort studies," said Zuojun Xu, professor at Peking Union Medical College. The safety picture supports his caution: seven patients discontinued for liver injury or dysfunction, and only 12 of 18 patients completed the high-dose regimen against 88 percent on placebo. Set against that are the field's public failures — Exscientia's discontinued EXS-21546 and BenevolentAI's BEN-2293 in atopic dermatitis.
Why this matters beyond pharma
Drug discovery is the most expensive stress test AI has ever been put through, and it is failing in an instructive way. The scaling story that worked for language does not transfer, because the binding constraint is not compute or parameters. It is data. Negative trial results go unpublished, failed compounds sit in corporate vaults, and every model trained on approved drugs is learning from survivors. You cannot fix survivorship bias by adding GPUs; you fix it by running the slow, costly wet-lab work that AI was supposed to replace.
A second lesson generalizes to every enterprise AI deployment. The Nature Reviews authors identify "technology push" rather than "science pull" as a root cause — models built because they could be built, benchmarked on tasks convenient to score, then handed to scientists whose actual decisions they were never designed to inform. Anyone who has watched a beautifully benchmarked model fail to change a single workflow will recognize the pattern.
What to watch
The honest ledger says what the Nature Reviews authors say: the front end is solved enough to be boring, and the back end is unmoved. What changes that is not another benchmark. It is a Phase 3 readout — a properly powered, geographically diverse trial in which an AI-originated molecule beats standard of care on a hard endpoint. Insilico has begun regulatory discussions on larger rentosertib cohorts. Watch that trial, watch whether the AI-native pipeline's Phase 2 rate moves off 40 percent as more compounds mature, and treat anyone conflating a generation demo with a clinical result as selling something.
“The sample size in each patient group was relatively limited, and these findings will need to be validated in larger cohort studies.”— Zuojun Xu, Professor, Peking Union Medical College