Most AI protein design papers test a few dozen molecules. In March 2026, NVIDIA and a consortium of drug companies and universities tested roughly one million.

The model is Proteina-Complexa, a fully atomistic binder generator released publicly on March 25, 2026, with code under Apache 2.0 and weights on NGC and Hugging Face. Its method paper, ‘Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time Compute,’ was accepted as an oral at ICLR 2026, led by Kieran Didi with Karsten Kreis as project lead. A companion wet-lab preprint, released March 16, carries 40-plus authors spanning NVIDIA, Manifold Bio, Novo Nordisk, Viva Biotech, Duke, Cambridge, LMU Munich, Oxford and Seoul National University.

Architecturally, Proteina-Complexa does two things at once that the field has historically split apart. It builds on NVIDIA’s La-Proteina partially latent flow-matching framework: backbone alpha-carbon atoms are modeled explicitly in 3D Cartesian space, while side chains and amino acid identity are compressed into a learned continuous latent space by an autoencoder. That means the model emits sequence and full atomic structure together, with no ProteinMPNN inverse-folding step bolted on afterward. Then, at inference, it searches. Beam search, Best-of-N, Feynman-Kac steering and Monte Carlo tree search prune and re-launch denoising trajectories, scored against structure-prediction confidence and interface hydrogen-bond energy. This is the ‘reasoning’ framing: spend more compute on hard targets, less on easy ones.

Training data came from over a million curated structures, including Teddymer, a synthetic dataset built by splitting 47 million AlphaFold Database monomers into structural domains and reassembling 10 million domain-domain pairs as stand-ins for real dimers, filtered to 3.5 million Foldseek clusters. It is an order of magnitude larger than the PDB.

The 127-target verdict

The centerpiece experiment ran on Manifold Bio’s multiplexed phage display platform: more than one million designed proteins screened against 127 targets in a single all-to-all experiment, measuring over 100 million protein-protein interactions. Proteina-Complexa contributed 467,176 sequences and produced on-design hits, binders that hit their intended target, for 86 of 127 targets (68%), with 74 of those yielding target-specific binders.

The head-to-head is the more interesting number. On 75 targets where at least one method scored, under matched compute budgets, Proteina-Complexa’s self-generated sequences hit at 2.45% averaged across targets, roughly 3x the next-best self-generated baseline (BoltzGen at 0.76%). In raw counts: 691 on-design hits from 25,707 tested sequences, against 514 for BoltzGen plus ProteinMPNN, 311 for BindCraft, 86 for RFdiffusion3 and 63 for RFdiffusion. Notably, Proteina-Complexa’s co-generated sequences beat ProteinMPNN redesign of its own backbones (691 versus 365), the only method in the panel where that held.

Focused campaigns produced the flashy affinities. Against PDGFR, 9,000 candidates were filtered to 192 for surface plasmon resonance; 122 of 191 expressed designs bound, a 63.5% hit rate, with affinities from 93.6 pM to 1.34 micromolar. Against ActRIIA, a muscle-wasting receptor relevant to cachexia and GLP-1-associated lean mass loss, 16 of 192 expressed designs bound (an 8.3% hit rate from raw designs), the tightest at 36 nM, and two blocked myostatin signaling in cells with IC50 values of 169 nM and 228 nM. Kinase work returned 40% enrichment for PAK1 mini-proteins and 50% for sub-31-residue CK1-delta peptides. A Nipah virus de novo design hit 56 nM. And 5 of 24 designs agglutinated type B red blood cells, which the authors call the first de novo proteins to bind a free carbohydrate, a class long considered inaccessible.

‘Manifold Bio’s platform uniquely enabled this massively multiplexed study, which establishes Proteina-Complexa as competitive with state-of-the-art methods,’ said Pierce Ogden, co-founder and CTO of Manifold Bio. ‘This study provides a practical demonstration of scaling laws in de novo protein design, with more designs producing more hits, and makes the case for inference scaling when experimental throughput can keep pace.’

Anthony Costa, NVIDIA’s director of digital biology, framed the architecture in explicitly LLM-inflected terms: ‘With test-time scaling, we’ve enabled the model to refine its logic before outputting a single sequence.’

What a million measurements actually buys

The economics here are real but narrower than the press release implies. Proteina-Complexa generates a candidate in 15.6 seconds versus 70.8 for RFdiffusion, and NVIDIA claims 30-60x end-to-end speedups. At that rate, the binding of hit discovery shifts from computation to assay capacity, which is precisely Ogden’s point. Programs that historically spent 12 to 18 months and low seven figures on hit generation via immunization or display libraries can, in principle, compress to weeks of GPU time plus one screening run.

But read the hit rates honestly. 2.45% is the number that reflects unfiltered scale; 63.5% is what you get after 9,000 designs are triaged down to 192 by a two-stage physicochemical and confidence filter. Those are not the same claim, and the gap is the whole story. Phage display is also a low-stringency readout: it establishes binding, not affinity, developability, expression at scale, immunogenicity or in vivo behavior. An August 2025 meta-analysis of 3,766 experimentally characterized binders found in silico filter precision varying from 0.1 to 1.0 depending on the target, which means filter performance is itself target-dependent and not transferable. And 68% target coverage inverts to 41 of 127 targets where a million designs produced nothing.

There is also a subtler finding buried in the tables: 126 of 127 targets picked up at least one specific off-design hit, a binder designed for something else. Latent promiscuity at this scale is a specificity problem waiting to be characterized.

What to watch: whether the ActRIIA or PDGFR series survives contact with animal pharmacology, since none of these molecules have cleared a single in vivo study; whether independent groups reproduce the 2.45% figure on targets NVIDIA did not choose; and whether the Teddymer trick, training on synthetic domain-domain interfaces, holds up as a general substitute for scarce experimental complexes. The wet-lab preprint is not yet peer reviewed.

“This study provides a practical demonstration of scaling laws in de novo protein design, with more designs producing more hits, and makes the case for inference scaling when experimental throughput can keep pace.”
— Pierce Ogden, Co-founder and CTO, Manifold Bio
~1 million
Designed binders experimentally tested across 133 targets
86 of 127
Targets with on-design hits in the multiplexed phage screen
2.45%
Average unfiltered hit rate, vs 0.76% for the next-best baseline
93.6 pM
Tightest measured affinity, a PDGFR binder