From AlphaFold to Co-Folding — What Structure Prediction Actually Does for Drug Discovery
TL;DR
Protein structure prediction crossed two thresholds. AlphaFold (2021) made static single-chain structure routine for many folded proteins — when they have evolutionary cousins. Since 2024 the frontier moved to co-folding: predicting a protein with a ligand in the pocket (AlphaFold3, Boltz, Chai), and in 2025 Boltz-2 reported binding affinity approaching free-energy-perturbation quality on selected public benchmarks at ~1000× lower cost — though a later leakage analysis showed a structure-blind baseline matching it there, so transferable affinity prediction isn’t yet proven. Affinity is one of the central quantities compounds are ranked on. But read the benchmarks: a low-RMSD pose (root-mean-square deviation — the Å gap between predicted and true atom positions) is necessary but not sufficient. Physical validity and novel-target generalization are the real bar, and one static structure still misses the dynamics that govern druggability. Here’s the honest map, from someone building cardiac drug-discovery agents on exactly these tools.
The story so far
AlphaFold2 did something real and specific: at CASP14 it hit a median backbone accuracy of 0.96 Å vs 2.8 Å for the next-best method [1]. RoseTTAFold reached “approaching” accuracy the same year, and did it open and fast [2]. A useful structural hypothesis went from months of experimental effort to minutes of computation — though experimental structure determination stays essential for ligand density, ordered waters, alternate states, and much mechanistic detail.
But “solved” has fine print worth memorizing. AlphaFold made accurate static, single-chain models routine when an evolutionary signal exists — accuracy falls off sharply below ~30 sequences in the multiple-sequence alignment (MSA) [1]. It did not solve ligands, dynamics, mutation effects, or MSA-poor targets. And it hands you a calibrated confidence — pLDDT (a per-residue self-confidence score) tracks true accuracy (r ≈ 0.76), and PAE (predicted aligned error) tells you which inter-domain and interface geometry to trust [1]. That confidence map is itself a deliverable: it tells you where the model is guessing.
Where we stand
Drug discovery doesn’t need a lone protein — it needs the bound system, a target with a molecule in the pocket. That’s the action since 2024.
- AlphaFold3 replaced the Evoformer with a diffusion model that jointly predicts proteins + nucleic acids + small molecules + ions [3]. On PoseBusters (428 protein-ligand structures) it reaches ~76% ligand RMSD < 2 Å from just sequence + a SMILES string — and beats classical docking (AutoDock Vina) even though docking is handed the crystal pocket (P = 2.3×10⁻¹³) [3].
- Then the ecosystem opened up — around AF3, not AF3 itself. Within ~7 months, permissively-licensed reproductions — Boltz-1 [4] and Chai-1 [5] — reached AF3-class accuracy (Chai-1 benchmarks 77% on PoseBusters, 81% when given the apo — unbound — structure; Boltz-1 claims AF3-level and shows it head-to-head against Chai-1, since AF3’s weights weren’t yet public) — an academic group can now run this locally. By 2026 several locally-deployable models — Boltz, Protenix [6], and RoseTTAFold3 [7] — had become competitive with AF3 on selected benchmarks, though no single open model matched it consistently across every modality (RF3, for instance, reports performance between AF3 and other open models on de-leaked antibody–antigen sets [7]). AF3 itself launched in May 2024 as a restricted server with no code or weights; DeepMind released inference code and weights in November 2024 but only for non-commercial use (Nature later corrected an article that had called AF3 “open source”) — so permissive models like Boltz still matter for commercial work and unrestricted local deployment.
- RoseTTAFold All-Atom [8] extends the open lineage to full assemblies and feeds RFdiffusion for de-novo binder design — closing the predict→design loop.
- For antibodies and orphan targets where MSAs fail, single-sequence protein-language-model approaches earn their keep: ESMFold is far faster (up to ~60× — it skips the MSA search entirely) [9], and OmegaFold beats AF2 on orphan proteins (TM 0.73 vs 0.60) and CDR-H3 loops (2.12 Å vs 2.98 Å) [10].
The 2025 inflection: affinity, not just pose. Knowing where a ligand sits is not knowing how tightly it binds — and affinity is one of the central quantities compounds are ranked on. Boltz-2 reported the first AI binding-affinity correlations approaching FEP quality: on the 4-target FEP+ subset (CDK2, TYK2, JNK1, P38) it hits average Pearson r = 0.66 — approaching FEP+ itself (r = 0.78) at >1000× lower cost (~20 s/ligand) [11]. Run out-of-the-box on the CASP16 affinity benchmark, it beat every official submission. For a screening pipeline, high-throughput ranking at library scale is a genuinely new capability. But how much of that number is real generalization? The authors are already honest about the ceiling: across 8 blinded industrial hit-to-lead assays, average Pearson fell to 0.39 (per-target range 0.165–0.634) [11] — “strong performance on public benchmarks does not always translate to real-world drug discovery.” And a June 2026 leakage analysis lands harder: a structure-blind, ligand-only baseline — no protein information at all — matches that r = 0.66 on the same FEP+ set (and collapses to r = 0.14 on genuinely novel ligands) [12], arguing that public affinity benchmarks reward memorizing correlated ligand activity over learning transferable protein–ligand energetics. Boltz-2 is a promising high-throughput ranking model; FEP-level generalization is not yet established. The benchmarks share the blame: leakage-corrected re-splits (LP-PDBBind) deflate the very scores everyone quotes [13].
Going deeper — RMSD is necessary, but not sufficient
First, mind what’s being compared. PoseBusters evaluated docking methods (DiffDock, Gold, Vina) that are handed a receptor structure; co-folding models (AF3, Boltz, Chai) do the harder job of predicting the complex from sequence — so raw success rates across the two aren’t a clean head-to-head. Within that caveat the lesson holds: ranking by RMSD alone is misleading. DiffDock looks best by RMSD on one set (72%) but drops to 47% once you check physical validity (stereochemistry, bond lengths, clashes); on the held-out benchmark of unseen complexes only ~12% of its poses are physically valid — near-zero on truly novel (<30%-identity) sequences — while classical Gold and Vina (55–58%) hold up [14]. The failure runs deeper than validity: a peer-reviewed 2026 study (Nature Structural & Molecular Biology), scoring 2,600 post-cutoff complexes, finds the leading co-folding models (AF3, Boltz, Chai) largely reuse poses similar to their training data, with accuracy falling as training-set similarity drops [15]. Deep-learning poses routinely have steric clashes, bad chirality, non-planar rings; MM force fields “contain docking-relevant physics missing from deep-learning methods” [14]. Boltz-1 openly documents its own hallucinations — entire chains stacked on each other [4]. And even a good-RMSD pose often misses the actual interaction fingerprint [16] — though a fingerprint from one crystal structure isn’t absolute truth either (alternate waters, protonation, tautomers, and valid alternate binding modes all shift it), so treat interaction recovery as a stronger additional check, not a replacement for energetic and prospective validation. Rule I live by: never accept a co-folding pose on RMSD alone — demand physical validity, novel-target generalization, and interaction recovery.
What still needs the wet lab (and the physics). One static structure is the standing limitation, and it’s the one that bites drug discovery hardest — real targets are ensembles: cryptic pockets, allostery, apo↔holo (unbound↔bound), disorder. AlphaFold gives one conformer; the diffusion models are “prone to hallucination” in unstructured regions [3]. Early fixes (Boltz-2’s MD-conditioning; AlphaFold-Metainference for disordered ensembles [17]) exist but aren’t solved. Still hard: absolute project-grade affinity, mutation/ΔΔG effects, glycan chemistry (42–46% success [3]), antibody–antigen interfaces (4–8% high-quality even with restraints [5]), and anything genuinely novel and MSA-poor. The encouraging counter-signal: prospective large-scale docking against unrefined AF2 models has found potent ligands at hit rates matching experimental structures [18] — though that was shown on two receptors (σ2 and 5-HT2A), so it’s a feasibility result, not proven equivalence across protein families. The AI structure is already useful for discovery, even where it’s imperfect.
What’s next
Overhyped: “AF3 replaces docking.” Half-true, and the honest half matters — AF3 beats docking on pose, but that benchmark ignores physical validity, says nothing about affinity, and co-folding overstates success when interaction fingerprints are scored. It changes docking; it doesn’t retire the affinity problem. “Boltz-2 = FEP for free” — it reports FEP-approaching numbers on curated public sets, but a structure-blind baseline matches them and it’s inconsistent on live medicinal-chemistry projects: a high-throughput ranker, not an FEP replacement. “AlphaFold solved protein structure” — it made static single-chain modeling routine given an MSA, not every monomer, membrane protein, or alternate state.
Not hyped enough: the confidence maps (pLDDT/PAE) that tell you where to trust the model — a quietly huge deliverable; open weights that put AF3-class modeling in any lab; and MSA-free speed that makes proteome- and library-scale structural triage tractable.
What to watch: the next unlock isn’t a prettier monomer — it’s ensembles and affinity you can trust on novel targets: generative conformational sampling for cryptic pockets and allostery — turning single-structure predictors into ensemble generators (AlphaFlow [19]) or learning the equilibrium distribution directly (DiG [20]), affinity models that survive blinded industrial assays, and co-folding that’s scored on interaction recovery, not RMSD. When a model reliably tells you how tight, on a target it has never seen, and which pose is physically real, the assay becomes a focused validation-and-learning loop rather than the first filter for every hypothesis — binding is still only one layer of therapeutic activity (selectivity, PK, toxicity, disease biology all stay experimental).
Sources
- Jumper et al., “Highly accurate protein structure prediction with AlphaFold” (AlphaFold2, Nature, 2021) — https://doi.org/10.1038/s41586-021-03819-2
- Baek et al., “Accurate prediction of protein structures and interactions using a three-track neural network” (RoseTTAFold, Science, 2021) — https://doi.org/10.1126/science.abj8754
- Abramson et al., “Accurate structure prediction of biomolecular interactions with AlphaFold 3” (Nature, 2024) — https://doi.org/10.1038/s41586-024-07487-w
- Wohlwend et al., “Boltz-1: Democratizing Biomolecular Interaction Modeling” (bioRxiv, 2024) — https://doi.org/10.1101/2024.11.19.624167
- Boitreaud et al., “Chai-1: Decoding the molecular interactions of life” (bioRxiv, 2024) — https://doi.org/10.1101/2024.10.10.615955
- Chen et al., “Protenix — Advancing Structure Prediction Through a Comprehensive AlphaFold3 Reproduction” (bioRxiv, 2025) — https://doi.org/10.1101/2025.01.08.631967
- Corley et al., “Accelerating Biomolecular Modeling with AtomWorks and RF3” (RoseTTAFold3, bioRxiv, 2025) — https://doi.org/10.1101/2025.08.14.670328
- Krishna et al., “Generalized biomolecular modeling and design with RoseTTAFold All-Atom” (Science, 2024) — https://doi.org/10.1126/science.adl2528
- Lin et al., “Evolutionary-scale prediction of atomic-level protein structure with a language model” (ESMFold, Science, 2023) — https://doi.org/10.1126/science.ade2574
- Wu et al., “High-resolution de novo structure prediction from primary sequence” (OmegaFold, bioRxiv, 2022) — https://doi.org/10.1101/2022.07.21.500999
- Passaro et al., “Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction” (bioRxiv, 2025) — https://doi.org/10.1101/2025.06.14.659707
- Mattsson & Walters, “Identifying and Addressing Systematic Data Leakage in Protein-Ligand Affinity Benchmarks” (bioRxiv, 2026) — https://doi.org/10.64898/2026.06.29.735309
- Li, Guan, Zhang & Head-Gordon, “Leak Proof PDBBind: A Reorganized Dataset of Protein-Ligand Complexes for More Generalizable Binding Affinity Prediction” (LP-PDBBind, arXiv, 2024) — https://arxiv.org/abs/2308.09639
- Buttenschoen, Morris & Deane, “PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences” (Chemical Science, 2023) — https://doi.org/10.1039/D3SC04185A
- Škrinjar et al., “Evaluating generalization in protein–ligand cofolding methods” (Nature Structural & Molecular Biology, 2026) — https://doi.org/10.1038/s41594-026-01797-5
- Errington et al., “Assessing interaction recovery of predicted protein–ligand poses” (2025) — https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12090448/
- Brotzakis et al., “AlphaFold prediction of structural ensembles of disordered proteins” (AlphaFold-Metainference, Nat. Commun., 2025) — https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11829000/
- Lyu et al., “AlphaFold2 structures template ligand discovery” (bioRxiv/Science, 2024) — https://doi.org/10.1101/2023.12.20.572662
- Jing, Berger & Jaakkola, “AlphaFold Meets Flow Matching for Generating Protein Ensembles” (AlphaFlow, arXiv, 2024) — https://arxiv.org/abs/2402.04845
- Zheng et al., “Towards Predicting Equilibrium Distributions for Molecular Systems with Deep Learning” (Distributional Graphormer / DiG, Nat. Mach. Intell., 2024) — https://doi.org/10.1038/s42256-024-00837-3
Sung, J. (2026). "From AlphaFold to Co-Folding — What Structure Prediction Actually Does for Drug Discovery." LatentCell. https://latentcell.ai/posts/protein-structure-predictionCC BY 4.0 — reuse with credit.Full formats (APA · MLA · BibTeX) →