Journal Article
Computer Vision

Investigating whether deep learning models for co-folding learn the physics of protein-ligand interactions

Matthew R. Masters, Amr H. Mahmoud, Markus A. Lill
October 6, 2025Nature Communications63 citations

63

Citations

2

Influential Citations

Nature Communications

Venue

2025

Year

Abstract

Abstract Co-folding models represent a major innovation in deep-learning-based protein-ligand structure prediction. The recent publications of RoseTTAFold All-Atom, AlphaFold3, and others have shown high-quality results on predicting the structures of proteins interacting with small-molecules, nucleic-acids, and other proteins. Despite these advanced capabilities and broad potential, the current study presents critical findings that question the adherence of these models to fundamental physical principles. Through adversarial examples based on established physical, chemical, and biological principles, we demonstrate notable discrepancies in protein-ligand structural predictions when subjected to biologically and chemically plausible perturbations. These discrepancies reveal a significant divergence from expected physical behaviors, indicating potential overfitting to particular data features within its training corpus. Our findings underscore the models’ limitations in generalizing effectively across diverse protein-ligand structures and highlight the necessity of integrating robust physical and chemical priors in the development of such predictive tools. The results advocate a measured reliance on deep-learning-based models for critical applications in drug discovery and protein engineering, where a deep understanding of the underlying physical and chemical properties is crucial.

Analysis

Why This Paper Matters

This paper strikes at the heart of a critical assumption in computational biology: that deep learning models trained on massive structural data implicitly learn the underlying physics of molecular interactions. As co-folding models like AlphaFold3 and RoseTTAFold All-Atom are increasingly adopted for drug discovery and protein engineering, the community has largely celebrated their impressive predictive accuracy. However, this work provides a sobering counterpoint by systematically demonstrating that these models often fail to respect basic physical principles when faced with subtle, biologically plausible perturbations.

The significance lies in the method: instead of merely reporting aggregate accuracy metrics, the authors craft adversarial examples grounded in real physical chemistry. This approach reveals that the models are not learning physics but rather memorizing statistical correlations in the training data. For practitioners deploying these models in high-stakes settings, this paper is a wake-up call that high accuracy on benchmarks does not guarantee physical plausibility or generalization to novel systems.

Technical Contributions

  • Adversarial perturbation framework: The authors design perturbations that are chemically and biologically plausible (e.g., small rotations, bond length changes) yet violate physical laws, allowing systematic testing of model robustness.
  • Physics-based evaluation metrics: Instead of relying solely on RMSD or other geometric measures, the study uses physical criteria such as steric clashes, bond angle deviations, and violation of van der Waals radii.
  • Cross-model comparison: The analysis spans multiple state-of-the-art co-folding models, providing a comprehensive view of the problem rather than targeting a single architecture.
  • Overfitting diagnosis: By showing that models fail on simple perturbations that should not affect physically grounded predictions, the authors provide strong evidence of overfitting to training data features.

Results

The paper reports that all tested co-folding models produce predictions with notable physical discrepancies under adversarial perturbations. For example, models predict steric clashes and unrealistic bond angles that would be energetically impossible in real protein-ligand complexes. The discrepancies are consistent across different model architectures, suggesting a systemic issue rather than a bug in one particular model. The authors do not provide numerical metrics in the abstract, but the qualitative findings are striking: models that achieve high accuracy on standard benchmarks fail dramatically on these physically motivated adversarial examples.

Significance

This work has immediate implications for the AI and computational biology communities. It challenges the prevailing narrative that deep learning models can replace physics-based simulation in drug discovery. For AI practitioners, it underscores the importance of evaluating models not just on held-out test sets but on their ability to generalize to physically meaningful variations. The paper advocates for a hybrid approach that combines deep learning with explicit physical priors, a direction that could lead to more robust and trustworthy models. In the broader context, this research serves as a cautionary tale about the limits of pattern recognition in scientific domains where physical laws are paramount.