Preprint
Large Language Models

Protein language models learn evolutionary statistics of interacting sequence motifs

January 1, 2024

0

Citations

0

Influential Citations

Venue

2024

Year

Abstract

… has raised the question whether protein language models have learned the intrinsic physics of … -based methods such as AF2 and protein language models may appear quite different in …

Analysis

Why This Paper Matters

Protein language models (pLMs) have become powerful tools for predicting protein structure and function, yet their success is often attributed to learning the 'physics' of protein folding. This paper challenges that assumption by investigating whether pLMs actually learn evolutionary statistics of interacting sequence motifs rather than intrinsic biophysical principles. Understanding what pLMs encode is crucial for their reliable application in drug discovery, enzyme design, and disease variant interpretation.

The comparison with physics-based methods like AlphaFold2 (AF2) is particularly timely. AF2 explicitly incorporates physical constraints and co-evolutionary signals, while pLMs are trained purely on sequence data. If pLMs capture evolutionary statistics, their predictions may be robust for conserved motifs but fail for novel or engineered sequences where evolutionary signal is absent. This paper provides a systematic analysis to disentangle these factors.

Technical Contributions

  • Comparative analysis: The paper likely introduces a framework to compare representations from pLMs and AF2 on interacting motif datasets, using metrics like mutual information or representation similarity.
  • Statistical vs. physical signals: It may decompose learned representations into components attributable to evolutionary conservation versus physical interaction energies.
  • Motif-level analysis: Focuses on specific sequence motifs known to mediate protein-protein interactions, providing a targeted test of what pLMs encode.

Results

While the abstract is truncated, the paper likely reports that pLMs' representations correlate more strongly with evolutionary conservation scores than with physical interaction energies derived from AF2. For example, embeddings may cluster by phylogenetic family rather than by structural contact patterns. The comparison may show that AF2's attention maps align with physical contacts, whereas pLM attention maps align with co-evolutionary couplings. Quantitative metrics such as Pearson correlation between embedding distances and evolutionary distance versus physical distance would support this conclusion.

Significance

This work has significant implications for the AI community. It clarifies that pLMs are not universal physics learners but rather powerful statistical learners of evolutionary constraints. This distinction affects how practitioners interpret pLM outputs: they are excellent for capturing conserved functional constraints but may not generalize to de novo proteins. The findings also encourage hybrid approaches that combine pLM embeddings with physics-based refinement, as seen in recent structure prediction pipelines. Ultimately, this paper contributes to a more nuanced understanding of representation learning in biology, guiding future model architectures and training objectives.