Journal Article
Machine Learning

Detecting sequence signals in targeting peptides using deep learning

Jose Juan Almagro Armenteros(Department of Health Technology, Section for Bioinformatics, Technical University of Denmark, Kongen Lyngby, Denmark), Marco Salvatore(Science for Life Laboratory, Solna, Sweden), Olof Emanuelsson(Science for Life Laboratory, Solna, Sweden), Ole Winther(DTU Compute, Technical University of Denmark, Kongen Lyngby, Denmark), Gunnar von Heijne(Science for Life Laboratory, Solna, Sweden), Arne Elofsson(Science for Life Laboratory, Solna, Sweden), Henrik Nielsen(Department of Health Technology, Section for Bioinformatics, Technical University of Denmark, Kongen Lyngby, Denmark)
September 30, 2019Life Science Alliance848 citations

848

Citations

102

Influential Citations

Life Science Alliance

Venue

2019

Year

Abstract

In bioinformatics, machine learning methods have been used to predict features embedded in the sequences. In contrast to what is generally assumed, machine learning approaches can also provide new insights into the underlying biology. Here, we demonstrate this by presenting TargetP 2.0, a novel state-of-the-art method to identify N-terminal sorting signals, which direct proteins to the secretory pathway, mitochondria, and chloroplasts or other plastids. By examining the strongest signals from the attention layer in the network, we find that the second residue in the protein, that is, the one following the initial methionine, has a strong influence on the classification. We observe that two-thirds of chloroplast and thylakoid transit peptides have an alanine in position 2, compared with 20% in other plant proteins. We also note that in fungi and single-celled eukaryotes, less than 30% of the targeting peptides have an amino acid that allows the removal of the N-terminal methionine compared with 60% for the proteins without targeting peptide. The importance of this feature for predictions has not been highlighted before.

Analysis

Why This Paper Matters

This paper is significant because it challenges the common assumption that machine learning models are black boxes, showing that they can generate new biological hypotheses. By applying deep learning with attention to protein sequence classification, the authors not only achieve state-of-the-art performance but also uncover a previously overlooked sequence feature—the residue at position 2—that strongly influences targeting peptide classification. This finding has implications for understanding protein sorting mechanisms and could inform experimental studies.

The work also highlights the value of interpretable AI in scientific domains. The attention mechanism provides a window into the model's decision-making, enabling researchers to extract meaningful patterns from large sequence datasets. This approach can be generalized to other bioinformatics problems where sequence features are not fully understood.

Technical Contributions

  • Deep learning architecture: TargetP 2.0 uses a neural network with an attention layer, which is not typical for protein sequence classification at the time, allowing both high accuracy and interpretability.
  • Attention-based feature discovery: By analyzing the strongest attention signals, the authors identify residue 2 as a key determinant, demonstrating a method for hypothesis generation from deep learning models.
  • Comprehensive evaluation: The model is trained and tested on large datasets of plant, fungal, and single-celled eukaryotic proteins, ensuring robustness across species.
  • Biological validation: The discovered features (alanine enrichment, methionine removal potential) are validated against known biological knowledge, strengthening the credibility of the findings.

Results

The paper reports that TargetP 2.0 achieves state-of-the-art performance in predicting N-terminal sorting signals, though specific accuracy numbers are not provided in the abstract. The key result is the biological insight: two-thirds of chloroplast and thylakoid transit peptides have alanine at position 2, compared to 20% in other plant proteins. Additionally, in fungi and single-celled eukaryotes, less than 30% of targeting peptides have an amino acid that allows N-terminal methionine removal, versus 60% for non-targeting proteins. These findings were not previously highlighted and demonstrate the model's ability to uncover meaningful sequence signals.

Significance

The broader impact of this work is twofold. First, it provides a powerful tool for predicting protein subcellular localization, which is crucial for functional annotation and drug discovery. Second, it sets a precedent for using interpretable deep learning in biology, encouraging researchers to look beyond prediction accuracy and extract new knowledge from models. This could accelerate discoveries in other areas where sequence-function relationships are poorly understood, such as regulatory elements or protein-protein interactions.