[Paper Review] How different are self and nonself?
This paper proposes that self and nonself peptides are statistically similar but highly inhomogeneous in sequence space, leading the immune system to target peptides nearly identical to self—especially those one amino acid away—due to overfitting to the finite set of self peptides. This explains strong immune responses to cancer neoantigens and the immunodominance of near-self peptides.
Biological and artificial networks routinely make reliable distinctions between similar inputs, and the rules for making these distinctions are learned. In some ways, self/nonself discrimination in the immune system is similar, being both reliable and (partly) learned through thymic selection. In contrast to other examples, we show that the distributions of self and nonself peptides are nearly identical but strongly inhomogeneous. Reliable discrimination is possible only because self-peptides are a particular finite sample drawn out of this distribution, and T cells can target the spaces in between these samples. In conventional learning problems, this would constitute overfitting and lead to disaster. Here, the strong inhomogeneities imply instead that the immune system gains by targeting peptides which are similar to self, with maximum sensitivity for sequences just one or two substitutions away. This prediction from the structure of the underlying distribution in sequence space agrees, for example, with the observed responses to mutation derived cancer neoantigens.
Motivation & Objective
- To understand why cytotoxic T cells can respond strongly to nonself peptides that differ from self by only one amino acid.
- To investigate the statistical structure of the 9-mer peptide distribution in the human proteome and pathogenic viruses.
- To determine how the inhomogeneous distribution of peptides in sequence space enables reliable self/nonself discrimination despite similarity.
- To explain the immunodominance of near-self peptides and the success of immune responses against cancer neoantigens.
Proposed method
- Used maximum entropy modeling to infer the statistical distribution of 9-mer peptides from the human proteome and viral proteomes.
- Incorporated constraints incrementally: amino acid frequencies, pairwise correlations, and higher-order moments to build increasingly accurate models.
- Modeled the distribution of viral peptides relative to the nearest human peptide in sequence space to assess similarity and coincidence rates.
- Compared model predictions with empirical data on immunogenicity from the Immune Epitope Database to validate predictions on near-self peptide recognition.
- Defined a 'shell model' of immunogenicity where immune responses are biased toward peptides just outside the densest regions of self-peptide space.
- Used the model to show that overfitting to the finite set of self peptides is not a flaw but a necessary strategy for effective nonself detection.
Experimental results
Research questions
- RQ1Why do cytotoxic T cells respond strongly to nonself peptides that differ from self by only one amino acid substitution?
- RQ2How does the statistical distribution of 9-mer peptides in the human proteome influence self/nonself discrimination?
- RQ3To what extent do viral peptides resemble human peptides in sequence space, and how does this affect immune recognition?
- RQ4Why is overfitting to self-peptides beneficial rather than detrimental in immune recognition?
- RQ5How does the inhomogeneous structure of the peptide distribution in sequence space enable reliable discrimination between self and nonself?
Key findings
- The distribution of 9-mer peptides in the human proteome is strongly inhomogeneous, with significant correlations extending across all nine positions despite short peptide length.
- More than 0.1% of viral 9-mer peptides are exact matches to human peptides, and over 1% are one amino acid away—10–100 times more frequent than expected under uniform randomness.
- The maximum entropy model accurately predicts the frequency of near-self matches, showing that the inhomogeneity of the peptide distribution underlies the observed biases.
- Peptides one amino acid away from self are more likely to be immunogenic than those with three or more substitutions, as confirmed by data from the Immune Epitope Database.
- The immune system gains by overfitting to the finite set of self peptides, as this allows it to target dense regions of nonself sequence space where near-self peptides are most abundant.
- The 'shell model' of immunogenicity explains that the highest density of nonself targets lies just outside the regions of highest self-peptide density, making near-self peptides optimal immune targets.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.