ArticleslgStudy

biology

Protein–protein interaction prediction

Protein–protein interaction prediction is a biology topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Protein–protein interaction prediction rather than just read about it. In short: Protein–protein interaction prediction is a field combining bioinformatics and structural biology in an attempt to identify and catalog physical interactions between pairs or groups of proteins. Understanding protein–protein interactions is important for the investigation of intracellular signaling pathways, modelling of protein complex structures and for gaining insights into various biochemical processes.

Protein–protein interaction prediction — main illustration
Protein–protein interaction prediction — illustration

Key takeaways

  • Protein–protein interaction prediction belongs to biology; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Protein–protein interaction prediction to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Protein–protein interaction prediction from memory before moving on to harder problems.

Reference excerpt

Protein–protein interaction prediction is a field combining bioinformatics and structural biology in an attempt to identify and catalog physical interactions between pairs or groups of proteins. Understanding protein–protein interactions is important for the investigation of intracellular signaling pathways, modelling of protein complex structures and for gaining insights into various biochemical processes. Experimentally, physical interactions between pairs of proteins can be inferred from a variety of techniques, including yeast two-hybrid systems, protein-fragment complementation assays (PCA), affinity purification/mass spectrometry, protein microarrays, fluorescence resonance energy transfer (FRET), and Microscale Thermophoresis (MST). Efforts to experimentally determine the interactome of numerous species are ongoing. Experimentally determined interactions usually provide the basis for computational methods to predict interactions, e.g. using homologous protein sequences across species. However, there are also methods that predict interactions de novo, without prior knowledge of existing interactions.

Methods Proteins that interact are more likely to co-evolve, therefore, it is possible to make inferences about interactions between pairs of proteins based on their phylogenetic distances. It has also been observed in some cases that pairs of interacting proteins have fused orthologues in other organisms. In addition, a number of bound protein complexes have been structurally solved and can be used to identify the residues that mediate the interaction so that similar motifs can be located in other organisms.

Phylogenetic profiling

The phylogenetic profile method is based on the hypothesis that if two or more proteins are concurrently present or absent across several genomes, then they are likely functionally related. Figure A illustrates a hypothetical situation in which proteins A and B are identified as functionally linked due to their identical phylogenetic profiles across 5 different genomes. The Joint Genome Institute provides an Integrated Microbial Genomes and Microbiomes database (JGI IMG) that has a phylogenetic profiling tool for single genes and gene cassettes.

Prediction of co-evolved protein pairs based on similar phylogenetic trees It was observed that the phylogenetic trees of ligands and receptors were often more similar than due to random chance. This is likely because they faced similar selection pressures and co-evolved. This method uses the phylogenetic trees of protein pairs to determine if interactions exist. To do this, homologs of the proteins of interest are found (using a sequence search tool such as BLAST) and multiple-sequence alignments are done (with alignment tools such as Clustal) to build distance matrices for each of the proteins of interest. The distance matrices should then be used to build phylogenetic trees. However, comparisons between phylogenetic trees are difficult, and current methods circumvent this by simply comparing distance matrices. The distance matrices of the proteins are used to calculate a correlation coefficient, in which a larger value corresponds to co-evolution. The benefit of comparing distance matrices instead of phylogenetic trees is that the results do not depend on the method of tree building that was used. The downside is that difference matrices are not perfect representations of phylogenetic trees, and inaccuracies may result from using such a shortcut. Another factor worthy of note is that there are background similarities between the phylogenetic trees of any protein, even ones that do not interact. If left unaccounted for, this could lead to a high false-positive rate. For this reason, certain methods construct a background tree using 16S rRNA sequences which they use as the canonical tree of life. The distance matrix constructed from this tree of life is then subtracted from the distance matrices of the proteins of interest. However, because RNA distance matrices and DNA distance matrices have different scale, presumably because RNA and DNA have different mutation rates, the RNA matrix needs to be rescaled before it can be subtracted from the DNA matrices. By using molecular clock proteins, the scaling coefficient for protein distance/RNA distance can be calculated. This coefficient is used to rescale the RNA matrix.

Rosetta stone (gene fusion) method The Rosetta Stone or Domain Fusion method is based on the hypothesis that interacting proteins are sometimes fused into a single protein. For instance, two or more separate proteins in a genome may be identified as fused into one single protein in another genome. The separate proteins are likely to interact and thus are likely functionally related. An example of this is the Human Succinyl coA Transferase enzyme, which is found as one protein in humans but as two separate proteins, Acetate coA Transferase alpha and Acetate coA Transferase beta, in Escherichia coli. In order to identify these sequences, a sequence similarity algorithm such as the one used by BLAST is necessary. For example, if we had the amino acid sequences of proteins A and B and the amino acid sequences of all proteins in a certain genome, we could check each protein in that genome for non-overlapping regions of sequence similarity to both proteins A and B. Figure B depicts the BLAST sequence alignment of Succinyl coA Transferase with its two separate homologs in E. coli. The two subunits have non-overlapping regions of sequence similarity with the human protein, indicated by the pink regions, with the alpha subunit similar to the first half of the protein and the beta similar to the second half. One limit of this method is that not all proteins that interact can be found fused in another genome, and therefore cannot be identified by this method. On the other hand, the fusion of two proteins does not necessitate that they physically interact. For instance, the SH2 and SH3 domains in the src protein are known to interact. However, many proteins possess homologs of these domains and they do not all interact.

… excerpt ends here. Continue reading the full article.

Illustrations

Protein–protein interaction prediction: Figure B. The Human succinyl-CoA-Transferase enzyme is represented by the two joint blue and green bars at the top of the image. The alpha subunit of the Acetate-CoA-Transferase enzyme is homologous with the first half of the enzyme, represents by the blue bar. The beta subunit of the Acetate-CoA-Transferase enzyme is homologous with the second half of the enzyme, represents by the green bar. This mage was adapted from Uetz, P. & Pohl, E. (2018) Protein–Protein and Protein–DNA Interactions. In: Wink, M. (ed.), Introduction to Molecular Biotechnology, 3rd ed. Wiley-VCH, in press.
Figure B. The Human succinyl-CoA-Transferase enzyme is represented by the two joint blue and green bars at the top of the image. The alpha subunit of the Acetate-CoA-Transferase enzyme is homologous with the first half of the enzyme, represents by the blue bar. The beta subunit of the Acetate-CoA-Transferase enzyme is homologous with the second half of the enzyme, represents by the green bar. This mage was adapted from Uetz, P. & Pohl, E. (2018) Protein–Protein and Protein–DNA Interactions. In: Wink, M. (ed.), Introduction to Molecular Biotechnology, 3rd ed. Wiley-VCH, in press.
Protein–protein interaction prediction: FigureC. Organization of the trp operon in three different species of bacteria: Escherichia coli, Haemophilus influenzae, Helicobacter pylori. Only the trpA and trpB genes are adjacent across all three organisms and are thus predicted to interact by the conserved gene neighborhood method. This image was adapted from Dandekar, T., Snel, B., Huynen, M., & Bork, P. (1998). Conservation of gene order: a fingerprint of proteins that physically interact. Trends in biochemical sciences, 23(9), 324-328.[1]
FigureC. Organization of the trp operon in three different species of bacteria: Escherichia coli, Haemophilus influenzae, Helicobacter pylori. Only the trpA and trpB genes are adjacent across all three organisms and are thus predicted to interact by the conserved gene neighborhood method. This image was adapted from Dandekar, T., Snel, B., Huynen, M., & Bork, P. (1998). Conservation of gene order: a fingerprint of proteins that physically interact. Trends in biochemical sciences, 23(9), 324-328.[1]

Worked examples

Example 1 — a first encounter with Protein–protein interaction prediction

Start with the simplest possible case. Write down what Protein–protein interaction prediction claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In biology, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Protein–protein interaction prediction before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Protein–protein interaction prediction ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Protein–protein interaction prediction

In research
Protein–protein interaction prediction appears in biology research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Protein–protein interaction prediction in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Protein–protein interaction prediction is common in secondary-school and first-year university syllabi. It links to neighbouring topics Proteomics, so understanding it makes those chapters shorter.
In everyday life
Look for Protein–protein interaction prediction outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.
Ask Teacher Smith questions about this articleOpens your AI tutor with a question about “Protein–protein interaction prediction” →

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Protein–protein interaction prediction in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Protein–protein interaction prediction means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Protein–protein interaction prediction out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Protein–protein interaction prediction in simple terms?

Protein–protein interaction prediction is a field combining bioinformatics and structural biology in an attempt to identify and catalog physical interactions between pairs or groups of proteins. Understanding protein–protein interactions is important for the investigation of intracellular signaling…

Why does Protein–protein interaction prediction matter?

Because it connects several biology ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Protein–protein interaction prediction?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Protein–protein interaction prediction.

Tags

  • Proteomics

Keep exploring