Protein–protein interaction prediction is a field combining bioinformatics and structural biology in an attempt to identify and catalog physical interactions between pairs or groups of proteins. Understanding protein–protein interactions is important for the investigation of intracellular signaling pathways, modelling of protein complex structures and for gaining insights into various biochemical processes. Experimentally, physical interactions between pairs of proteins can be inferred from a variety of techniques, including yeast two-hybrid systems, protein-fragment complementation assays (PCA), affinity purification/mass spectrometry, protein microarrays, fluorescence resonance energy transfer (FRET), and Microscale Thermophoresis (MST). Efforts to experimentally determine the interactome of numerous species are ongoing. Experimentally determined interactions usually provide the basis for computational methods to predict interactions, e.g. using homologous protein sequences across species. However, there are also methods that predict interactions de novo, without prior knowledge of existing interactions.
Methods Proteins that interact are more likely to co-evolve, therefore, it is possible to make inferences about interactions between pairs of proteins based on their phylogenetic distances. It has also been observed in some cases that pairs of interacting proteins have fused orthologues in other organisms. In addition, a number of bound protein complexes have been structurally solved and can be used to identify the residues that mediate the interaction so that similar motifs can be located in other organisms.
Phylogenetic profiling
The phylogenetic profile method is based on the hypothesis that if two or more proteins are concurrently present or absent across several genomes, then they are likely functionally related. Figure A illustrates a hypothetical situation in which proteins A and B are identified as functionally linked due to their identical phylogenetic profiles across 5 different genomes. The Joint Genome Institute provides an Integrated Microbial Genomes and Microbiomes database (JGI IMG) that has a phylogenetic profiling tool for single genes and gene cassettes.
Prediction of co-evolved protein pairs based on similar phylogenetic trees It was observed that the phylogenetic trees of ligands and receptors were often more similar than due to random chance. This is likely because they faced similar selection pressures and co-evolved. This method uses the phylogenetic trees of protein pairs to determine if interactions exist. To do this, homologs of the proteins of interest are found (using a sequence search tool such as BLAST) and multiple-sequence alignments are done (with alignment tools such as Clustal) to build distance matrices for each of the proteins of interest. The distance matrices should then be used to build phylogenetic trees. However, comparisons between phylogenetic trees are difficult, and current methods circumvent this by simply comparing distance matrices. The distance matrices of the proteins are used to calculate a correlation coefficient, in which a larger value corresponds to co-evolution. The benefit of comparing distance matrices instead of phylogenetic trees is that the results do not depend on the method of tree building that was used. The downside is that difference matrices are not perfect representations of phylogenetic trees, and inaccuracies may result from using such a shortcut. Another factor worthy of note is that there are background similarities between the phylogenetic trees of any protein, even ones that do not interact. If left unaccounted for, this could lead to a high false-positive rate. For this reason, certain methods construct a background tree using 16S rRNA sequences which they use as the canonical tree of life. The distance matrix constructed from this tree of life is then subtracted from the distance matrices of the proteins of interest. However, because RNA distance matrices and DNA distance matrices have different scale, presumably because RNA and DNA have different mutation rates, the RNA matrix needs to be rescaled before it can be subtracted from the DNA matrices. By using molecular clock proteins, the scaling coefficient for protein distance/RNA distance can be calculated. This coefficient is used to rescale the RNA matrix.
Rosetta stone (gene fusion) method The Rosetta Stone or Domain Fusion method is based on the hypothesis that interacting proteins are sometimes fused into a single protein. For instance, two or more separate proteins in a genome may be identified as fused into one single protein in another genome. The separate proteins are likely to interact and thus are likely functionally related. An example of this is the Human Succinyl coA Transferase enzyme, which is found as one protein in humans but as two separate proteins, Acetate coA Transferase alpha and Acetate coA Transferase beta, in Escherichia coli. In order to identify these sequences, a sequence similarity algorithm such as the one used by BLAST is necessary. For example, if we had the amino acid sequences of proteins A and B and the amino acid sequences of all proteins in a certain genome, we could check each protein in that genome for non-overlapping regions of sequence similarity to both proteins A and B. Figure B depicts the BLAST sequence alignment of Succinyl coA Transferase with its two separate homologs in E. coli. The two subunits have non-overlapping regions of sequence similarity with the human protein, indicated by the pink regions, with the alpha subunit similar to the first half of the protein and the beta similar to the second half. One limit of this method is that not all proteins that interact can be found fused in another genome, and therefore cannot be identified by this method. On the other hand, the fusion of two proteins does not necessitate that they physically interact. For instance, the SH2 and SH3 domains in the src protein are known to interact. However, many proteins possess homologs of these domains and they do not all interact.
… excerpt ends here. Continue reading the full article.


![Protein–protein interaction prediction: FigureC. Organization of the trp operon in three different species of bacteria: Escherichia coli, Haemophilus influenzae, Helicobacter pylori. Only the trpA and trpB genes are adjacent across all three organisms and are thus predicted to interact by the conserved gene neighborhood method. This image was adapted from Dandekar, T., Snel, B., Huynen, M., & Bork, P. (1998). Conservation of gene order: a fingerprint of proteins that physically interact. Trends in biochemical sciences, 23(9), 324-328.[1]](https://upload.wikimedia.org/wikipedia/commons/thumb/e/ef/Trp_operon_organization_across_three_different_bacterial_species.png/500px-Trp_operon_organization_across_three_different_bacterial_species.png?utm_source=en.wikipedia.org&utm_campaign=parser&utm_content=thumbnail)
