PrePPI – Structure-based prediction of protein-protein interactomes and networks

Numerous databases of protein-protein interactions (PPIs) are available in the literature. Some, such as STRING [1], BioGRID [2], APID [1], and HINT [3], are based primarily on curation from multiple sources while others are derived primarily from high throughput experimental techniques, such as Affinity Purification-Mass Spectroscopy (AP-MS) 4, 5 and Yeast Two Hybrid (Y2H) [6]. The type of interaction – physical versus indirect – cannot easily be extracted from many of these resources. Moreover, the scale of protein interactome space – 9 million possible pairwise protein-protein interaction (PPI) combinations for the Escherichia coli K12 (E coli) proteome, 18 million for Saccharomyces cerevisiae (yeast) and 200 million for human – has so far limited the full interactome coverage of even the most efficient experimental approaches. Moreover, although some databases contain entries from the Protein Data Bank (PDB) [7], coverage of PPI complexes in the PDB is far from complete.

Experimentally derived databases do not typically provide structural models of complexes although a number of databases attempt to fill this gap. Interactome3D [8] contains PDB structures and high-confidence homology models thus increasing overall structural coverage. Interactome INSIDER [9] provides predictions for interfacial residues in experimentally determined PPI complexes while CM2D3 [10] uses comparative modeling, docking and AlphaFold2 to provide structural models for physical interactions extracted from experimental databases. Thus, although comprising extremely valuable resources, these structural databases do not contain de novo predictions.

Some computational approaches, such as Topsy-Turvy [11], are efficient enough to be applied to entire interactomes but do not provide structural models of PPI complexes. AlphaFold-based 12, 13 methods have been used to predict the structures of binary complexes but are too computationally intensive to be applied to full interactomes and, moreover, are of uncertain reliability when used to determine whether two proteins interact. To address these problems, Cong and coworkers [14] reported a computational pipeline that screens the human interactome for PPIs where the last step involves AlphaFold2 predictions of atomistic models. The pipeline yielded about 7,000 high quality de novo predictions and a total of ∼18,000 predictions when information from high throughput experiments were included in the pipeline. About 5,500 of the predicted PPIs have not, to date, been experimentally observed. These predictions revealed new biological insights but did not significantly expand the structural coverage of the human protein interactome.

In contrast to existing experimental and computational approaches, the PrePPI 15, 16, 17, 18 computational pipeline was designed to screen up to billions of potential binary physical interactions and to output a large number of predictions. For example, as described below and depicted in Fig. 1 for the human interactome, the current PrePPI pipeline screens all possible PPIs in the human, yeast and E. coli proteomes yielding, at a false positive rate (FPR) of < 0.005, 735,000 PPIs for human, 65,000 for yeast and 39,000 for E. coli. These predictions, most of which are novel in the sense that they do not appear in curated databases, are available in our online database. PrePPI models provide a powerful means of both hypothesis generation and hypothesis testing because the domains involved as well as interfacial contacts for each PPI are provided. The focus on binary interactions leads to a “bottom up” approach to PPI prediction since multi-protein complexes as well as non-physical genetic interactions will inevitably have an origin in binary interactions either within a complex or through a network of PPIs.

The PrePPI database has been completely redesigned to improve ease of use and to incorporate new functionalities. A number of new features are of particular note. First, the PrePPI database now contains interactomes for three organisms − E. coli, yeast and human – with more planned. Second, in addition to PPIs involving two structured domains, PrePPI now includes predicted interactions between structured peptide recognition domains (PRDs) and short linear motifs (SLiMs) [19] which are functionally related in the Eukaryotic Linear Motif (ELM) database [20] as Pfam domains and regular expressions, respectively (Fig. 1). In the new website, PrePPI PRD-SLiM predictions are linked to PDB complexes in Propedia [21] that have similar Pfam domains and peptidic motifs, thus providing a structural context for these interactions. Third, we have found that clustering the PrePPI interactomes reveals functionally coherent clusters that enable the identification of previously unassigned functions for many proteins and provide networks of binary PPIs associated with specific biological functions [22]. These clusters and their function annotations are conveniently accessible with a visual display that can be interactively interrogated. Fourth, the new website offers visual display of PrePPI-predicted structural models of complexes with the latest implementation of the RCSB PDB Mol* 3D Viewer [23] for further structural analysis. Finally, users can query the database for either single proteins or protein pairs. Together, these functionalities allow researchers to interactively explore PrePPI interactomes and apply the associated analyses to their biological questions. All results – from high-confidence interactomes to structural models to protein sequence features to clustered subnetworks and annotations – can be downloaded. The database and all tools can be accessed on our website, https://honigcomplab.c2b2.columbia.edu/PrePPI, and a set of video tutorials guides users through the various functionalities.

Comments (0)

No login
gif