Search bioRxiv⌕ Search

Biology subjects

El-Hendi, F.

Publications and source records attributed to El-Hendi, F..

3 recordsLinked to original sources

Structural alphabets approach performance of structural alignment in remote homology detection

MotivationRemote homology detection (RHD) is central to fold recognition and protein function annotation. While structural alignments provide a gold standard, they are computationally expensive. Encoding protein structures as sequences over structural alphabets offers a scalable alternative, but the relative performance of simple secondary-structure alphabets versus higher-resolution representations remains unclear. ResultsWe systematically compare 20-letter (3Di), 8-letter (Q8), and 3-letter (Q3) structural alphabets across three large-scale fold recognition benchmarks of increasing difficulty, using both advanced and basic sequence alignment algorithms. All three alphabets perform close to structural alignment gold standards and substantially outperform sequence-based methods. Remarkably, the minimal Q3 alphabet, distinguishing only helices, strands, and loops, achieves robust performance. We further demonstrate the practical utility of this finding in a protein function annotation task for a newly sequenced genome. Data AvailabilityBenchmark data are freely available at https://doi.org/10.6084/m9.figshare.c.8208161. Contactmichael.schroeder@tu-dresden.de Supplementary InformationSupplementary data are available online at the journal website.

bioinformatics↗

Much ado about nothing: Modelling amino acid replacement with predicted protein structures

Substitution matrices like BLOSUM62 model the likelihood of replacement of amino acids in evolution. Substitution matrices are used in protein sequence alignment tasks. Since the introduction of BLOSUM62 over three decades ago, many matrices have been released. Yet, to date, no effort uses large amounts of 3D structures predicted by AlphaFold. Here, we define AFSM, the AlphaFold Substitution Matrix derived from over 20,000 predicted 3D structures following the BLOSUM methodology. We benchmark AFSM against BLOSUM62 and 16 other matrices on five tasks in multiple sequence alignment (MSA) and protein homology search. Our analysis surprisingly reveals that all matrices perform similarly. Only when there are few sequences in an MSA, then BLOSUM62 and AFSM perform better than using no matrix. This suggests that substitution matrices were most beneficial when there was little sequence data. We corroborate this argument by showing that embeddings, which are computed from billions of sequences, perform better than substitution matrices, when sequence data is sparse. Taken together, this suggests that structural data does not improve BLOSUM62. But increased sequence data makes extrapolation with substitution matrices obsolete. Nonetheless, BLOSUM62 continues to capture chemists intuition on amino acids by providing numerical values implicitly reflecting physicochemical properties, and it remains indispensable for direct comparison of two sequences.

bioinformatics↗

Protein secondary structure and remote homology detection

1A protein can be represented by its primary, secondary, or tertiary structure. With recent advances in AI, there is now as much tertiary as primary structural data available. Fast and accurate search methods exist for both types of data, with searches over both representations being highly precise. However, primary structure data can sometimes be incomplete. As a result, tertiary structure has become the gold standard for remote homology detection. How does secondary structure perform in remote homology detection? Secondary structure interprets proteins as a sequence using an alphabet representing helices, strands, or loops. It shares its sequential nature with primary structure while retaining topological information similar to tertiary structure. To assess the effectiveness of secondary structure in remote homology detection, we devised a challenging classification task aimed at determining the superfamily membership of very distantly related protein domains. We used benchmarks from the CATH and SCOP databases and evaluated sequence and structure alignment algorithms on primary, secondary, and tertiary structures. As expected, both basic and advanced sequence alignment algorithms applied to primary structure achieved high precision, but their overall area under the curve was lower compared to the gold standard of structural alignment using tertiary structure. Surprisingly, a simple string comparison algorithm applied to secondary structure performed close to the gold standard. This result supports the hypothesis that key structural information is already encoded in secondary structure and suggests that secondary structure may be a promising representation to use when high-confidence structural data is unavailable, such as in cases involving protein flexibility and disorder.

bioinformatics↗