Search bioRxiv⌕ Search

Biology subjects

Hossen, M. B.

Publications and source records attributed to Hossen, M. B..

3 recordsLinked to original sources

Structural alphabets approach performance of structural alignment in remote homology detection

MotivationRemote homology detection (RHD) is central to fold recognition and protein function annotation. While structural alignments provide a gold standard, they are computationally expensive. Encoding protein structures as sequences over structural alphabets offers a scalable alternative, but the relative performance of simple secondary-structure alphabets versus higher-resolution representations remains unclear. ResultsWe systematically compare 20-letter (3Di), 8-letter (Q8), and 3-letter (Q3) structural alphabets across three large-scale fold recognition benchmarks of increasing difficulty, using both advanced and basic sequence alignment algorithms. All three alphabets perform close to structural alignment gold standards and substantially outperform sequence-based methods. Remarkably, the minimal Q3 alphabet, distinguishing only helices, strands, and loops, achieves robust performance. We further demonstrate the practical utility of this finding in a protein function annotation task for a newly sequenced genome. Data AvailabilityBenchmark data are freely available at https://doi.org/10.6084/m9.figshare.c.8208161. Contactmichael.schroeder@tu-dresden.de Supplementary InformationSupplementary data are available online at the journal website.

bioinformatics↗

Protein secondary structure and remote homology detection

1A protein can be represented by its primary, secondary, or tertiary structure. With recent advances in AI, there is now as much tertiary as primary structural data available. Fast and accurate search methods exist for both types of data, with searches over both representations being highly precise. However, primary structure data can sometimes be incomplete. As a result, tertiary structure has become the gold standard for remote homology detection. How does secondary structure perform in remote homology detection? Secondary structure interprets proteins as a sequence using an alphabet representing helices, strands, or loops. It shares its sequential nature with primary structure while retaining topological information similar to tertiary structure. To assess the effectiveness of secondary structure in remote homology detection, we devised a challenging classification task aimed at determining the superfamily membership of very distantly related protein domains. We used benchmarks from the CATH and SCOP databases and evaluated sequence and structure alignment algorithms on primary, secondary, and tertiary structures. As expected, both basic and advanced sequence alignment algorithms applied to primary structure achieved high precision, but their overall area under the curve was lower compared to the gold standard of structural alignment using tertiary structure. Surprisingly, a simple string comparison algorithm applied to secondary structure performed close to the gold standard. This result supports the hypothesis that key structural information is already encoded in secondary structure and suggests that secondary structure may be a promising representation to use when high-confidence structural data is unavailable, such as in cases involving protein flexibility and disorder.

bioinformatics↗

The Rad52 superfamily as seen by AlphaFold

1Rad52, a highly conserved eukaryotic protein, plays a crucial role in DNA repair, especially in double-strand break repair. Recent findings reveal that its distinct structural features, including a characteristic {beta}-sheet and {beta}-hairpin motif, are shared with the lambda phage single-strand annealing proteins, Red{beta}, indicating a common superfamily. Our analysis of over 10,000 single-strand annealing proteins (SSAPs) across all kingdoms of life supports this hypothesis, confirming their possession of the characteristic motif despite variations in size and composition. We found that archaea, representing only 1% of the studied proteins, exhibit most of these variations. Through the examination of four representative archaeal SSAPs, we elucidate the structural relationship between eukaryotic and bacterial SSAPs, highlighting differences in {beta}-sheet size and {beta}-hairpin complexity. Furthermore, we identify an archaeal SSAP with a structure nearly identical to the human variant and screen over 100 million unannotated proteins for potential SSAP candidates. Our computational analysis complements existing sequence with structural evidence supporting the suggested orthology among five SSAP families across all kingdoms: Rad52, Red{beta}, RecT, Erf, and Sak3.

bioinformatics↗