Search bioRxiv⌕ Search

Biology subjects

Hayford, R. K.

Publications and source records attributed to Hayford, R. K..

2 recordsLinked to original sources

PanEffect: A pan-genome visualization tool for variant effects in maize

Understanding the effects of genetic variants is crucial for accurately predicting traits and phenotypic outcomes. Recent advances have utilized protein language models to score all possible missense variant effects at the proteome level for a single genome, but a reliable tool is needed to explore these effects at the pan-genome level. To address this gap, we introduce a new tool called PanEffect. We implemented PanEffect at MaizeGDB to enable a comprehensive examination of the potential effects of coding variants across 51 maize genomes. The tool allows users to visualize over 550 million possible amino acid substitutions in the B73 maize reference genome and also to observe the effects of the 2.3 million natural variations in the maize pan-genome. Each variant effect score, calculated from the Evolutionary Scale Modeling (ESM) protein language model, shows the log-likelihood ratio difference between B73 and all variants in the pan-genome. These scores are shown using heatmaps spanning benign outcomes to strong phenotypic consequences. Additionally, PanEffect displays secondary structures and functional domains along with the variant effects, offering additional functional and structural context. Using PanEffect, researchers now have a platform to explore protein variants and identify genetic targets for crop enhancement. Availability and implementation: The PanEffect code is freely available on GitHub (https://github.com/Maize-Genetics-and-Genomics-Database/PanEffect). A maize implementation of PanEffect and underlying datasets are available at MaizeGDB (https://www.maizegdb.org/effect/maize/).

bioinformatics↗

FASSO: An AlphaFold based method to assign functional annotations by combining sequence and structure orthology

Methods to predict orthology play an important role in bioinformatics for phylogenetic analysis by identifying orthologs within or across any level of biological classification. Sequence-based reciprocal best hit approaches are commonly used in functional annotation since orthologous genes are expected to share functions. The process is limited as it relies solely on sequence data and does not consider structural information and its role in function. Previously, determining protein structure was highly time-consuming, inaccurate, and limited to the size of the protein, all of which resulted in a structural biology bottleneck. With the release of AlphaFold, there are now over 200 million predicted protein structures, including full proteomes for dozens of key organisms. The reciprocal best structural hit approach uses protein structure alignments to identify structural orthologs. We propose combining both sequence- and structure-based reciprocal best hit approaches to obtain a more accurate and complete set of orthologs across diverse species, called Functional Annotations using Sequence and Structure Orthology (FASSO). Using FASSO, we annotated orthologs between five plant species (maize, sorghum, rice, soybean, Arabidopsis) and three distance outgroups (human, budding yeast, and fission yeast). We inferred over 270,000 functional annotations across the eight proteomes including annotations for over 5,600 uncharacterized proteins. FASSO provides confidence labels on ortholog predictions and flags potential misannotations in existing proteomes. We further demonstrate the utility of the approach by exploring the annotation of the maize proteome.

bioinformatics↗