Search bioRxiv⌕ Search

Biology subjects

Segura Aba, K.

Publications and source records attributed to Segura Aba, K..

3 recordsLinked to original sources

Global Genotype by Environment Prediction Competition Reveals That Diverse Modeling Strategies Can Deliver Satisfactory Maize Yield Estimates

Predicting phenotypes from a combination of genetic and environmental factors is a grand challenge of modern biology. Slight improvements in this area have the potential to save lives, improve food and fuel security, permit better care of the planet, and create other positive outcomes. In 2022 and 2023 the first open-to-the-public Genomes to Fields (G2F) initiative Genotype by Environment (GxE) prediction competition was held using a large dataset including genomic variation, phenotype and weather measurements and field management notes, gathered by the project over nine years. The competition attracted registrants from around the world with representation from academic, government, industry, and non-profit institutions as well as unaffiliated. These participants came from diverse disciplines include plant science, animal science, breeding, statistics, computational biology and others. Some participants had no formal genetics or plant-related training, and some were just beginning their graduate education. The teams applied varied methods and strategies, providing a wealth of modeling knowledge based on a common dataset. The winners strategy involved two models combining machine learning and traditional breeding tools: one model emphasized environment using features extracted by Random Forest, Ridge Regression and Least-squares, and one focused on genetics. Other high-performing teams methods included quantitative genetics, classical machine learning/deep learning, mechanistic models, and model ensembles. The dataset factors used, such as genetics; weather; and management data, were also diverse, demonstrating that no single model or strategy is far superior to all others within the context of this competition.

genetics↗

Prediction of plant complex traits via integration of multi-omics data

The formation of complex traits is the consequence of genotype and activities at multiple molecular levels. However, connecting genotypes and these activities to complex traits remains challenging. Here, we investigated whether integrating different omics data could improve trait prediction. We built prediction models using genomic, transcriptomic, and methylomic data from the Arabidopsis 1001 Genomes Project for six Arabidopsis traits, and found that transcriptome- and methylome-based models had performances comparable to those of genome-based models. However, when comparing models for flowering time prediction, we found that models built using different omics data identified different benchmark genes. Nine novel genes identified as important for flowering time from our models were experimentally validated as regulating flowering. In addition, we found that gene contributions to flowering time prediction are accession-dependent and that distinct genes contribute to trait prediction in different genetic backgrounds. Models integrating multi-omics data performed best and revealed known and novel gene interactions, extending knowledge about existing regulatory networks underlying flowering time determination. These results demonstrate the feasibility of revealing molecular mechanisms underlying complex traits through multi-omics data integration.

bioinformatics↗

The topological shape of gene expression across the evolution of flowering plants

Since they emerged ~125 million years ago, flowering plants have evolved to dominate the terrestrial landscape and survive in the most inhospitable environments on earth. At their core, these adaptations have been shaped by changes in numerous, interconnected pathways and genes that collectively give rise to emergent biological phenomena. Linking gene expression to morphological outcomes remains a grand challenge in biology, and new approaches are needed to begin to address this gap. Here, we implemented topological data analysis (TDA) to summarize the high dimensionality and noisiness of gene expression data using lens functions that delineate plant tissue and stress responses. Using this framework, we created a topological representation of the shape of gene expression across plant evolution, development, and environment for the phylogenetically diverse flowering plants. The TDA-based Mapper graphs form a well-defined gradient of tissues from leaves to seeds, or from healthy to stressed samples, depending on the lens function. This suggests there are distinct and conserved expression patterns across angiosperms that delineate different tissue types or responses to biotic and abiotic stresses. Genes that correlate with the tissue lens function are enriched in central processes such as photosynthetic, growth and development, housekeeping, or stress responses. Together, our results highlight the power of TDA for analyzing complex biological data and reveal a core expression backbone that defines plant form and function. Significance statementA grand challenge in biology is to link gene expression to phenotypes across evolution, development, and the environment, but efforts have been hindered by biological complexity and dataset heterogeneity. Here, we implemented topological data analysis across thousands of gene expression datasets in phylogenetically diverse flowering plants. We created a topological representation of gene expression across plants and observed well-defined gradients of tissues from leaves to seeds, or from healthy to environmentally stressed. Using this framework, we identified a core and deeply conserved expression backbone that defines plant form and function, with key patterns that delineate plant tissues, abiotic, and biotic stresses. Our results highlight the power of topological approaches for analyzing complex biological datasets.

plant biology↗