Search bioRxiv⌕ Search

Biology subjects

Ferebee, T.

Publications and source records attributed to Ferebee, T..

2 recordsLinked to original sources

Current genomic deep learning architectures generalize across grass species but not alleles

Non-coding regions of the genome are just as important as coding regions for understanding the mapping from genotype to phenotype. Interpreting deep learning models trained on RNA-seq is an emerging method to highlight functional sites within non-coding regions. Most of the work on RNA abundance models has been done within humans and mice, with little attention paid to plants. Here, we benchmark four genomic deep learning model architectures with genomes and RNA-seq data from 18 species closely related to maize and sorghum within the Andropogoneae. The Andropogoneae are a tribe of C4 grasses that have adapted to a wide range of environments worldwide since diverging 18 million years ago. Hundreds of millions of years of evolution across these species has produced a large, diverse pool of training alleles across species sharing a common physiology. As model input, we extracted 1,026 base pairs upstream of each genes translation start site. We held out maize as our test set and two closely related species as our validation set, training each architecture on the remaining Andropogoneae genomes. Within a panel of 26 maize lines, all architectures predict expression across genes moderately well but poorly across alleles. DanQ consistently ranked highest or second highest among all architectures yet performance was generally very similar across architectures despite orders of magnitude differences in size. This suggests that state-of-the-art supervised genomic deep learning models are able to generalize moderately well across related species but not sensitively separate alleles within species, the latter of which agrees with recent work within humans. We are releasing the preprocessed data and code for this work as a community benchmark to evaluate new architectures on our across-species and across-allele tasks.

genomics↗

Elucidating the patterns of pleiotropy and its biological relevance in maize

Pleiotropy - when a single gene controls two or more seemingly unrelated traits - has been shown to impact genes with effects on flowering time, leaf architecture, and inflorescence morphology in maize. However, the genome-wide impact of true biological pleiotropy across all maize phenotypes is largely unknown. Here we investigate the extent to which biological pleiotropy impacts phenotypes within maize through GWAS summary statistics reanalyzed from previously published metabolite, field, and expression phenotypes across the Nested Association Mapping population and Goodman Association Panel. Through phenotypic saturation of 120,597 traits, we obtain over 480 million significant quantitative trait nucleotides. We estimate that only 1.56-32.3% of intervals show some degree of pleiotropy. We then assessed the relationship between pleiotropy and various biological features such as gene expression, chromatin accessibility, sequence conservation, and enrichment for gene ontology terms. We find very little relationship between pleiotropy and these variables when compared to permuted pleiotropy. We hypothesize that biological pleiotropy of common alleles is not widespread in maize and is highly impacted by nuisance terms such as population structure and linkage disequilibrium. Natural selection on large standing natural variation in maize populations may target wide- and large-effect variants, leaving the prevalence of detectable pleiotropy relatively low. Author SummaryThe genetic basis of complex traits has been thought to exhibit pleiotropy, which is the notion that a single locus can control two or more unrelated traits. Widespread reports in the human disease literature show genomic signatures of pleiotropic loci across many traits. However, little is known about the prevalence and behavior of pleiotropy in maize across a large number of phenotypes. Using association mapping of common alleles in over one hundred thousand traits, we determine how pleiotropic each region was and use these quantitative scores to functionally characterize each region of the genome. Our results show little evidence that pleiotropy is a common phenomenon in maize. We observed that maize does not exhibit the same pleiotropic characteristics as human diseases in terms of prevalence, gene expression, chromatin accessibility, or sequence conservation. Rather than pervasive pleiotropy, we hypothesize that strong selection on large and wide effect loci and the need for trait independence at the gene level keep the prevalence of pleiotropy low, thus, allowing for the adaptation of maize varieties to novel environments and conditions.

plant biology↗