Search bioRxivSearch

Biology subjects

Buckler, E. S.

Publications and source records attributed to Buckler, E. S..

13 recordsLinked to original sources

Metabolome-scale genome-wide association studies reveal chemical diversity and genetic control of maize specialized metabolites

One Sentence SummaryHPLC-MS metabolite profiling of maize seedlings, in combination with genome-wide association studies, identifies numerous quantitative trait loci that influence the accumulation of foliar metabolites.\n\nAbstractCultivated maize (Zea mays) retains much of the genetic and metabolic diversity of its wild ancestors. Non-targeted HPLC-MS metabolomics using a diverse panel of 264 maize inbred lines identified a bimodal distribution in the prevalence of foliar metabolites. Although 15% of the detected mass features were present in >90% of the inbred lines, the majority were found in <50% of the samples. Whereas leaf bases and tips were differentiated primarily by flavonoid abundance, maize varieties (stiff-stalk, non-stiff-stalk, tropical, sweet corn, and popcorn) were differentiated predominantly by benzoxazinoid metabolites. Genome-wide association studies (GWAS), performed for 3,991 mass features from the leaf tips and leaf bases, showed that 90% have multiple significantly associated loci scattered across the genome. Several quantitative trait locus hotspots in the maize genome regulate the abundance of multiple, often metabolically related mass features. The utility of maize metabolite GWAS was demonstrated by confirming known benzoxazinoid biosynthesis genes, as well as by mapping isomeric variation in the accumulation of phenylpropanoid hydroxycitric acid esters to a single linkage block in a citrate synthase-like gene. Similar to gene expression databases, this metabolomic GWAS dataset constitutes an important public resource for linking maize metabolites with biosynthetic and regulatory genes.

genetics

RNA polymerase mapping in plants identifies enhancers enriched in causal variants

Promoter-proximal pausing and divergent transcription at promoters and enhancers, which are prominent features in animals, have been reported to be absent in plants based on a study of Arabidopsis thaliana. Here, our PRO-Seq analysis in cassava (Manihot esculenta) identified peaks of transcriptionally-engaged RNA polymerase II (Pol2) at both 5 and 3 ends of genes, consistent with paused or slowly-moving Pol2, and divergent transcription at potential intragenic enhancers. A full genome search for bi-directional transcription using an algorithm for enhancer detection developed in mammals (dREG) identified many enhancer candidates. These sites show distinct patterns of methylation and nucleotide variation based on genomic evolutionary rate profiling characteristic of active enhancers. Maize GRO-Seq data showed RNA polymerase occupancy at promoters and enhancers consistent with cassava but not Arabidopsis. Furthermore, putative enhancers in maize identified by dREG significantly overlapped with sites previously identified on the basis of open chromatin, histone marks, and methylation. We show that SNPs within these divergently transcribed intergenic regions predict significantly more variation in fitness and root composition than SNPs in chromosomal segments randomly ascertained from the same intergenic distribution, suggesting a functional importance of these sites on cassava. The findings shed new light on plant transcription regulation and its impact on development and plasticity.

genomics

Evolutionarily informed deep learning methods: Predicting transcript abundance from DNA sequence

Deep learning methodologies have revolutionized prediction in many fields, and show potential to do the same in molecular biology and genetics. However, applying these methods in their current forms ignores evolutionary dependencies within biological systems and can result in false positives and spurious conclusions. We developed two novel approaches that account for evolutionary relatedness in machine learning models: 1) gene-family guided splitting, and 2) ortholog contrasts. The first approach accounts for evolution by constraining the models training and testing sets to include different gene families. The second, uses evolutionarily informed comparisons between orthologous genes to both control for and leverage evolutionary divergence during the training process. The two approaches were explored and validated within the context of mRNA expression level prediction, and have prediction auROC values ranging from 0.72 to 0.94. Model weight inspections showed biologically interpretable patterns, resulting in the novel hypothesis that the 3 UTR is more important for fine tuning mRNA abundance levels while the 5 UTR is more important for large scale changes.

molecular biology

Leveraging mutational burden for complex trait prediction in sorghum

Sorghum (Sorghum bicolor (L.) Moench) is a major staple food cereal for millions of people worldwide. The sorghum genome, like other species, accumulates deleterious mutations, likely impacting its fitness. Though selection keeps deleterious mutations rare, their complete removal from the genome is impeded due to lack of recombination, drift, and their coupling with favorable loci. To study how deleterious mutations impact agronomic phenotypes, we identified putative deleterious mutations among ~5.5M segregating variants of 229 diverse sorghum lines. We provide the whole-genome estimate of the deleterious burden in sorghum, showing that about 33% of nonsynonymous substitutions are putatively deleterious. The pattern of mutation burden varies appreciably among racial groups; the caudatum shows higher mutation burden while the guinea has lower burden. Across racial groups, the mutation burden correlated negatively with biomass, plant height, Specific Leaf Area (SLA), and tissue starch content, suggesting deleterious burden decreases trait fitness. Putatively deleterious variants explain roughly half of the genetic variance. However, there is only moderate improvement in total heritable variance explained for biomass (7.6%) and plant height (5.2%). There is no advantage in total heritable variance for SLA and starch. The contribution of putatively deleterious variants to phenotypic diversity therefore appears to be dependent on the genetic architecture of traits. Overall, our results suggest that including putatively deleterious variants in models do not significantly improve breeding accuracy because of extensive linkage. However, knowledge of deleterious variants could be leveraged for sorghum breeding through genome editing.

genetics

Ethylene signaling regulates natural variation in the abundance of antifungal acetylated diferuloylsucroses and Fusarium graminearum resistance in maize seedling roots

O_LIThe production and regulation of defensive specialized metabolites plays a central role in pathogen resistance in maize (Zea mays) and other plants. Therefore, identification of genes involved in plant specialized metabolism can contribute to improved disease resistance.\nC_LIO_LIWe used comparative metabolomics to identify previously unknown antifungal metabolites in maize seedling roots, and investigated the genetic and physiological mechanisms underlying their natural variation using quantitative trait locus (QTL) mapping and comparative transcriptomics approaches.\nC_LIO_LITwo maize metabolites, smilaside A (3,6-diferuloyl-3',6'-diacetylsucrose) and smiglaside C (3,6-diferuloyl-2',3',6'-triacetylsucrose), that may contribute to maize resistance against Fusarium graminearum and other fungal pathogens were identified. Elevated expression of an ethylene receptor gene, ETHYLENE INSENSITIVE 2 (ZmEIN2), co-segregated with decreased smilaside A/smiglaside C ratio. Pharmacological and genetic manipulation of ethylene availability and sensitivity in vivo indicated that, whereas ethylene was required for the production of both metabolites, the smilaside A/smiglaside C ratio was negatively regulated by ethylene sensitivity. This ratio, rather than the absolute abundance of these two metabolites, was important for maize seedling root defense against F. graminearum.\nC_LIO_LIEthylene signaling regulates the relative abundance of the two F. graminearum-resistance-related metabolites and affects resistance against F. graminearum in maize seedling roots.\nC_LI

plant biology

Tripsacum de novo transcriptome assemblies reveal parallel gene evolution with maize after ancient polyploidy

Plant genomes reduce in size following a whole genome duplication event, and one gene in a duplicate gene pair can lose function in absence of selective pressure to maintain duplicate gene copies. Maize and its sister genus, Tripsacum, share a genome duplication event that occurred 5 to 26 million years ago. Because few genomic resources for Tripsacum exist, it is unknown whether Tripsacum grasses and maize have maintained a similar set of genes under purifying selection. Here we present high quality de novo transcriptome assemblies for two species: Tripsacum dactyloides and Tripsacum floridanum. Genes with experimental protein evidence in maize were good candidates for genes under purifying selection in both genera because pseudogenes by definition do not produce protein. We tested whether 15,160 maize genes with protein evidence are resisting gene loss and whether their Tripsacum homologs are also resisting gene loss. Protein-encoding maize transcripts and their Tripsacum homologs have higher GC content, higher gene expression levels, and more conserved expression levels than putatively untranslated maize transcripts and their Tripsacum homologs. These results indicate that gene loss is occurring in a similar fashion in both genera after a shared ancient polyploidy event. The Tripsacum transcriptome assemblies provide a high quality genomic resource that can provide insight into the evolution of maize, an highly valuable crop worldwide.\n\nCore ideasO_LIMaize genes with protein evidence have higher expression and GC content\nC_LIO_LITripsacum homologs of maize genes exhibit the same trends as in maize\nC_LIO_LIMaize proteome genes have more highly correlated gene expression with Tripsacum\nC_LIO_LIExpression dominance for homeologs occurs similarly between maize and Tripsacum\nC_LIO_LIA similar set of genes may be decaying into pseudogenes in maize and Tripsacum\nC_LI

genomics

Quantitative Genetic Analysis of the Maize Leaf Microbiome

The degree to which an organism can affect its associated microbial communities (\"microbiome\") varies by organism and habitat, and in many cases is unknown. We address this question by analyzing the metabolically active bacteria of the maize phyllosphere across 300 diverse maize lines growing in a common environment. We performed comprehensive heritability analysis for 49 community diversity metrics, 380 bacterial clades (individual operational taxonomic units and higher-level groupings), and 9042 predicted metagenomic functions. We find that only a few few bacterial clades (5) and diversity metrics (2) are significantly heritable, while a much larger number of metabolic functions (200) are. Many of these associations appear to be driven by the amount of Methylobacteria present in each sample, and we find significant enrichment for traits relating to short-chain carbon metabolism, secretion, and nitrotoluene degradation. Genome-wide association analysis identifies a small number of associated loci for these heritable traits, including two loci (on maize chromosomes 7 and 10) that affect a large number of traits even after correcting for correlations among traits. This work is among the most comprehensive analyses of the maize phyllosphere to date. Our results indicate that while most of the maize phyllosphere composition is driven by environmental factors and/or stochastic founder events, a subset of bacterial taxa and metabolic functions is nonetheless significantly impacted by host plant genetics. Additional work will be needed to identify the exact nature of these interactions and what effects they may have on the phenotype of host plants.

genetics

k-mer grammar uncovers maize regulatory architecture

Only a small percentage of the genome sequence is involved in regulation of gene expression, but to biochemically identify this portion is expensive and laborious. In species like maize, with diverse intergenic regions and lots of repetitive elements, this is an especially challenging problem. While regulatory regions are rare, they do have characteristic chromatin contexts and sequence organization (the grammar) with which they can be identified. We developed a computational framework to exploit this sequence arrangement. The models learn to classify regulatory regions based on sequence features - k-mers. To do this, we borrowed two approaches from the field of natural language processing: (1) \"bag-of-words\" which is commonly used for differentially weighting key words in tasks like sentiment analyses, and (2) a vector-space model using word2vec (vector-k-mers), that captures semantic and linguistic relationships between words. We built \"bag-of-k-mers\" and \"vector-k-mers\" models that distinguish between regulatory and non-regulatory regions with an accuracy above 90%. Our \"bag-of-k-mers\" achieved higher overall accuracy, while the \"vector-k-mers\" models were more useful in highlighting key groups of sequences within the regulatory regions. These models now provide powerful tools to annotate regulatory regions in other maize lines beyond the reference, at low cost and with high accuracy.

plant biology

Genetic Analysis of Lodging in Diverse Maize Hybrids

Damage caused by lodging is a significant problem in corn production that results in estimated annual yield losses of 5-20%. Over the past 100 years, substantial maize breeding efforts have increased lodging resistance by artificial selection. However, less research has focused on understanding the genetic architecture underlying lodging. Lodging is a problematic trait to evaluate since it is greatly influenced by environmental factors such as wind, rain, and insect infestation, which make replication difficult. In this study over 1,723 diverse inbred maize genotypes were crossed to a common tester and evaluated in five environments over multiple years. Natural lodging due to severe weather conditions occurred in all five environments. By testing a large population of genetically diverse maize lines in multiple field environments, we detected significant correlations for this highly environmentally influenced trait across environments and with important agronomic traits such as yield and plant height. This study also permitted the mapping of quantitative trait loci (QTL) for lodging. Several QTL identified in this study overlapped with loci previously mapped for stalk strength in related maize inbred lines. QTL intervals mapped in this study also overlapped candidate genes implicated in the regulation of lignin and cellulose synthesis.

genetics

rAmpSeq: Using repetitive sequences for robust genotyping

Repetitive sequences have been used for DNA fingerprinting and genotyping for more than a quarter century. Now, with our knowledge of whole genome sequences, repetitive sequences can be used to identify polymorphisms that can be mapped and scored in a systematic manner. We have developed a simple, robust platform for designing primers, PCR amplification, and high throughput cloning that allows hundreds to thousands of markers to be scored for less than $5 per sample. Conserved regions were used to design PCR primers for amplifying thousands of middle repetitive regions of the maize (Zea mays ssp. mays) genome. Bioinformatic scans were then used to identify DNA sequence polymorphisms in the low copy intervening sequences. When used in conjunction with simple DNA preps, optimized PCR conditions, high multiplex Illumina indexing and a bioinformatic marker calling platform tailored for repetitive sequences, this methodology provides a cost effective genotyping strategy for large-scale genomic selection projects. We show detailed results from four maize primer sets that produced between 1,335-3,225 good coverage loci with 1056 that segregated appropriately in a bi-parental family. This approach could have wide applicability to breeding and conservation biology, where hundreds of thousands of samples need to be genotyped for very minimal cost.

genomics

Identifying the diamond in the rough: a study of allelic diversity underlying flowering time adaptation in maize landraces

Landraces (traditional varieties) of crop species are a reservoir of useful genetic diversity, yet remain untapped due to the genetic linkage between the few useful alleles with hundreds of undesirable alleles1. We integrated two approaches to characterize the genetic diversity of over 3000 maize landraces from across the Americas. First, we mapped the genomic regions controlling latitudinal and altitudinal adaptation, identifying 1498 genes. Second, we developed and used F-One Association Mapping (FOAM) to directly map genes controlling flowering time across 22 environments, identifying 1,005 genes. In total 65% of the SNPs associated with altitude were also associated with flowering time. In particular, we observed many of the significant SNPs were contained in large structural variants (inversions, centromeres, and pericentromeric regions): 29.4% for flowering time, 58.4% for altitude and 13.1% for latitude. The combined mapping results indicate that while floral regulatory network genes contribute substantially to field variation, over 90% of contributing genes likely have indirect effects. Our strategy can be used to harness the diversity of maize and other plant and animal species.

genetics

A large scale joint analysis of flowering time reveals independent temperate adaptations in maize

Modulating days to flowering is a key mechanism in plants for adapting to new environments, and variation in days to flowering drives population structure by limiting mating. To elucidate the genetic architecture of flowering across maize, a quantitative trait, we mapped flowering in five global populations, a diversity panel (Ames) and four half-sib mapping designs, Chinese (CNNAM), US (USNAM), and European Dent (EUNAM-Dent) and Flint (EUNAM-Flint). Using whole-genome projected SNPs, we tested for joint association using GWAS, resampling GWAS and two regional approaches; Regional Heritability Mapping (RHM) (1, 2) and a novel method, Boosted Regional Heritability Mapping (BRHM). Direct overlap in significant regions detected between populations and flowering candidate genes was limited, but whole-genome cross-population predictive abilities were [&le;]0.78. Poor predictive ability correlated with increased population differentiation (r = 0.41), unless the parents were broadly sampled from across the North American temperate-tropical germplasm gradient; uncorrected GWAS results from populations with broadly sampled parents were well predicted by temperate-tropical FSTs in machine learning. Machine learning between GWAS results also suggested shared architecture between the American panels and, more distantly, the European panels, but not the Chinese panel. Machine learning approaches can reconcile non-linear relationships, but the combined predictive ability of all of the populations did not significantly enhance prediction of physiological candidates. While the North American-European temperate adaption is well studied, this study suggest independent temperate adaptation evolved in the Chinese panel, most likely in China after 1500, a finding supported by differential gene ontology term enrichment between populations.

genetics

Incomplete dominance of deleterious alleles contribute substantially to trait variation and heterosis in maize

AbstractDeleterious alleles have long been proposed to play an important role in patterning phenotypic variation and are central to commonly held ideas explaining the hybrid vigor observed in the offspring by crossing two inbred parents. We test these ideas using evolutionary measures of sequence conservation to ask whether incorporating information about putatively deleterious alleles can inform genomic selection (GS) models and improve phenotypic prediction. We measured a number of agronomic traits in both the inbred parents and hybrids of an elite maize partial diallel population and re-sequenced the parents of the population. Inbred elite maize lines vary for more than 350,000 putatively deleterious sites, but show a lower burden of such sites than a comparable set of traditional landraces. Our modeling reveals widespread evidence for incomplete dominance at these loci, and supports theoretical models that more damaging variants are usually more recessive. We identify haplotype blocks using an identity-by-decent (IBD) analysis and perform genomic prediction analyses in which we weigh blocks on the basis of segregating putatively deleterious variants. Cross-validation results show that incorporating sequence conservation in genomic selection improves prediction accuracy for grain yield and other fitness-related traits as well as heterosis for those traits. Our results provide empirical support for an important role for incomplete dominance of deleterious alleles in explaining heterosis and demonstrate the utility of incorporating functional annotation in phenotypic prediction and plant breeding.\n\nAuthor SummaryA key long-term goal of biology is understanding the genetic basis of phenotypic variation. Although most new mutations are likely disadvantageous, their prevalence and importance in explaining patterns of phenotypic variation is controversial and not well understood. In this study we combine whole genome-sequencing and field evaluation of a maize mapping population to investigate the contribution of deleterious mutations to phenotype. We show that a priori prediction of deleterious alleles correlates well with effect sizes for grain yield and that variants predicted to be more damaging are on average more recessive. We develop a simple model allowing for variation in the heterozygous effects of deleterious mutations and demonstrate its improved ability to predict both phenotypes and hybrid vigor. Our results help reconcile alternative explanations for hybrid vigor and highlight the use of leveraging evolutionary history to facilitate breeding for crop improvement.

genetics