Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Comparative genomics of ten new Caenorhabditis species

The nematode Caenorhabditis elegans has been central to the understanding of metazoan biology. However, C. elegans is but one species among millions and the significance of this important model organism will only be fully revealed if it is placed in a rich evolutionary context. Global sampling efforts have led to the discovery of over 50 putative species from the genus Caenorhabditis, many of which await formal species description. Here, we present species descriptions for ten new Caenorhabditis species. We also present draft genome sequences for nine of these new species, along with a transcriptome assembly for one. We exploit these whole-genome data to reconstruct the Caenorhabditis phylogeny and use this phylogenetic tree to dissect the evolution of morphology in the genus. We show unexpected complexity in the evolutionary history of key developmental pathway genes. The genomic data also permit large scale analysis of gene structure, which we find to be highly variable within the genus. These new species and the associated genomic resources will be essential in our attempts to understand the evolutionary origins of the C. elegans model.

evolutionary biology

Characterization of emetic and diarrheal Bacillus cereus strains from a 2016 foodborne outbreak using whole-genome sequencing: addressing the microbiological, epidemiological, and bioinformatic challenges

The Bacillus cereus group comprises multiple species capable of causing emetic or diarrheal foodborne illness. Despite being responsible for tens of thousands of illnesses each year in the U.S. alone, whole-genome sequencing (WGS) has not been routinely employed to characterize B. cereus group isolates from foodborne outbreaks. Here, we describe the first WGS-based characterization of isolates linked to an outbreak caused by members of the B. cereus group. In conjunction with a 2016 outbreak traced to a supplier of refried beans served by a fast food restaurant chain in upstate New York, a total of 33 B. cereus group strains were obtained from human cases (n =7) and food samples (n = 26). Emetic (n = 30) and diarrheal (n = 3) isolates were most closely related to B. paranthracis (clade III) and B. cereus sensu stricto (clade IV), respectively. WGS indicated that the 30 emetic isolates (24 and 6 from food and humans, respectively) were closely-related and formed a well-supported clade relative to publicly-available emetic clade III genomes with an identical sequence type (ST 26). When compared to publicly-available emetic clade III ST 26 B. cereus group genomes, the 30 emetic clade III isolates from this outbreak differed from each other by a mean of 8.3 to 11.9 core single nucleotide polymorphisms (SNPs), while differing from publicly-available genomes by a mean of 301.7 to 528.0 core SNPs, depending on the SNP calling methodology used. Using a WST-1 cell proliferation assay, the strains isolated from this outbreak had only mild detrimental effects on HeLa cell metabolic activity compared to reference diarrheal strain B. cereus ATCC 14579. Based on both WGS and epidemiological data, we hypothesize that the outbreak was a single source outbreak caused by emetic clade III B. cereus belonging to the B. paranthracis species. In addition to showcasing how WGS can be used to characterize B. cereus group strains linked to a foodborne outbreak, we also discuss potential microbiological and epidemiological challenges presented by B. cereus group outbreaks, and we offer recommendations for analyzing WGS data from the isolates associated with them.

microbiology

Genome-wide association studies on endometriosis and endometriosis-related infertility

Endometriosis affects [~]10% of women of reproductive age. It is characterized by the growth of endometrial-like tissue outside the uterus and is frequently associated with severe pain and infertility. We performed the largest endometriosis genome-wide association study (GWAS) to date, with 37,183 cases and 251,258 controls. All women were of European ancestry. We also performed the first GWAS of endometriosis-related infertility, including 2,969 cases and 3,770 controls. Our endometriosis GWAS study replicated, at genome-wide significance, seven loci identified in previous endometriosis GWASs (CELA3A-CDC42, SYNE1, KDR, FSHB-ARL14EP, GREB1, ID4, and CEP112) and identified seven new candidate loci with genome-wide significance (NGF, ATP1B1-F5, CD109, HEY2, OSR2-VPS13B, WT1, and TEX11-SLC7A3). No loci demonstrated genome-wide significance for endometriosis-related infertility, however, the three most strongly associated loci (MCTP1, EPS8L3-CSF1, and LPIN1) were in or near genes associated with female fertility or embryonic lethality in model organisms. These results reveal new candidate genes with potential involvement in the pathophysiology of endometriosis and endometriosis-related infertility.

genetics

Elevated pyrimidine dimer formation at distinct genomic bases underlie promoter mutation hotspots in UV-exposed cancers

Sequencing of whole cancer genomes has revealed an abundance of recurrent mutations in gene-regulatory promoter regions, in particular in melanoma where strong mutation hotspots are observed adjacent to ETS-family transcription factor (TF) binding sites. While sometimes interpreted as functional driver events, these mutations have also been suggested to be due to locally inhibited DNA repair or, alternatively, locally increased propensity for UV damage. Here, we provide evidence that base-specific elevations in the efficacy of UV lesion formation underlie these mutations. First, we find that low-dose UV light induces mutations preferably at a known ETS promoter hotspot in cultured cells even in the absence of global or transcription-coupled nucleotide excision repair (NER), ruling out inhibited repair. Further, by genome-wide mapping of cyclobutane pyrimidine dimers (CPDs) shortly after UV exposure and thus before DNA repair, we find that ETS-related mutation hotspots exhibit a strong base-specific increase in CPD formation frequency. Analysis of a large whole genome cohort illustrates the widespread contribution of this effect to recurrent mutations in melanoma. While inhibited NER underlies a general increase in somatic mutation burden in regulatory regions, we conclude that the most recurrently mutated individual DNA bases arise instead due to locally favorable conditions for UV damage formation, thus explaining a key phenomenon in whole-genome cancer analyses.

cancer biology

Assessing genomic admixture between cryptic Plutella moth species following secondary contact

Cryptic species are genetically distinct taxa without obvious variation in morphology and are occasionally discovered using molecular or sequence datasets of populations previously thought to be a single species. The world-wide Brassica pest, Plutella xylostella (diamondback moth), has been a problematic insect in Australia since 1882, yet a morphologically cryptic species with apparent endemism (P. australiana) was only recognized in 2013. Plutella xylostella and P. australiana are able to hybridize under laboratory conditions, and it was unknown whether introgression of adaptive traits could occur in the field to improve fitness and potentially increase pressure on agriculture. Phylogenetic reconstruction of 29 nuclear genomes confirmed P. xylostella and P. australiana are divergent, and molecular dating with 13 mitochondrial genes estimated a common Plutella ancestor 1.96{+/-}0.175 MYA. Sympatric Australian populations and allopatric Hawaiian P. xylostella populations were used to test whether neutral or adaptive introgression had occurred between the two Australian species. We used three approaches to test for genomic admixture in empirical and simulated datasets including i) the f3 statistic at the level of the population, ii) pairwise comparisons of Neis absolute genetic divergence (dXY) between populations and iii) changes in phylogenetic branch lengths between individuals across 50 kb genomic windows. These complementary approaches all supported reproductive isolation of the Plutella species in Australia, despite their ability to hybridize. Finally, we highlight the most divergent genomic regions between the two cryptic Plutella species and find they contain genes involved with processes including digestion, detoxification and DNA binding.

evolutionary biology

Comprehensive Analysis of Indels in Whole-genome Microsatellite Regions and Microsatellite Instability across 21 Cancer Types

Microsatellites are repeats of 1-6bp units and [~]10 million microsatellites have been identified across the human genome. Microsatellites are vulnerable to DNA mismatch errors, and have thus been used to detect cancers with mismatch repair deficiency. To reveal the mutational landscape of the microsatellite repeat regions at the genome level, we analyzed approximately 20.1 billion microsatellites in 2,717 whole genomes of pan-cancer samples across 21 tissue types. Firstly, we developed a new insertion and deletion caller (MIMcall) that takes into consideration the error patterns of different types of microsatellites. Among the 2,717 pan-cancer samples, our analysis identified 31 samples, including colorectal, uterus, and stomach cancers, with higher microsatellite mutation rate ([≥] 0.03), which we defined as microsatellite instability (MSI) cancers in genome-wide level. Next, we found 20 highly-mutated microsatellites that can be used to detect MSI cancers with high sensitivity. Third, we found that replication timing and DNA shape were significantly associated with mutation rates of the microsatellites. Analysis of germline variation of the microsatellites suggested that the amount of germline variations and somatic mutation rates were correlated. Lastly, analysis of mutations in mismatch repair genes showed that somatic SNVs and short indels had larger functional impact than germline mutations and structural variations. Our analysis provides a comprehensive picture of mutations in the microsatellite regions, and reveals possible causes of mutations, as well as provides a useful marker set for MSI detection.

cancer biology

Classification and monomer-by-monomer annotation of suprachromosomal family 1 alpha satellite higher-order repeats in hg38 human genome assembly

In the latest hg38 human genome assembly, centromeric gaps has been filled in by alpha satellite (AS) reference models (RMs) which are statistical representations of homogeneous higher-order repeat (HOR) arrays that make up the bulk of the centromeric regions. We studied these models to compose an atlas of human HORs where each monomer of a HOR could be characterized and represented by a number of its polymorphic sequence variants. We further used these data and HMMER sequence analysis platform to annotate AS HORs in the assembly. This led to discovery and annotation of a new type of low copy number highly divergent HORs which were not represented by RMs. The annotation can be viewed as UCSC Genome Browser custom track (the HOR-track) and used together with our previous annotation of AS SFs in the same assembly where each AS monomer can be viewed in its genomic context together with its classification into one of the 5 major SFs (the SF-track). To catalog the diversity of AS HORs in the human genome we introduced a new naming system. Each HOR received a name which showed its SF, chromosomal location and index number. Here we present the first installment of the HOR-track covering only the 17 HORs that belong to SF1 which forms live functional centromeres in chromosomes 1, 3, 5, 6, 7, 10, 12, 16 and 19 and also a large number of minor dead HOR domains, both homogeneous (pseudo) and divergent (relic). The 4 newly discovered divergent SF1 HORs have provided the missing links in SF1 early evolution and substantiated its partition into 2 generations, archaic and modern, which we reported earlier. Additionally, we demonstrated that monomer-by-monomer HOR annotation was useful for mapping and quantification of various structural variants of AS HORs which would be important for studies of inter-individual polymorphism of AS including centromeric functional epialleles.

bioinformatics

Genome-Wide Control of Population Structure and Relatedness in Genetic Association Studies via Linear Mixed Models with Orthogonally Partitioned Structure

Linear mixed models (LMMs) have become the standard approach for genetic association testing in the presence of sample structure. However, the performance of LMMs has primarily been evaluated in relatively homogeneous populations of European ancestry, despite many of the recent genetic association studies including samples from worldwide populations with diverse ancestries. In this paper, we demonstrate that existing LMM methods can have systematic miscalibration of association test statistics genome-wide in samples with heterogenous ancestry, resulting in both increased type-I error rates and a loss of power. Furthermore, we show that this miscalibration arises due to varying allele frequency differences across the genome among populations. To overcome this problem, we developed LMM-OPS, an LMM approach which orthogonally partitions diverse genetic structure into two components: distant population structure and recent genetic relatedness. In simulation studies with real and simulated genotype data, we demonstrate that LMM-OPS is appropriately calibrated in the presence of ancestry heterogeneity and outperforms existing LMM approaches, including EMMAX, GCTA, and GEMMA. We conduct a GWAS of white blood cell (WBC) count in an admixed sample of 3,551 Hispanic/Latino American women from the Womens Health Initiative SNP Health Association Resource where LMM-OPS detects genome-wide significant associations with corresponding p-values that are one or more orders of magnitude smaller than those from competing LMM methods. We also identify a genome-wide significant association with regulatory variant rs2814778 in the DARC gene on chromosome 1, which generalizes to Hispanic/Latino Americans a previous association with reduced WBC count identified in African Americans.

genetics

Pan-cancer whole genome analyses of metastatic solid tumors

Metastatic cancer is one of the major causes of death and is associated with poor treatment efficiency. A better understanding of the characteristics of late stage cancer is required to help tailor personalised treatment, reduce overtreatment and improve outcomes. Here we describe the largest pan-cancer study of metastatic solid tumor genomes, including 2,520 whole genome-sequenced tumor-normal pairs, analyzed at a median depth of 106x and 38x respectively, and surveying over 70 million somatic variants. Metastatic lesions were found to be very diverse, with mutation characteristics reflecting those of the primary tumor types, although with high rates of whole genome duplication events (56%). Metastatic lesions are relatively homogeneous with the vast majority (96%) of driver mutations being clonal and up to 80% of tumor suppressor genes bi-allelically inactivated through different mutational mechanisms. For 62% of all patients, genetic variants that may be associated with outcome of approved or experimental therapies were detected. These actionable events were distributed across various mutation types underlining the importance of comprehensive genomic tumor profiling for cancer precision medicine.

cancer biology

NOJAH: Not Just Another Heatmap for Genome-Wide Cluster Analysis

Since their inception, several tools have been developed for cluster analysis and heatmap construction. The application of such tools to the number and types of genome-wide data available from next generation sequencing (NGS) technologies requires the adaptation of statistical concepts, such as in defining a most variable gene set, and more intricate cluster analyses method to address multiple omic data types. Additionally, the growing number of publicly available datasets has created the desire to estimate the statistical significance of a gene signature derived from one dataset to similarly group samples based on another dataset. The currently available number of tools and their combined use for generating heatmaps, along with the several adaptations of statistical concepts for addressing the higher dimensionality of genome-wide NGS-derived data, has created a further challenge in the ability to replicate heatmap results. We introduce NOJAH (NOt Just Another Heatmap), an interactive tool that defines and implements a workflow for genome-wide cluster analysis and heatmap construction by creating and combining several tools into a single user interface. NOJAH includes several newly developed scripts for techniques that though frequently applied are not sufficiently documented to allow for replicability of results. These techniques include: defining a most variable gene set (a.k.a., core genes), estimating the statistical significance of a gene signature to separate samples into clusters, and performing a result merging integrated cluster analysis. With only a user uploaded dataset, NOJAH provides as output, among other things, the minimum documentation required for replicating heatmap results. Additionally, NOJAH contains five different existing R packages that are connected in the interface by their functionality as part of a defined workflow for genome-wide cluster analysis. The NOJAH application tool is available at http://bbisr.shinyapps.winship.emory.edu/NOJAH/ with corresponding source code available at https://github.com/bbisr-shinyapps/NOJAH/.

bioinformatics

Genotyping by low-coverage whole-genome sequencing in intercross pedigrees from outbred founders: a cost efficient approach

BackgroundExperimental intercrosses between outbred founder populations are powerful resources for mapping loci contributing to complex traits (Quantitative Trait Loci or QTL). Here, we present an approach and accompanying software for high-resolution genotype imputation in such populations using whole-genome high coverage sequence data on founder individuals ([~]30x) and low coverage sequence data on intercross individuals ([~]0.4x). The method is illustrated in a large F2 pedigree between lines of chickens that have been divergently selected for 40 generations for the same trait (body weight at 8 weeks of age).\n\nResultsDescribed is how hundreds of individuals were whole-genome sequenced in a cost- and time-efficient manner using a Tn5-based library preparation protocol optimized for this application. In total, 7.6M markers segregated in this pedigree and 10.0 to 13.7% were informative for imputing the founder line genotypes within the F0-F2 families. The genotypes imputed from low coverage sequence data were consistent with the founder line genotypes estimated using SNP and microsatellite markers both at individual imputed sites (92%) and across the genome of individual chickens (93%). The resolution of the recombination breakpoints was high with 50% being resolved within <10kb.\n\nConclusionsA method for genotype imputation from low-coverage whole-genome sequencing in outbred intercrosses is described and evaluated. By applying it to an outbred chicken F2 cross it is illustrated that it provides high quality, high-resolution genotypes in a time and cost efficient manner.

genetics

Impact of cancer mutational signatures on transcription factor motifs in the human genome

BackgroundSomatic mutations in cancer genomes occur through a variety of molecular mechanisms, which contribute to different mutational patterns. To summarize these, mutational signatures have been defined using a large number of cancer genomes, and related to distinct mutagenic processes. Each cancer genome can be compared to this reference dataset and its exposure to one or the other signature be determined. Given the very different mutational patterns of these signatures, we anticipate that they will have distinct impact on genomic elements, in particular motifs for transcription factor binding sites (TFBS).\n\nResultsIn this work, we build the link between mutational signatures and TFBS motif alterations. We investigated and computed the theoretical impact of mutational signatures on 512 TFBS motifs, hence translating the trinucleotide mutation frequencies of the signatures into alteration frequencies of specific TFBS motifs, leading either to creation of disruption of these motifs. We further build a theoretical prediction of the alteration patterns for different cancer types based on the exposure of these cancer types to the mutation signatures. For certain motifs, a high correlation is observed between the TFBS motif creation and disruption events related to the information content of the motif.\n\nConclusionOur results show that the mutational signatures have different impact on the binding motifs of transcription factors and that for certain high complexity motifs there is a strong correlation between creation and disruption, related to the information content of the motif. This study represents a background estimation of the alterations due purely to mutational signatures in the absence of additional contributions, e.g. from evolutionary processes.

bioinformatics

Temperature preference biases parental genome retention during hybrid evolution

Interspecific hybridization can introduce genetic variation that aids in adaptation to new or changing environments. Here we investigate how the environment, and more specifically temperature, interacts with hybrid genomes to alter parental genome representation over time. We evolved Saccharomyces cerevisiae x Saccharomyces uvarum hybrids in nutrient-limited continuous culture at 15{degrees}C for 200 generations. In comparison to previous evolution experiments at 30{degrees}C, we identified a number of temperature specific responses, including the loss of the S. cerevisiae allele in favor of the cryotolerant S. uvarum allele for several portions of the hybrid genome. In particular, we discovered a genotype by environment interaction in the form of a reciprocal loss of heterozygosity event on chromosome XIII. Which species haplotype is lost or maintained is dependent on the parental species temperature preference and the temperature at which the hybrid was evolved. We show that a large contribution to this directionality is due to temperature sensitivity at a single locus, the high affinity phosphate transporter PHO84. This work helps shape our understanding of what forces impact genome evolution after hybridization, and how environmental conditions may favor or disfavor hybrids over time.

evolutionary biology

Telomeric dysfunction triggers an unstable growth arrest leading to irreparable genomic lesions and entry into cellular senescence

Replicative senescence is the permanent growth arrest caused by gradual telomere attrition occurring at each round of genome replication. Critically shortened telomeres lose their protective shelterin complex and t-loop structure revealing uncapped chromosome ends that are recognized as DNA double-strand breaks causing a p53-dependent DNA damage response (DDR) towards proliferation arrest. Because telomeres are heterogeneous in length within a single cell, the number of short telomeres necessary for senescence onset remains ill defined. Using controlled Tin2-mediated shelterin inactivation, we show that telomere uncapping is not sufficient to trigger senescence. While uncapping generates expected telomeric DNA damage detection, the associated weak DDR allows a rapid bypass of the primary growth arrest and re-entry into the cell cycle despite dysfunctional telomeres. During the ensuing mitosis, fused telomeres lead to additional DNA breaks and to genomic instability including chromosomes bridges or micronuclei, which sustain a secondary entry into stable growth arrest. The loss of p53 prevented both primary and secondary growth arrest, leading to amplified genomic instablility. Our results support a new multistep model for entry into telomere-mediated replicative senescence in normal cells, which is not directly induced by telomere uncapping, but rather by an amplification of DNA lesions caused by telomere fusions that leads to permanent irreparable genome damage.

cell biology

Genomic architecture of parallel ecological divergence: beyond a single environmental contrast

The genetic basis of parallel ecological divergence provides important clues to the operation of natural selection and the predictability of evolution. Many examples exist where binary environmental contrasts seem to drive parallel divergence. However, this simplified view can conceal important components of parallel divergence because environmental variation is often more complex. Here, we disentangle the genetic basis of parallel divergence across two axes of environmental differentiation (crab-predation vs. wave-action and low-shore vs. high-shore habitat contrasts) in the marine snail Littorina saxatilis, a well established natural system of parallel ecological divergence. We used whole-genome resequencing across multiple instances of these two environmental axes, at local and regional scales from Spain to Sweden. Overall, sharing of genetic differentiation is generally low but it is highly heterogeneous across the genome and increases at smaller spatial scales. We identified genomic regions, both overlapping and non-overlapping with recently described candidate chromosomal inversions, that are differentially involved in adaptation to each of the environmental axis. Thus, the evolution of parallel divergence in L. saxatilis is largely determined by the joint action of geography, history, genomic architecture and congruence between environmental axes. We argue that the maintenance of standing variation, perhaps as balanced polymorphism, and/or the re-distribution of adaptive variants via gene flow can facilitate parallel divergence in multiple directions as an adaptive response to heterogeneous environments.

evolutionary biology

A piggyBac-based toolkit for inducible genome editing in mammalian cells

We describe the development and application of a novel series of vectors that facilitate CRISPR-Cas9-mediated genome editing in mammalian cells, which we call CRISPR-Bac. CRISPR-Bac leverages the piggyBac transposon to randomly insert CRISPR-Cas9 components into mammalian genomes. In CRISPR-Bac, a single piggyBac cargo vector containing a doxycycline-inducible Cas9 or catalytically-dead Cas9 (dCas9) variant and a gene conferring resistance to Hygromycin B is co-transfected with a plasmid expressing the piggyBac transposase. A second cargo vector, expressing a single-guide RNA (sgRNA) of interest, the reverse-tetracycline TransActivator (rtTA), and a gene conferring resistance to G418, is also cotransfected. Subsequent selection on Hygromycin B and G418 generates polyclonal cell populations that stably express Cas9, rtTA, and the sgRNA(s) of interest. Using Mus musculus-derived embryonic and trophoblast stem cells, we show that CRISPR-Bac can be used to knockdown proteins of interest, to create targeted genetic deletions with high efficiency, and to activate or repress transcription of protein-coding genes and an imprinted long noncoding RNA. The ratio of sgRNA-to-Cas9-to-transposase can be adjusted in transfections to alter the average number of cargo insertions into the genome. sgRNAs targeting multiple genes can be inserted in a single transfection. CRISPR-Bac is a versatile platform for genome editing that simplifies the generation of mammalian cells that stably express the CRISPR-Cas9 machinery.

molecular biology

Assessing mitochondrial function in angiosperms with highly divergent mitochondrial genomes

Angiosperm mitochondrial (mt) genes are generally slow-evolving, but multiple lineages have undergone dramatic accelerations in rates of nucleotide substitution and extreme changes in mt genome structure. While molecular evolution in these lineages has been investigated, very little is known about their mt function. Here, we develop a new protocol to characterize respiration in isolated plant mitochondria and apply it to species of Silene with mt genomes that are rapidly evolving, highly fragmented, and exceptionally large ([~]11 Mbp). This protocol, complemented with traditional measures of plant fitness, cytochrome c oxidase activity assays, and fluorescence microscopy, was used to characterize inter-and intraspecific variation in mt function. Contributions of the individual \"classic\" OXPHOS complexes, the alternative oxidase, and external NADH dehydrogenases to overall mt respiratory flux were found to be similar to previously studied angiosperms with more typical mt genomes. Some differences in mt function could be explained by inter-and intraspecific variation, possibly due to local adaptation or environmental effects. Although this study suggests that these Silene species with peculiar mt genomes still show relatively normal mt function, future experiments utilizing the protocol developed here can explore such questions in a more detailed and comparative framework.

plant biology

Metabolome-scale genome-wide association studies reveal chemical diversity and genetic control of maize specialized metabolites

One Sentence SummaryHPLC-MS metabolite profiling of maize seedlings, in combination with genome-wide association studies, identifies numerous quantitative trait loci that influence the accumulation of foliar metabolites.\n\nAbstractCultivated maize (Zea mays) retains much of the genetic and metabolic diversity of its wild ancestors. Non-targeted HPLC-MS metabolomics using a diverse panel of 264 maize inbred lines identified a bimodal distribution in the prevalence of foliar metabolites. Although 15% of the detected mass features were present in >90% of the inbred lines, the majority were found in <50% of the samples. Whereas leaf bases and tips were differentiated primarily by flavonoid abundance, maize varieties (stiff-stalk, non-stiff-stalk, tropical, sweet corn, and popcorn) were differentiated predominantly by benzoxazinoid metabolites. Genome-wide association studies (GWAS), performed for 3,991 mass features from the leaf tips and leaf bases, showed that 90% have multiple significantly associated loci scattered across the genome. Several quantitative trait locus hotspots in the maize genome regulate the abundance of multiple, often metabolically related mass features. The utility of maize metabolite GWAS was demonstrated by confirming known benzoxazinoid biosynthesis genes, as well as by mapping isomeric variation in the accumulation of phenylpropanoid hydroxycitric acid esters to a single linkage block in a citrate synthase-like gene. Similar to gene expression databases, this metabolomic GWAS dataset constitutes an important public resource for linking maize metabolites with biosynthetic and regulatory genes.

genetics