Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

G-quadruplex forming sequences in the genome of all known human viruses: a comprehensive guide

G-quadruplexes are non-canonical nucleic acid structures that control transcription, replication, and recombination in organisms. G-quadruplexes are present in eukaryotes, prokaryotes, and viruses. In the latter, mounting evidence indicates their key biological activity. Since data on viruses are scattered, we here present a comprehensive analysis of putative G-quadruplexes in the genome of all known viruses that can infect humans. We show that the presence, distribution, and location of G-quadruplexes are features characteristic of each virus class and family. Our statistical analysis proves that their presence within the viral genome is orderly arranged, as indicated by the possibility to correctly assign up to two-thirds of viruses to their exact class based on the G-quadruplex classification. For each virus we provide: i) the list of all G-quadruplexes formed by GG-, GGG- and GGGG-islands present in the genome (positive and negative strands), ii) their position in the viral genome along with the known function of that region, iii) the degree of conservation among strains of each G-quadruplex in its genome context, iv) the statistical significance of G-quadruplex formation. This information is accessible from a database (http://www.medcomp.medicina.unipd.it/main_site/doku.php?id=g4virus) to allow the easy and interactive navigation of the results. The availability of these data will greatly expedite research on G-quadruplex in viruses, with the possibility to accelerate finding therapeutic opportunities to numerous and some fearsome human diseases.

bioinformatics

A histone H4K20 methylation-mediated chromatin compaction threshold ensures genome integrity by limiting DNA replication licensing

The decompaction and re-establishment of chromatin organization immediately after mitosis is essential for genome regulation. The mechanisms underlying chromatin structure control in daughter cells are not fully understood. Here, we show that a chromatin compaction threshold in cells exiting mitosis ensures genome integrity by limiting replication licensing in G1 phase. Upon mitotic exit, appropriate chromatin relaxation is safeguarded by SET8-dependent methylation of histone H4 on lysine 20. Thus, in the absence of either SET8 or the H4K20 residue, substantial genome-wide chromatin decompaction occurs which allows excessive loading of the Origin Recognition Complex (ORC) in the daughter cells. ORC overloading stimulates aberrant recruitment of the MCM2-7 complex that promotes single-stranded DNA formation and DNA damage. Restoring chromatin compaction restrains excess replication licensing and the loss of genome integrity. Our findings identify a cell cycle-specific mechanism whereby fine-tuned chromatin relaxation suppresses excessive detrimental replication licensing and maintains genome integrity at the cellular transition from mitosis to G1 phase.

molecular biology

The genome and metabolome of the tobacco tree, Nicotiana glauca: a potential renewable feedstock for the bioeconomy

BackgroundGiven its tolerance to stress and its richness in particular secondary metabolites, the tobacco tree, Nicotiana glauca, has been considered a promising biorefinery feedstock that would not be competitive with food and fodder crops.\n\nResultsHere we present a 3.5 Gbp draft sequence and annotation of the genome of N. glauca spanning 731,465 scaffold sequences, with an N50 size of approximately 92 kbases. Furthermore, we supply a comprehensive transcriptome and metabolome analysis of leaf development comprising multiple techniques and platforms.\n\nThe genome sequence is predicted to cover nearly 80% of the estimated total genome size of N. glauca. With 73,799 genes predicted and a BUSCO score of 94.9%, we have assembled the majority of gene-rich regions successfully. RNA-Seq data revealed stage-and/or tissue-specific expression of genes, and we determined a general trend of a decrease of tricarboxylic acid cycle metabolites and an increase of terpenoids as well as some of their corresponding transcripts during leaf development.\n\nConclusionThe N. glauca draft genome and its detailed transcriptome, together with paired metabolite data, constitute a resource for future studies of valuable compound analysis in tobacco species and present the first steps towards a further resolution of phylogenetic, whole genome studies in tobacco.

plant biology

Streamlined, recombinase-free genome editing with CRISPR-Cas9 in Lactobacillus plantarum reveals barriers to efficient editing

Lactic-acid bacteria such as Lactobacillus plantarum are commonly used for fermenting foods and as probiotics, where increasingly sophisticated genome-editing tools are currently being employed to elucidate and enhance these microbes beneficial properties. The most advanced tools to-date require heterologous single-stranded DNA recombinases to integrate short oligonucleotides followed by using CRISPR-Cas9 to eliminate cells harboring unedited sequences. Here, we show that encoding the recombineering template on a replicating plasmid allowed efficient genome editing with CRISPR-Cas9 in multiple L. plantarum strains without a recombinase. This strategy accelerated the genome-editing pipeline and could efficiently introduce a stop codon in ribB, silent mutations in ackA, and a complete deletion of lacM. In contrast, oligo-mediated recombineering with CRISPR-Cas9 proved far less efficient in at least one instance. We also observed unexpected outcomes of our recombinase-free method, including an ~1.3-kb genomic deletion when targeting ribB in one strain, and reversion of a point mutation in the recombineering template in another strain. Our method therefore can streamline targeted genome editing in different strains of L. plantarum, although the best means of achieving efficient editing may vary based on the selected sequence modification, gene, and strain.

synthetic biology

Simulation-based approaches to characterize the effect of sequencing depth on the quantity and quality of metagenome-assembled genomes

We applied simulation-based approaches to characterize how microbial community structure influences the amount of sequencing effort to reconstruct metagenomes that are assembled from short read sequences. An initial analysis evaluated the quantity, completion, and contamination of complete-metagenome-assembled genome (complete-MAG) equivalents, a bioinformatic-pipeline normalized metric for MAG quantity, as a function of sequencing effort, on four preexisting sequence read datasets taken from a maize soil, an estuarine sediment, the surface ocean, and the human gut. These datasets were subsampled to varying degrees of completeness in order to simulate the effect of sequencing effort on MAG retrieval. Modeling suggested that sequencing efforts beyond what is typical in published experiments (1 to 10 Gbp) would generate diminishing returns in terms of MAG binning. A second analysis explored the theoretical relationship between sequencing effort and the proportion of available metagenomic DNA sequenced during a sequencing experiment as a function of community richness, evenness, and genome size. Simulations from this analysis demonstrated that while community richness and evenness influenced the amount of sequencing required to sequence a community metagenome to exhaustion, the effort necessary to sequence an individual genome to a target fraction of exhaustion was only dependent on the relative abundance of the corresponding organism and its genome size. A software tool, GRASE, was created to assist investigators further explore this relationship. Re-evaluation of the relationship between sequencing effort and binning success in the context of the relative abundance of genomes, as opposed to base pairs, provides a framework to design sequencing experiments based on the relative abundance of microbes in an environment rather than arbitrary levels of sequencing effort.

bioinformatics

Whole genome scan reveals the multigenic basis of recent tidal marsh adaptation in a sparrow

Natural selection acts on functional molecular variation to create local adaptation, the \"good fit\" we observe between an organisms phenotype and its environment. Genomic comparisons of lineages in the earliest stages of adaptive divergence have high power to reveal genes under natural selection because molecular signatures of selection on functional loci are maximally detectable when overall genomic divergence is low. We conducted a scan for local adaptation genes in the North American swamp sparrow (Melospiza georgiana), a species that includes geographically connected populations that are differentially adapted to freshwater vs. brackish tidal marshes. The brackish tidal marsh form has rapidly evolved tolerance for salinity, a deeper bill, and darker plumage since colonizing coastal habitats within the last 15,000 years. Despite their phenotypic differences, background genomic divergence between these populations is very low, rendering signatures of natural selection associated with this recent coastal adaptation highly detectable. We recovered a multigenic snapshot of ecological selection via a whole genome scan that revealed robust signatures of selection at 31 genes with functional connections to bill shape, plumage melanism and salt tolerance. As in Darwins finches, BMP signaling appears responsible for changes in bill depth, a putative magic trait for ecological speciation. A signal of selection at BNC2, a melanocyte transcription factor responsible for human skin color saturation, implicates a shared genetic mechanism for sparrow plumage color and human skin tone. Genes for salinity tolerance constituted the majority of adaptive candidates identified in this genome scan (23/31) and included vasoconstriction hormones that can flexibly modify osmotic balance in tune with the tidal cycle by influencing both drinking behavior and kidney physiology. Other salt tolerance genes had potential pleiotropic effects on bill depth and melanism (6/31), offering a mechanistic explanation for why these traits have evolved together in coastal swamp sparrows, and in other organisms that have converged on the same \"salt marsh syndrome\". As a set, these candidates capture the suite of physiological changes that coastal swamp sparrows have evolved in response to selection pressures exerted by a novel and challenging habitat.

evolutionary biology

Mango: Distributed Visualization for Genomic Analysis

The decreasing cost of DNA sequencing over the past decade has led to an explosion of available sequencing datasets, leaving us with terabytes to petabytes of data to explore and analyze. It is critical for analysts in research and clinical settings to be able to develop new data-driven hypotheses from these datasets through bias identification, analysis of data quality, and testing different algorithms and parameter settings. However, current interactive tools for sequence analysis are designed to run on single machines that do not scale to the size of modern genomic datasets, and rely on precomputed static views, rather than allowing direct interaction with the primary dataset. Mango is a genomic sequence visualization and analysis platform that removes these constraints regarding scalability and staticity by leveraging the power of multi-node compute clusters in the cloud to allow interactive analysis over terabytes of sequencing data. Mango provides both a genome browser graphical user interface and programmable notebook form factor to allow users of varying analytical experience to explore large sequencing datasets on both private clusters and in the cloud. These tools provide a flexible environment for interactive exploration of genomic datasets, while surpassing the computational limits of single-node genomic visualization tools.

bioinformatics

Evidence for a Recombinant Origin of HIV-1 group M from Genomic Variation

Reconstructing the early dynamics of the HIV-1 pandemic can provide crucial insights into the socioeconomic drivers of emerging infectious diseases in human populations, including the roles of urbanization and transportation networks. Current evidence indicates that the global pandemic comprising almost entirely of HIV-1/M originated around the 1920s in central Africa. However, these estimates are based on molecular clock estimates that are assumed to apply uniformly across the virus genome. There is growing evidence that recombination has played a significant role in the early history of the HIV-1 pandemic, such that different regions of the HIV-1 genome have different evolutionary histories. In this study, we have conducted a dated-tip analysis of all near full-length HIV-1/M genome sequences that were published in the GenBank database. We used a sliding window approach similar to the bootscanning method for detecting breakpoints in intersubtype recombinant sequences. We found evidence of substantial variation in estimated root dates among windows, with an estimated mean time to the most recent common ancestor (tMRCA) of 1922. Estimates were significantly autocorrelated, which was more consistent with an early recombination event than with stochastic error variation in phylogenetic reconstruction and dating analyses. A piecewise regression analysis supported the existence of at least one recombination breakpoint in the HIV-1/M genome with interval-specific means around 1929 and 1913, respectively. This analysis demonstrates that a sliding window approach can accommodate early recombination events outside the established nomenclature of HIV-1/M subtypes, although it is difficult to incorporate the earliest available samples due to their limited genome coverage.

evolutionary biology

Dynamics of genomic change during evolutionary rescue in the seed beetle Callosobruchus maculatus

Rapid adaptation can be necessary to prevent extinction when populations are exposed to extremely marginal or stressful environments. Factors that affect the likelihood of evolutionary rescue from extinction have been identified, but much less is known about the evolutionary dynamics and genomic basis of successful evolutionary rescue, particularly in multicellular organisms. We conducted an evolve and resequence experiment to investigate the dynamics and repeatability of evolutionary rescue at the genetic level in the cowpea seed beetle, Callosobruchus maculatus, when it is experimentally shifted to a stressful host plant, lentil (Lens culinaris). Low survival (~ 1%) at the onset of the experiment caused population decline. But adaptive evolution quickly rescued the population with survival rates climbing to 69% by the F5 generation and 90% by the F10 generation. Population genomic data showed that rescue likely was caused by rapid evolutionary change at multiple loci, with many alleles fixing or nearly fixing within five generations of selection on lentil. By comparing estimates of selection across five lentil-adapted C. maculatus populations (two new sublines and three long-established lines), we found that adaptation to lentil involves a mixture of parallel and idiosyncratic evolutionary changes. Parallelism was particularly pronounced in sublines that were formed after the parent line had passed through an initial bottleneck. Overall, our results suggest that evolutionary rescue in this system is driven by very strong selection on a modest number of loci, and these results provide empirical evidence that ecological dynamics during evolutionary rescue cause distinct evolutionary trajectories and genomic signatures relative to adaptation in less stressful environments.\n\nImpact StatementEvolutionary adaptation is an ongoing process in most populations, but when populations occupy particularly stressful or marginal environments, adaptation can be necessary to prevent extinction. Adaptation that reverses demographic decline and allows for population persistence is termed evolutionary rescue. Evolutionary rescue can prevent species loss from climate change or other environmental stresses, but it can also thwart attempts to control or eradicate agricultural pests and pathogens. Many factors affect the likelihood of evolutionary rescue, but little is known about the underlying evolutionary dynamics, particularly molecular evolutionary changes in multicellular organisms. Here we use a powerful combination of experimental evolution and genomics to track the evolutionary dynamics and genomic outcomes of evolutionary rescue. We focus on the seed beetle Callosobruchus maculatus, which is both an agricultural pest and a convenient model system. We specifically examine how this species is able to persist on a novel and very poor crop host, lentil.\n\nWe show that evolution in an experimental seed beetle populations increases survival on lentil from ~1% to >80% in fewer than a dozen generations. This rapid adaptive evolutionary change at the trait (i.e., phenotypic) level was associated with equally rapid evolution at the molecular level, with some gene variants (i.e., alleles) showing frequency shifts of around 30% in a single generation. In contrast to most other experimental evolution studies in multicellular organisms (particularly Drosophila fruit flies), we find that gene variants at multiple loci rapidly fix, that is, reach a frequency of 100%, during adaptation to lentil. Our results suggest that the dynamics and genetics of adaptation to severe conditions could be distinct from adaptation under more benign conditions. By comparing outcomes of adaptation across multiple lines and sublines, we show that repeated rapid adaptation at the trait level does not necessarily involve the same evolutionary changes at the molecular level. This limited parallelism was likely driven by extreme population bottlenecks caused by low survival in the early generations on lentil. Indeed, evolutionary changes in sublines formed after recovery from a common bottleneck were highly parallel. This coupling of demographic (i.e., ecological) and evolutionary changes during evolutionary rescue may therefore limit the predictability of evolution. Because colonization of novel environments may often occur after a bottleneck, our results could be of general significance for understanding patterns of parallel (and non-parallel) evolutionary change in nature.

evolutionary biology

Genome-scale metabolic construction of the stress-tolerant hybrid yeast Zygosaccharomyces parabailii

Genome-scale metabolic models are powerful tools to understand and engineer cellular systems facilitating their use as cell factories. This is especially true for microorganisms with known genome sequences from which nearly complete sets of enzymes and metabolic pathways are determined, or can be inferred. Yeasts are highly diverse eukaryotes whose metabolic traits have long been exploited in industry, and although many of their genome sequences are available, few genome-scale metabolic models have so far been produced. For the first time, we reconstructed the genome-scale metabolic model of the hybrid yeast Zygosaccharomyces parabailii, which is a member of the Z. bailii sensu lato clade notorious for stress-tolerance and therefore relevant to industry. The model comprises 3096 reactions, 2091 metabolites, and 2413 genes. Our own laboratory data were then used to establish a biomass synthesis reaction, and constrain the extracellular environment. Through constraint-based modeling, our model reproduces the co-consumption and catabolism of acetate and glucose posing it as a promising platform for understanding and exploiting the metabolic potential of Z. parabailii.

systems biology

Genome-scale sequence disruption following biolistic transformation in rice and maize

We biolistically transformed linear 48 kb phage lambda and two different circular plasmids into rice and maize and analyzed the results by whole genome sequencing and optical mapping. While some transgenic events showed simple insertions, others showed extreme genome damage in the form of chromosome truncations, large deletions, partial trisomy, and evidence of chromothripsis and breakage-fusion bridge cycling. Several transgenic events contained megabase-scale arrays of introduced DNA mixed with genomic fragments assembled by non-homologous or microhomology-mediated joining. Damaged regions of the genome, assayed by the presence of small fragments displaced elsewhere, were often repaired without a trace, presumably by homology-dependent repair (HDR). The results suggest a model whereby successful biolistic transformation relies on a combination of end joining to insert foreign DNA and HDR to repair collateral damage caused by the microprojectiles. The differing levels of genome damage observed among transgenic events may reflect the stage of the cell cycle and the availability of templates for HDR.

plant biology

Cross-species alcohol dependence-associated gene networks: Co-analysis of mouse brain gene expression and human genome-wide association data

Genome-wide association studies on alcohol dependence, by themselves, have yet to account for the estimated heritability of the disorder and provide incomplete mechanistic understanding of this complex trait. Integrating brain ethanol-responsive gene expression networks from model organisms with human genetic data on alcohol dependence could aid in identifying dependence-associated genes and functional networks in which they are involved. This study used a modification of the Edge-Weighted Dense Module Searching for genome-wide association studies (EW-dmGWAS) approach to co-analyze whole-genome gene expression data from ethanol-exposed mouse brain tissue, human protein-protein interaction databases and alcohol dependence-related genome-wide association studies. Results revealed novel ethanol-regulated and alcohol dependence-associated gene networks in prefrontal cortex, nucleus accumbens, and ventral tegmental area. Three of these networks were overrepresented with genome-wide association signals from an independent dataset. These networks were significantly overrepresented for gene ontology categories involving several mechanisms, including actin filament-based activity, transcript regulation, Wnt and Syndecan-mediated signaling, and ubiquitination. Together, these studies provide novel insight for brain mechanisms contributing to alcohol dependence.

genetics

Diverse endogenous retroviruses generate structural variation between human genomes via LTR recombination

Human endogenous retroviruses (HERVs) occupy a substantial fraction of the genome and impact cellular function with both beneficial and deleterious consequences. The vast majority of HERV sequences descend from ancient retroviral families no longer capable of infection or genomic propagation. In fact, most are no longer represented by full-length proviruses but by solitary long terminal repeats (solo LTRs) that arose via non-allelic recombination events between the two LTRs of a proviral insertion. Because LTR-LTR recombination events may occur long after proviral insertion but are challenging to detect in resequencing data, we hypothesize that this mechanism produces an underappreciated amount of genomic variation in the human population. To test this idea, we develop a computational pipeline specifically designed to capture such dimorphic HERV alleles from short-read genome sequencing data. When applied to 279 individuals sequenced as part of the Simons Genome Diversity Project, the pipeline retrieves most of the dimorphic variants previously reported for the HERV-K(HML2) subfamily as well as dozens of additional candidates, including members of the HERV-H and HERV-W families. We experimentally validate several of these candidates, including the first reported instance of an unfixed HERV-W provirus. These data indicate that human proviral content exhibit more extensive interindividual variation than previously recognized. These findings have important implications for our understanding of the contribution of HERVs to human physiology and disease.

genetics

Allele-specific genome editing using CRISPR-Cas9 causes off-target mutations in diploid yeast

Targeted DNA double-strand breaks (DSBs) with CRISPR-Cas9 have revolutionized genetic modification by enabling efficient genome editing in a broad range of eukaryotic systems. Accurate gene editing is possible with near-perfect efficiency in haploid or (predominantly) homozygous genomes. However, genomes exhibiting polyploidy and/or high degrees of heterozygosity are less amenable to genetic modification. Here, we report an up to 99-fold lower gene editing efficiency when editing individual heterozygous loci in the yeast genome. Moreover, Cas9-mediated introduction of a DSB resulted in large scale loss of heterozygosity affecting DNA regions up to 360 kb that resulted in introduction of nearly 1700 off-target mutations, due to replacement of sequences on the targeted chromosome by corresponding sequences from its non-targeted homolog. The observed patterns of loss of heterozygosity were consistent with homology directed repair. The extent and frequency of loss of heterozygosity represent a novel mutagenic side-effect of Cas9-mediated genome editing, which would have to be taken into account in eukaryotic gene editing. In addition to contributing to the limited genetic amenability of heterozygous yeasts, Cas9-mediated loss of heterozygosity could be particularly deleterious for human gene therapy, as loss of heterozygous functional copies of anti-proliferative and pro-apoptotic genes is a known path to cancer.

molecular biology

A versatile platform strain for high-fidelity multiplex genome editing

Precision genome editing accelerates the discovery of the genetic determinants of phenotype and the engineering of novel behaviors in organisms. Advances in DNA synthesis and recombineering have enabled high-throughput engineering of genetic circuits and biosynthetic pathways via directed mutagenesis of bacterial chromosomes. However, the highest recombination efficiencies have to date been reported in persistent mutator strains, which suffer from reduced genomic fidelity. The absence of inducible transcriptional regulators in these strains also prevents concurrent control of genome engineering tools and engineered functions. Here, we introduce a new recombineering platform strain, BioDesignER, which incorporates (1) a refactored {lambda}-Red recombination system that reduces toxicity and accelerates multi-cycle recombination, (2) genetic modifications that boost recombination efficiency, and (3) four independent inducible regulators to control engineered functions. These modifications resulted in single-cycle recombineering efficiencies of up to 25% with a seven-fold increase in recombineering fidelity compared to the widely used recombineering strain EcNR2. To facilitate genome engineering in BioDesignER, we have curated eight context-neutral genomic loci, termed Safe Sites, for stable gene expression and consistent recombination efficiency. BioDesignER is a platform to develop and optimize engineered cellular functions and can serve as a model to implement comparable recombination and regulatory systems in other bacteria.

synthetic biology

Evolution of Replication Origins in Vertebrate Genomes: Rapid Turnover Despite Selective Constraints.

BackgroundThe replication programme of vertebrate genomes is driven by the chro-mosomal distribution and timing of activation of tens of thousands of replication origins. Genome-wide studies have shown the frequent association of origins with promoters and CpG islands, and their enrichment in G-quadruplex sequence motifs (G4). However, the genetic determinants driving their activity remain poorly understood. To gain insight on the functional constraints operating on replication origins and their spatial distribution, we conducted the first evolutionary comparison of genome-wide origins maps across vertebrates.\n\nResultsWe generated a high resolution genome-wide map of chicken replication origins (the first of a bird genome), and performed an extensive comparison with human and mouse maps. The analysis of intra-species polymorphism revealed a strong depletion of genetic diversity on an ~ 40 bp region centred on the replication initiation loci. Surprisingly, this depletion in genetic diversity was not linked to the presence of G4 motifs, nor to the association with promoters or CpG islands. In contrast, we also showed that origins experienced a rapid turnover during vertebrates evolution, since pairwise comparisons of origin maps revealed that only 4 to 24% of them were conserved between any two species.\n\nConclusionsThis study unravels the existence of a novel genetic determinant of replication origins, the precise functional role of which remains to be determined. Despite the importance of replication initiation activity for the fitness of organisms, the distribution of replication origins along vertebrate chromosomes is highly flexible.

evolutionary biology

Pan-genome-scale network reconstruction: a framework to increase the quantity and quality of metabolic network reconstructions throughout the tree of life

A genome-scale network reconstruction (GENRE) represents the knowledgebase of an organism and can be used in a variety of applications. The drop in genome sequencing costs has led to an increase in sequenced genomes, but the number of curated GENRE s has not kept pace. This gap hinders our ability to study physiology across the tree of life. Furthermore, our analysis of yeast GENRE s has found they contain significant commission and omission errors, especially in central metabolism. To address these quantity and quality issues for GENRE s, we propose open and transparent curation of the pan-genome, pan-reactome, pan-metabolome, and pan-phenome for taxons by research communities, rather than for a single species. We outline our approach with a Fungi pan-GENRE by integrating AYbRAH, our ortholog database, and AYbRAHAM, our new fungal reaction database. This pan-GENRE was used to compile 33 yeast/fungi GENRE s in the Dikarya subkingdom, spanning 600 million years. The fungal pan-GENRE contains 1547 orthologs, 2726 reactions, 2226 metabolites, and 10 compartments. The strain GENRE s have a wider genomic and metabolic than previous yeast and fungi GENRE s. Metabolic simulations show the amino acid yields from glucose differs between yeast lineages, indicating metabolic networks have evolved in yeasts. Curating ortholog and reaction databases for a taxon can be used to increase the quantity and quality of strain GENRE s. This pan-GENRE framework provides the ability to scale high-quality GENRE s to more branches in the tree of life.

bioinformatics

Genome wide association with quantitative resistance phenotypes in Mycobacterium tuberculosis reveals novel resistance genes and regulatory regions

Drug resistance is threatening attempts at tuberculosis epidemic control. Molecular diagnostics for drug resistance that rely on the detection of resistance-related mutations could expedite patient care and accelerate progress in TB eradication. We performed minimum inhibitory concentration testing for 12 anti-TB drugs together with Illumina whole genome sequencing on 1452 clinical Mycobacterium tuberculosis (MTB) isolates. We then used a linear mixed model to evaluate genome wide associations between mutations in MTB genes or noncoding regions and drug resistance, followed by validation of our findings in an independent dataset of 792 patient isolates. Novel associations at 13 genomic loci were confirmed in the validation set, with 2 involving noncoding regions. We found promoter mutations to have smaller average effects on resistance levels than gene body mutations in genes where both can contribute to resistance. Enabled by a quantitative measure of resistance, we estimated the heritability of the resistance phenotype to 11 anti-TB drugs and identify a lower than expected contribution from known resistance genes. We also report the proportion of variation in resistance levels explained by the novel loci identified here. This study highlights the complexity of the genomic mechanisms associated with the MTB resistance phenotype, including the relatively large number of potentially causative or compensatory loci, and emphasizes the contribution of the noncoding portion of the genome.

evolutionary biology