Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

SMRT Genome Assembly Corrects Reference Errors, Resolving the Genetic Basis of Virulence in Mycobacterium tuberculosis

The genetic basis of virulence in Mycobacterium tuberculosis has been investigated through genome comparisons of its virulent (H37Rv) and attenuated (H37Ra) sister strains. Such analysis, however, relies heavily on the accuracy of the sequences. While the H37Rv reference genome has had several corrections to date, that of H37Ra is unmodified since its original publication. Here, we report the assembly and finishing of the H37Ra genome from single-molecule, real-time (SMRT) sequencing. Our assembly reveals that the number of H37Ra-specific variants is less than half of what the Sanger-based H37Ra reference sequence indicates, undermining and, in some cases, invalidating the conclusions of several studies. PE_PPE family genes, which are intractable to commonly-used sequencing platforms because of their repetitive and GC-rich nature, are overrepresented in the set of genes in which all reported H37Ra-specific variants are contradicted. We discuss how our results change the picture of virulence attenuation and the power of SMRT sequencing for producing high-quality reference genomes.

Genomics

The frequency spectrum of chromatin accessibility in Yorubans points toward a significant role of random genetic drift in shaping the chromatin landscape

The function of non-coding variation in the human genome is hotly debated. While much of the genome appears to be involved in some kind of molecular activity, a relatively small portion of the genome appears to be conserved across mammalian species. To try to understand part of this seeming paradox, we examined chromatin accessibility as a model molecular phenotype. We modeled chromatin state as either open or closed as looked at the frequency of open chromatin across 70 Yoruban cell lines. We saw that most regions of chromatin accessibility occurred in only a small number of individuals, although there are a number of regions that are accessible across the entire panel. To delve further into understanding the evolutionary mechanisms, we examined nucleotide diversity in and around accessible regions. We found that in the open chromatin access, low frequency regions had decreased nucleotide diversity, however, they were situated within regions of elevated nucleotide diversity. These results point toward a role of random mutation and genetic drift shaping the distribution of accessible regions in the human genome.

Evolutionary Biology

Evaluating the genetic diagnostic power of exome sequencing: Identifying missing data.

A hurdle of exome sequencing is its limited capacity to represent the entire exome. To ascertain the diagnostic power of this approach we determined the extent of coverage per individual sample. Using alignment data (BAM files) from 15 exome samples, sequences of any length that were below a determined sequencing depth coverage (DP) were detected and annotated with the Ensembl exon database using MIST, a novel software tool. Samples sequenced at 50X mean coverage had, on average, up to 50% of the Ensembl annotated exons with at least one nucleotide (L=1) with a DP<20, improving to 35% at 100X mean coverage. In addition, almost 15% of annotated exons were never sequenced (L=50, DP<1) at 50x mean coverage, reaching down to 5% at 100x. The diagnostic utility of this approach was tested for hypertrophic cardiomyopathy, a genetically heterogeneous disease, where exome sequencing covered as much as 80% of all candidate genes exons at DP[&ge;]20. This report stresses the value of identifying, precisely, which sequences are below a specific depth in an individuals exome, and provides a useful tool to assess the potential and pitfalls of exome sequencing in a diagnostic or gene discovery setting.\n\nCOMPETING INTERESTSThe authors declare no conflicts of interest.

Bioinformatics

Ohana, a tool set for population genetic analyses of admixture components

MotivationStructure methods are highly used population genetic methods for classifying individuals in a sample fractionally into discrete ancestry components.\n\nContributionWe introduce a new optimization algorithm of the classical Structure model in a maximum likelihood framework. Using analyses of real data we show that the new optimization algorithm finds higher likelihood values than the state-of-the-art method in the same computational time. We also present a new method for estimating population trees from ancestry components using a Gaussian approximation. Using coalescence simulations modeling populations evolving in a tree-like fashion, we explore the adequacy of the Structure model and the Gaussian assumption for identifying ancestry components correctly and for inferring the correct tree. In most cases, ancestry components are inferred correctly, although sample sizes and times since admixture can influence the inferences. Similarly, the popular Gaussian approximation tends to perform poorly when branch lengths are long, although the tree topology is correctly inferred in all scenarios explored. The new methods are implemented together with appropriate visualization tools in the computer package Ohana.\n\nAvailabilityOhana is publicly available at https://github.com/jade-cheng/ohana. Besides its source code and installation instructions, we also provide example workflows in the project wiki site.\n\nContactjade.cheng@birc.au.dk

Bioinformatics

Local genetic effects on gene expression across 44 human tissues

Expression quantitative trait locus (eQTL) mapping provides a powerful means to identify functional variants influencing gene expression and disease pathogenesis. We report the identification of cis-eQTLs from 7,051 post-mortem samples representing 44 tissues and 449 individuals as part of the Genotype-Tissue Expression (GTEx) project. We find a cis-eQTL for 88% of all annotated protein-coding genes, with one-third having multiple independent effects. We identify numerous tissue-specific cis-eQTLs, highlighting the unique functional impact of regulatory variation in diverse tissues. By integrating large-scale functional genomics data and state-of-the-art fine-mapping algorithms, we identify multiple features predictive of tissue-specific and shared regulatory effects. We improve estimates of cis-eQTL sharing and effect sizes using allele specific expression across tissues. Finally, we demonstrate the utility of this large compendium of cis-eQTLs for understanding the tissue-specific etiology of complex traits, including coronary artery disease. The GTEx project provides an exceptional resource that has improved our understanding of gene regulation across tissues and the role of regulatory variation in human genetic diseases.

Genomics

DNA Compass: a secure, client-side site for navigating personal genetic information

MotivationMillions of individuals have access to raw genomic data using direct-to-consumer companies. The advent of large-scale sequencing projects, such as the Precision Medicine Initiative, will further increase the number of individuals with access to their own genomic information. However, querying genomic data requires a computer terminal - an impediment for the general public.\n\nResultsDNA Compass is a website designed to empower the public by enabling simple navigation of personal genomic data. Users can query the status of their genomic variants for over 400 conditions or tens of millions of documented SNPs. DNA Compass presents the relevant genotypes of the user side-by-side with explanatory scientific resources. The genotypes data never leaves the users computer, a feature that provides improved security and performance. Nearly 2500 unique users have used our tool, mainly from the general genetic genealogy community, demonstrating its utility.\n\nAvailabilityDNA Compass is freely available on https://compass.dna.land.\n\nContactyaniv@cs.columbia.edu

Genomics

MR-Base: a platform for systematic causal inference across the phenome using billions of genetic associations

Published genetic associations can be used to infer causal relationships between phenotypes, bypassing the need for individual-level genotype or phenotype data. We have curated complete summary data from 1094 genome-wide association studies (GWAS) on diseases and other complex traits into a centralised database, and developed an analytical platform that uses these data to perform Mendelian randomization (MR) tests and sensitivity analyses (MR-Base, http://www.mrbase.org). Combined with curated data of published GWAS hits for phenomic measures, the MR-Base platform enables millions of potential causal relationships to be evaluated. We use the platform to predict the impact of lipid lowering on human health. While our analysis provides evidence that reducing LDL-cholesterol, lipoprotein(a) or triglyceride levels reduce coronary disease risk, it also suggests causal effects on a number of other non-vascular outcomes, indicating potential for adverse-effects or drug repositioning of lipid-lowering therapies.

epidemiology

Genetic Information Relationship Network (GIRN): A Force-directed graphing tool for gene expression analysis

We present a web-based tool, Genetic Information Relationship Network (GIRN), for mapping genes to a global protein interaction network for humans, and for further annotating this map with gene ontology, pathway, disease, and user-generated terms. These annotations can be adjusted according to enrichment within a subset of genes. Additionally, drugs that interact with genes in the subset can be added to the graph. The maps are force-directed graphs in which highly connected nodes tend toward the center, and highly interconnected nodes tend toward each other. Icons of different shapes and colors are employed to indicate whether the node represents a protein, a gene, or one of the various types of annotation. Coordinated interaction of genes and functional inference can be identified by visual inspection. Each node in the graph is associated with a menu of links to external data sources. Collectively, these tools provide an efficient portal to gene-associated public data for any group of genes specified by the user. The site can be found at www.voxvill.org/relnet.

Bioinformatics

The human functional genome defined by genetic diversity

Large scale efforts to sequence whole human genomes provide extensive data on the non-coding portion of the genome. We used variation information from 11,257 human genomes to describe the spectrum of sequence conservation in the population. We established the genome-wide variability for each nucleotide in the context of the surrounding sequence in order to identify departure from expectation at the population level (context-dependent conservation). We characterized the population diversity for functional elements in the genome and identified the coordination of conserved sequences of distal and cis enhancers, chromatin marks, promoters, coding and intronic regions. The most context-dependent conserved regions of the genome are associated with unique functional annotations and a genomic organization that spreads up to one megabase. Importantly, these regions are enriched by over 100-fold of non-coding pathogenic variants. This analysis of human genetic diversity thus provides a detailed view of sequence conservation, functional constraint and genomic organization of the human genome. Specifically, it identifies highly conserved non-coding sequences that are not captured by analysis of interspecies conservation and are greatly enriched in disease variants.

genomics

Level-Based Analysis of Genetic Algorithms and Other Search Processes

Understanding how the time-complexity of evolutionary algorithms (EAs) depend on their parameter settings and characteristics of fitness landscapes is a fundamental problem in evolutionary computation. Most rigorous results were derived using a handful of key analytic techniques, including drift analysis. However, since few of these techniques apply effortlessly to population-based EAs, most time-complexity results concern simplified EAs, such as the (1 + 1) EA.\n\nThis paper describes the level-based theorem, a new technique tailored to population-based processes. It applies to any non-elitist process where o spring are sampled independently from a distribution depending only on the current population. Given conditions on this distribution, our technique provides upper bounds on the expected time until the process reaches a target state.\n\nWe demonstrate the technique on several pseudo-Boolean functions, the sorting problem, and approximation of optimal solutions in combina-torial optimisation. The conditions of the theorem are often straightfor-ward to verify, even for Genetic Algorithms and Estimation of Distribution Algorithms which were considered highly non-trivial to analyse. Finally, we prove that the theorem is nearly optimal for the processes considered. Given the information the theorem requires about the process, a much tighter bound cannot be proved.

evolutionary biology

A positive association between population genetic differentiation and speciation rates in New World birds

Although an implicit assumption of speciation biology is that population differentiation is an important stage of evolutionary diversification, its true significance remains largely untested. If population differentiation within a species is related to its speciation rate over evolutionary time, the causes of differentiation could also be driving dynamics of organismal diversity across time and space. Alternatively, geographic variants might be short-lived entities with rates of formation that are unlinked to speciation rates, in which case the causes of differentiation would have only ephemeral impacts. Combining population genetics datasets including 17,746 individuals from 176 New World bird species with speciation rates estimated from phylogenetic data, we show that the population differentiation rates within species predict their speciation rates over long timescales. Although relatively little variance in speciation rate is explained by population differentiation rate, the relationship between the two is robust to diverse strategies of sampling and analyzing both population-level and species-level datasets. Population differentiation occurs at least three to five times faster than speciation, suggesting that most populations are ephemeral. Population differentiation and speciation rates are more tightly linked in tropical species than temperate species, consistent with a history of more stable diversification dynamics through time in the Tropics. Overall, our results suggest investigations into the processes responsible for population differentiation can reveal factors that contribute to broad-scale patterns of diversity.

evolutionary biology

The Impact of Lethal Recessive Alleles on Bottlenecks with Implications for Conservation Genetics

When a bottleneck occurs, lethal recessive alleles from the ancestral population provide a genetic load. The purging of lethal recessive mutations may prolong the bottleneck, or even cause the population to become extinct. But the purging is of short duration, it will be over before near neutral deleterious alleles accumulate. Lethal recessive alleles from the parental population and near neutral deleterious mutations which occur during a bottleneck are temporally separated threats to the survival of a population. Breeding individuals from a large population into a small endangered population will provide the benefit of viable alleles to replace near neutral deleterious alleles but also the cost of lethal recessive mutations from the large population.

evolutionary biology

An interaction map of circulating metabolites, immune gene networks and their genetic regulation

The interaction between metabolism and the immune system plays a central role in many cardiometabolic diseases. We integrated blood transcriptomic, metabolomic, and genomic profiles from two population-based cohorts, including a subset with 7-year follow-up sampling. We identified topologically robust gene networks enriched for diverse immune functions including cytotoxicity, viral response, B cell, platelet, neutrophil, and mast cell/basophil activity. These immune gene modules showed complex patterns of association with 158 circulating metabolites, including lipoprotein subclasses, lipids, fatty acids, amino acids, and CRP. Genome-wide scans for module expression quantitative trait loci (mQTLs) revealed five modules with mQTLs of both cis and trans effects. The strongest mQTL was in ARHGEF3 (rs1354034) and affected a module enriched for platelet function. Mast cell/basophil and neutrophil function modules maintained their metabolite associations during 7-year follow-up, while our strongest mQTL in ARHGEF3 also displayed clear temporal stability. This study provides a detailed map of natural variation at the blood immuno-metabolic interface and its genetic basis, and facilitates subsequent studies to explain inter-individual variation in cardiometabolic disease.

genomics

Cross-tissue integration of genetic and epigenetic data offers insight into autism spectrum disorder

Epigenetics is an emerging area of investigation for Autism Spectrum Disorder (ASD). Integration of epigenetic information with ASD genetic results may elucidate functional insights not possible via either source of information in isolation. We used concurrent genotype and DNA methylation (DNAm) data from cord blood and peripheral blood from preschool-aged children to identify SNPs associated with DNA methylation, or methylation quantitative trait loci (meQTLs), and combined this with publicly available fetal brain and lung meQTL lists to assess enrichment of ASD GWAS results for tissue-specific meQTLs. ASD-associated SNPs were enriched for fetal brain (OR = 3.55; p < 0.001) and peripheral blood meQTLs (OR = 1.58; p < 0.001). The CpG site targets of ASD meQTLs across cord, blood, and brain tissues were enriched for immune-related pathways, consistent with other expression and DNAm results in ASD, and revealing pathways not implicated by genes identified from ASD rare variant work. Further, DNaseI hypersensitive sites and the STAT1 and TAF1 transcription factor binding sites were enriched for meQTL target CpGs of SNPs associated with psychiatric conditions. This joint analysis of genotype and DNAm demonstrates the potential utility of both brain and blood-based DNAm for insights into ASD and psychiatric phenotypes more broadly.

genomics

Review: Population Structure in Genetic Studies: Confounding Factors and Mixed Models

A genome-wide association study (GWAS) seeks to identify genetic variants that contribute to the development and progression of a specific disease. Over the past 10 years, new approaches using mixed models have emerged to mitigate the deleterious effects of population structure and relatedness in association studies. However, developing GWAS techniques to effectively test for association while correcting for population structure is a computational and statistical challenge. Using laboratory mouse strains as an example, our review characterizes the problem of population structure in association studies and describes how it can cause false positive associations. We then motivate mixed models in the context of unmodeled factors.

bioinformatics

Defining the genetic architecture of stripe rust resistance in the barley accession HOR1428

Puccinia striiformis f. sp. hordei, the causal agent of barley stripe rust, is a destructive fungal pathogen that significantly affects barley cultivation. A major constraint in breeding resistant cultivars is the lack of mapping information of resistance (R) genes and their introgression into adapted germplasm. A considerable number of R genes have been described in barley to P. striiformis f. sp. hordei, but only a few loci have been mapped. Previously, Chen and Line (1999) reported two recessive seedling resistance loci in the Ethiopian landrace HOR 1428. In this study, we map two loci that confer resistance to P. striiformis f. sp. hordei in HOR 1428, which are located on chromosomes 3H and 5H. Both loci act as additive effect QTLs, each explaining approximately 20% of the phenotypic variation. We backcrossed HOR 1428 to the cv. Manchuria and selected based on markers flanking the RpsHOR128-5H locus. Saturation of the RpsHOR1428-5H locus with markers in the region found KASP marker K_1_0292 in complete coupling with resistance to P. striiformis f. sp. hordei and was designated Rps9. Isolation of Rps9 and flanking markers will facilitate the deployment of this genetic resource into existing programs for P. striiformis f. sp. hordei resistance.

plant biology

FerriTag: A Genetically-Encoded Inducible Tag for Correlative Light-Electron Microscopy

A current challenge is to develop tags to precisely visualize proteins in cells by light and electron microscopy. Here, we introduce FerriTag, a genetically-encoded chemically-inducible tag for correlative light-electron microscopy (CLEM). FerriTag is a fluorescent recombinant electron-dense ferritin particle that can be attached to a protein-of-interest using rapamycin-induced heterodimerization. We demonstrate the utility of FerriTag for CLEM by labeling proteins associated with various intracellular structures including mitochondria, plasma membrane, and clathrin-coated pits and vesicles. FerriTagging has a high signal-to-noise ratio and a labeling resolution of 10 {+/-} 5 nm. We demonstrate how FerriTagging allows nanoscale mapping of protein location relative to a subcellular structure, and use it to detail the distribution of huntingtin-interacting protein 1 related (HIP1R) in clathrin-coated pits.

cell biology

A genetic basis for molecular asymmetry at vertebrate electrical synapses

Neural network function is based upon the patterns and types of connections made between neurons. Neuronal synapses are adhesions specialized for communication and they come in two types, chemical and electrical. Communication at chemical synapses occurs via neurotransmitter release whereas electrical synapses utilize gap junctions for direct ionic and metabolic coupling. Electrical synapses are often viewed as symmetrical structures, with the same components making both sides of the gap junction. By contrast, we show that a broad set of electrical synapses in zebrafish, Danio rerio, require two gap-junction-forming Connexins for formation and function. We find that one Connexin functions presynaptically while the other functions postsynaptically in forming the channels. We also show that these synapses are required for the speed and coordination of escape responses. Our data identify a genetic basis for molecular asymmetry at vertebrate electrical synapses and show they are required for appropriate behavioral performance.

neuroscience