Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Layering genetic circuits to build a single cell, bacterial half adder

Gene regulation in biological systems is impacted by the cellular and genetic context-dependent effects of the biological parts which comprise the circuit. Here, we have sought to elucidate the limitations of engineering biology from an architectural point of view, with the aim of compiling a set of engineering solutions for overcoming failure modes during the development of complex, synthetic genetic circuits. Using a synthetic biology approach that is supported by computational modelling and rigorous characterisation, AND, OR and NOT biological logic gates were layered in both parallel and serial arrangements to generate a repertoire of Boolean operations that include NIMPLY, XOR, half adder and half subtractor logics in single cell. Subsequent evaluation of these near-digital biological systems revealed critical design pitfalls that triggered genetic context dependent effects, including 5 UTR interference and uncontrolled switch-on behaviour of {sigma}54 promoter. Importantly, this work provides a representative case study to the debugging of genetic context dependent effects through principles elucidated herein, thereby providing a rational design framework to program single prokaryotic cell with diversified digital operations.

Synthetic Biology

A resurrection experiment finds evidence of both reduced genetic diversity and potential adaptive evolution in the agricultural weed Ipomoea purpurea

Despite the negative economic and ecological impact of weeds, relatively little is known about the evolutionary mechanisms that influence their persistence in agricultural fields. Here, we use a resurrection ecology approach and compare the genetic and phenotypic divergence of temporally sampled seed progenies of Ipomoeapurpurea, an agricultural weed that is resistant to glyphosate, the most widely used herbicide in current-day agriculture. We found striking reductions in allelic diversity between cohorts sampled nine years apart (2003 vs 2012), suggesting that populations of this species sampled from agricultural fields have experienced genetic bottleneck events that have led to lower neutral genetic diversity. Heterozygosity excess tests indicate that this bottleneck may have occurred prior to 2003. Further, a greenhouse assay of individuals sampled from the field as seed found that populations of this species, on average, exhibited modest increases in herbicide resistance over time. Our results show that populations of this noxious weed, capable of adapting to strong selection imparted by herbicide application, may lose genetic variation as a result of this or other environmental factors. We likely uncovered only modest increases in resistance between sampling cohorts due to a strong and previously identified fitness cost of resistance in this species, along with the potential that non-resistant migrants germinate from the seed bank.

Evolutionary Biology

Whole genome sequencing of field isolates reveals extensive genetic diversity in Plasmodium vivax from Colombia

Plasmodium vivax is the most prevalent malarial species in South America and exerts a substantial burden on the populations it affects. The control and eventual elimination of P. vivax are global health priorities. Genomic research contributes to this objective by improving our understanding of the biology of P. vivax and through the development of new genetic markers that can be used to monitor efforts to reduce malaria transmission.\n\nHere we analyze whole-genome data from eight field samples from a region in Cordoba, Colombia where malaria is endemic. We find considerable genetic diversity within this population, a result that contrasts with earlier studies suggesting that P. vivax had limited diversity in the Americas. We also identify a selective sweep around a substitution known to confer resistance to sulphadoxine-pyrimethamine (SP). This is the first observation of a selective sweep for SP resistance in this species. These results indicate that P. vivax has been exposed to SP pressure even when the drug is not in use as a first line treatment for patients afflicted by this parasite. We identify multiple non-synonymous substitutions in three other genes known to be involved with drug resistance in Plasmodium species. Finally, we found extensive microsatellite polymorphisms. Using this information we developed 18 polymorphic and easy to score microsatellite loci that can be used in epidemiological investigations in South America.\n\nAuthor SummaryAlthough P. vivax is not as deadly as the more widely studied P. falciparum, it remains a pressing global health problem. Here we report the results of a whole-genome study of P. vivax from Cordoba, Colombia, in South America. This parasite is the most prevalent in this region. We show that the parasite population is genetically diverse, which is contrary to expectations from earlier studies from the Americas. We also find molecular evidence that resistance to an anti-malarial drug has arisen recently in this region. This selective sweep indicates that the parasite has been exposed to a drug that is not used as first-line treatment for this malaria parasite. In addition to extensive single nucleotide and microsatellite polymorphism, we report 18 new genetic loci that might be helpful for fine-scale studies of this species in the Americas.

Genomics

Further genetic diversification in multiple tumors and an evolutionary perspective on therapeutics

AbstractsThe genetic diversity within a single tumor can be extremely large, possibly with mutations at all coding sites (Ling et al. 2015). In this study, we analyzed 12 cases of multiple hepatocellular carcinoma (HCC) tumors by sequencing and genotyping several samples from each case. In 10 cases, tumors are clonally related by a process of cell migration and colonization. They permit a detailed analysis of the evolutionary forces (mutation, migration, drift and natural selection) that influence the genetic diversity both within and between tumors. In 23 inter-tumor comparisons, the descendant tumor usually shows a higher growth rate than the parent tumor. In contrast, neutral diversity dominates within-tumor observations such that adaptively growing clones are rarely found. The apparent adaptive evolution between tumors can be explained by the inherent bias for detecting larger tumors that have a growth advantage. Beyond these tumors are a far larger number of clones which, growing at a neutral rate and too small to see, can nevertheless be verified by molecular means. Given that the estimated genetic diversity is often very large, therapeutic strategies need to take into account the pre-existence of many drug-resistance mutations. Importantly, these mutations are expected to be in the very low frequency range in the primary tumors (and become frequent in the relapses, as is indeed reported (1-3). In conclusion, tumors may often harbor a very large number of mutations in the very low frequency range. This duality provides both a challenge and an opportunity for designing strategies against drug resistance (4-8).\n\nOne Sentence SummaryThe total genetic diversity across all tumors of a single patient, with large number of low frequency mutations driven by neutral and adaptive forces, presents both a challenge and an opportunity for new cancer therapeutics.

Cancer Biology

General methods for evolutionary quantitative genetic inference from generalised mixed models.

Methods for inference and interpretation of evolutionary quantitative genetic parameters, and for prediction of the response to selection, are best developed for traits with normal distributions. Many traits of evolutionary interest, including many life history and behavioural traits, have inherently non-normal distributions. The generalised linear mixed model (GLMM) framework has become a widely used tool for estimating quantitative genetic parameters for non-normal traits. However, whereas GLMMs provide inference on a statistically-convenient latent scale, it is sometimes desirable to express quantitative genetic parameters on the scale upon which traits are expressed. The parameters of a fitted GLMMs, despite being on a latent scale, fully determine all quantities of potential interest on the scale on which traits are expressed. We provide expressions for deriving each of such quantities, including population means, phenotypic (co)variances, variance components including additive genetic (co)variances, and parameters such as heritability. We demonstrate that fixed effects have a strong impact on those parameters and show how to deal for this effect by averaging or integrating over fixed effects. The expressions require integration of quantities determined by the link function, over distributions of latent values. In general cases, the required integrals must be solved numerically, but efficient methods are available and we provide an implementation in an R package, QGglmm. We show that known formulae for quantities such as heritability of traits with Binomial and Poisson distributions are special cases of our expressions. Additionally, we show how fitted GLMM can be incorporated into existing methods for predicting evolutionary trajectories. We demonstrate the accuracy of the resulting method for evolutionary prediction by simulation, and apply our approach to data from a wild pedigreed vertebrate population.

Evolutionary Biology

A robust example of collider bias in a genetic association study

A recent paper by Aschard et al described the potential for \"collider bias\" when adjusting for heritable covariates in genetic association studies1. However, in their examples the authors acknowledged that they could not exclude the possibility of a true biological explanation for the genetic association seen only in the adjusted model. Furthermore, the extent to which this bias could create a completely spurious genetic association, rather than just modify the magnitude of the effect,2 remains unclear.\n\nCollider bias describes the artificial association created between two uncorrelated exposures (A and B) when a shared outcome (X) is included in the model as a covariate (Figure 1). We sought to definitively illustrate collider bias by deliberately inducing it to generate a biologically implausible SNP-phenotype association. Both sex (A) and autosomal genetic determinants ...

Genomics

Dissecting the genetic basis of a complex cis-regulatory adaptation

Although single genes underlying several evolutionary adaptations have been identified, the genetic basis of complex, polygenic adaptations has been far more challenging to pinpoint. Here we report that the budding yeast Saccharomyces paradoxus has recently evolved resistance to citrinin, a naturally occurring mycotoxin. Applying a genome-wide test for selection on cis-regulation, we identified five genes involved in the citrinin response that are constitutively up-regulated in S. paradoxus. Four of these genes are necessary for resistance, and are also sufficient to increase the resistance of a sensitive strain when over-expressed. Moreover, cis-regulatory divergence in the promoters of these genes contributes to resistance, while exacting a cost in the absence of citrinin. Our results demonstrate how the subtle effects of individual regulatory elements can be combined, via natural selection, into a complex adaptation. Our approach can be applied to dissect the genetic basis of polygenic adaptations in a wide range of species.\n\nAuthor SummaryAdaptation via natural selection has been a subject of great interest for well over a century, yet we still have little understanding of its molecular basis. What are the genetic changes that are actually being selected? While single genes underlying several adaptations have been identified, the genetic basis of complex, polygenic adaptations has been far more challenging to pinpoint. The complex trait that we study here is the resistance of Saccharomyces yeast to a mycotoxin called citrinin, which is produced by many other species of fungi, and is a common food contaminant for both humans and livestock. We found that sequence changes in the promoters of at least three genes have contributed to citrinin resistance, by up-regulating their transcription even in the absence of citrinin. Higher expression of these genes confers a fitness advantage in the presence of citrinin, while exacting a cost in its absence-- a fitness tradeoff. Our results provide a detailed view of a complex adaptation, and our approach can be applied to polygenic adaptations in a wide range of species.

Evolutionary Biology

Adaptation to heavy-metal contaminated environments proceeds via selection on pre-existing genetic variation

Anthropogenic environmental changes create evolutionary pressures on populations to adapt to novel stresses. It is as yet unclear, when populations respond to these selective pressures, the extent to which this results in convergent genetic evolution and whether convergence is due to independent mutations or shared ancestral variation. We address these questions using a classic example of adaptation by natural selection by investigating the rapid colonization of the plant species Mimulus guttatus to copper contaminated soils. We use field-based reciprocal transplant experiments to demonstrate that mine alleles at a major copper tolerance locus, Tol1, are strongly selected in the mine environment. We assemble the genome of a mine adapted genotype and identify regions of this genome in tight genetic linkage to Tol1. We discover a set of a multicopper oxidase genes that are genetically linked to Tol1 and exhibit large differences in expression between tolerant and non-tolerant genotypes. We overexpressed this gene in M. guttatus and A. thaliana and found the introduced gene contributes to enhanced copper tolerance. We identify convergent adaptation loci that are additional to Tol1 by measuring genome-wide differences in allele frequency between pairs of mine and off-mine populations and narrow these regions to specific candidate genes using differences in protein sequence and gene expression. Furthermore, patterns of genetic variation at the two most differentiated candidate loci are consistent with selection acting upon alleles that predates the existence of the copper mine habitat. These results suggest that adaptation to the mine habitat occurred via selection on ancestral variation, rather than independent de novo mutations or migration between populations.

Evolutionary Biology

Nested Russian Doll-like Genetic Mobility Drives Rapid Dissemination of the Carbapenem Resistance Gene blaKPC

The recent widespread emergence of carbapenem resistance in Enterobacteriaceae is a major public health concern, as carbapenems are a therapy of last resort in this family of common bacterial pathogens. Resistance genes can mobilize via various mechanisms including conjugation and transposition, however the importance of this mobility in short-term evolution, such as within nosocomial outbreaks, is currently unknown. Using a combination of short- and long-read whole genome sequencing of 281 blaKPC-positive Enterobacteriaceae isolated from a single hospital over five years, we demonstrate rapid dissemination of this carbapenem resistance gene to multiple species, strains, and plasmids. Mobility of blaKPC occurs at multiple nested genetic levels, with transmission of blaKPC strains between individuals, frequent transfer of blaKPC plasmids between strains/species, and frequent transposition of the blaKPC transposon Tn4401 between plasmids. We also identify a common insertion site for Tn4401 within various Tn2-like elements, suggesting that homologous recombination between Tn2-like elements has enhanced the spread of Tn4401 between different plasmid vectors. Furthermore, while short-read sequencing has known limitations for plasmid assembly, various studies have attempted to overcome this with the use of reference-based methods. We also demonstrate that as a consequence of the genetic mobility observed herein, plasmid structures can be extremely dynamic, and therefore these reference-based methods, as well as traditional partial typing methods, can produce very misleading conclusions. Overall, our findings demonstrate that non-clonal resistance gene dissemination can be extremely rapid, presenting significant challenges for public health surveillance and achieving effective control of antibiotic resistance.\n\nImportanceIncreasing antibiotic resistance is a major threat to human health, as highlighted by the recent emergence of multi-drug resistant \"superbugs\". Here, we tracked how one important multi-drug resistance gene spread in a single hospital over five years. This revealed high levels of resistance gene mobility to multiple bacterial species, which was facilitated by various different genetic mechanisms. The mobility occurred at multiple nested genetic levels, analogous to a Russian doll set where smaller dolls may be carried along inside larger dolls. Our results challenge traditional views that drug-resistance outbreaks are due to transmission of a single pathogenic strain. Instead, outbreaks can be \"gene-based\", and we must therefore focus on tracking specific resistance genes and their context rather than only specific bacteria.

Microbiology

Genetic diversity on the human X chromosome does not support a strict pseudoautosomal boundary

Unlike the autosomes, recombination between the X chromosome and Y chromosome often thought to be constrained to two small pseudoautosomal regions (PARs) at the tips of each sex chromosome. The PAR1 spans the first 2.7 Mb of the proximal arm of the human sex chromosomes, while the much smaller PAR2 encompasses the distal 320 kb of the long arm of each sex chromosome. In addition to the PAR1 and PAR2, there is a human-specific X-transposed region that was duplicated from the X to the Y. The X-transposed region is often not excluded from X-specific analyses, unlike the PARs, because it is not thought to routinely recombine. Genetic diversity is expected to be higher in recombining regions than in non-recombining regions because recombination reduces the effect of linked selection. In this study, we investigate patterns of genetic diversity in noncoding regions across the entire X chromosome of a global sample of 26 unrelated genetic females. We observe that genetic diversity in the PAR1 is significantly greater than the non-recombining regions (nonPARs). However, rather than an abrupt drop in diversity at the pseudoautosomal boundary, there is a gradual reduction in diversity from the recombining through the non-recombining region, suggesting that recombination between the human sex chromosomes spans across the currently defined pseudoautosomal boundary. In contrast, diversity in the PAR2 is not significantly elevated compared to the nonPAR, suggesting that recombination is not obligatory in the PAR2. Finally, diversity in the X-transposed region is higher than the surrounding nonPAR regions, providing evidence that recombination may occur with some frequency between the X and Y in the XTR.

Evolutionary Biology

HUMAN LONGEVITY IS INFLUENCED BY MANY GENETIC VARIANTS: EVIDENCE FROM 75,000 UK BIOBANK PARTICIPANTS

Variation in human lifespan is 20 to 30% heritable but few genetic variants have been identified. We undertook a Genome Wide Association Study (GWAS) using age at death of parents of middle-aged UK Biobank participants of European decent (n=75,244 with fathers and/or mothers data). Genetic risk scores for 19 phenotypes (n=777 proven variants) were also tested.\n\nGenotyped variants (n=845,997) explained 10.2% (SD=1.3%) of combined parental longevity. In GWAS, a locus in the nicotine receptor CHRNA3 - previously associated with increased smoking and lung cancer - was associated with paternal age at death, with each protective allele (rs1051730[G]) being associated with 0.03 years later age at fathers death (p=3x10-8). Offspring of longer lived parents had more protective alleles (lower genetic risk scores) for coronary artery disease, systolic blood pressure, body mass index, cholesterol and triglyceride levels, type-1 diabetes, inflammatory bowel disease and Alzheimers disease. In candidate gene analyses, variants in the TOMM40/APOE locus were associated with longevity (including rs429358, p=3x10-5), but FOXO variants were not associated.\n\nThese results support a multiple protective factors model for achieving longer lifespans in humans, with a prominent role for cardiovascular-related pathways. Several of these genetically influenced risks, including blood pressure and tobacco exposure, are potentially modifiable.

Genomics

Multi-dimensional structure function relationships in human β-cardiac myosin from population scale genetic variation.

Myosin motors are the fundamental force-generating element of muscle contraction. Variation in the human {beta}-cardiac myosin gene (MYH7) can lead to hypertrophic cardiomyopathy (HCM), a heritable disease characterized by cardiac hypertrophy, heart failure, and sudden cardiac death. How specific myosin variants alter motor function or clinical expression of disease remains incompletely understood. Here, we combine structural models of myosin from multiple stages of its chemomechanical cycle, exome sequencing data from population cohorts of 60,706 and 42,930 individuals, and genetic and phenotypic data from 2,913 HCM patients to elucidate novel structure-function relationships within {beta}-cardiac myosin. We first developed computational models of the human {beta}-cardiac myosin protein before and after the myosin power stroke. Then, using a spatial scan statistic modified to analyze genetic variation in protein three-dimensional space, we found significant enrichment of disease-associated variants in the converter, a kinetic domain that transduces force from the catalytic domain to the lever arm to accomplish the power stroke. Focusing our analysis on surface-exposed residues, we identified another region enriched for disease-associated variants that contains both the converter domain and residues on a single flat surface on the myosin head described as a myosin mesa. This surface is prominent in the pre-stroke model, but substantially reduced in size following the power stroke. Notably, HCM patients with variants in the enriched regions have earlier presentation and worse outcome than those with variants in other regions. In summary, this study provides a model for the combination of protein structure, large-scale genetic sequencing and detailed phenotypic data to reveal insight into time-shifted protein structures and genetic disease.

Genomics

DNA Methylation profiles of diverse Brachypodium distachyon aligns with underlying genetic diversity

DNA methylation, a common modification of genomic DNA, is known to influence the expression of transposable elements as well as some genes. Although commonly viewed as an epigenetic mark, evidence has shown that underlying genetic variation, such as transposable element polymorphisms, often associate with differential DNA methylation states. To investigate the role of DNA methylation variation, transposable element polymorphism, and genomic diversity, whole genome bisulfite sequencing was performed on genetically diverse lines of the model cereal Brachypodium distachyon. Although DNA methylation profiles are broadly similar, thousands of differentially methylated regions are observed between lines. An analysis of novel transposable element indel variation highlighted hundreds of new polymorphisms not seen in the reference sequence. DNA methylation and transposable element variation is correlated with the genome-wide amount of genetic variation present between samples. However, there was minimal evidence that novel transposon insertion or deletions are associated with nearby differential methylation. This study highlights the importance of genetic variation when assessing DNA methylation variation between samples and provides a valuable map of DNA methylation across diverse re-sequenced accessions of this model cereal species.\n\nReviewer Link to deposited dataAll data is publicly available in the NCBI short read archive under BioProject PRJNA281014. Data tables are available for download at https://drive.google.com/file/d/0BzBxfoxlBCneNkY1TFJDU29iSUU/view?usp=sharing

Genomics

Complex ancient genetic structure and cultural transitions in southern African populations.

The characterization of the structure of southern Africa populations has been the subject of numerous genetic, medical, linguistic, archaeological and anthropological investigations. Current diversity in the subcontinent is the result of complex episodes of genetic admixture and cultural contact between the early inhabitants and the migrants that have arrived in the region over the last 2,000 years. Here we analyze 1,856 individuals from 91 populations, comprising novel and available genotype data to characterize the genetic ancestry profiles of 631 individuals from 51 southern African populations. Combining local ancestry and allele frequency analyses we identify a tripartite, ancient, Khoesan-related genetic structure, which correlates with geography, but not with linguistic affiliation or subsistence strategy. The fine mapping of these components in southern African populations reveals admixture dynamics and episodes of cultural reversion involving several Khoesan groups and highlights different mixtures of ancestral components in Bantu speakers and Coloured individuals.

Genomics

On the importance of skewed offspring distributions and background selection in viral population genetics

Many features of virus populations make them excellent candidates for population genetic study, including a very high rate of mutation, high levels of nucleotide diversity, exceptionally large census population sizes, and frequent positive selection. However, these attributes also mean that special care must be taken in population genetic inference. For example, highly skewed offspring distributions, frequent and severe population bottleneck events associated with infection and compartmentalization, and strong purifying selection all affect the distribution of genetic variation but are often not taken in to account. Here, we draw particular attention to multiple-merger coalescent events and background selection, discuss potential mis-inference associated with these processes, and highlight potential avenues for better incorporating them in to future population genetic analyses.

Evolutionary Biology

Genetic and chemical differentiation of Campylobacter coli and Campylobacter jejuni lipooligosaccharide pathways

Despite the importance of lipooligosaccharides (LOS) in the pathogenicity of campylobacteriosis, little is known about the genetic and phenotypic diversity of LOS in C. coli. In this study, we investigated the distribution of LOS locus classes among a large collection of unrelated C. coli isolates sampled from several different host species. Furthermore, we paired C. coli genomic information and LOS chemical composition for the first time to identify mechanisms consistent with the generation of LOS phenotypic heterogeneity. After classifying three new LOS locus classes, only 85% of the 144 isolates tested were assigned to a class, suggesting higher genetic diversity than previously thought. This genetic diversity is at the basis of a completely unexplored LOS structure heterogeneity. Mass spectrometry analysis of the LOS of nine isolates, representing four different LOS classes, identified two features distinguishing C. coli LOS from C. jejunis. GlcN-GlcN disaccharides were present in the lipid A backbone in contrast to the GlcN3N-GlcN backbone observed in C. jejuni. Moreover, despite that many of the genes putatively involved in Qui3pNAcyl were absence in the genomes of various isolates, this rare sugar was found in the outer core of all C. coli. Therefore, regardless the high genetic diversity of LOS biosynthes is locus in C. coli, we identified species-specific phenotypic features of C. coli LOS which might explain differences between C. jejuni and C. coli in terms of population dynamics and host adaptation.\n\nDepositories (where applicable)The whole genome sequences of C. coli are publicly available on the RAST server (http://rast.nmpdr.org) with guest account (login and password guest) under IDs: 195.91, 195.96-195.119, 195.124-195.126, 195.128-195.130, 195.133, 195.134, 6666666.94320

Microbiology

Coalescent inferences in conservation genetics: should the exception become the rule?

Genetic estimates of effective population size (Ne) are an established means to develop informed conservation policies. Another key goal to pursue the conservation of endangered species is keeping the connectivity across fragmented environments, to which genetic inferences of gene flow and dispersal greatly contribute. Most current statistical tools for estimating such population demographic parameters are based on Kingman's coalescent (KC). However, KC is inappropriate for taxa displaying skewed reproductive variance, a property widely observed in natural species. Coalescent models that consider skewed reproductive success-called multiple merger coalescent (MMCs)-have been shown to substantially improve estimates of Ne when the distribution of offspring per capita is highly skewed. MMCs predictions of standard population genetic parameters, including the rate of loss of genetic variation and the fixation probability of strongly selected alleles, substantially depart from KC predictions. These extended models also allow studying gene genealogies in a spatial continuum, providing a novel theoretical framework to investigate spatial connectivity. Therefore, development of statistical tools based on MMC's should substantially improve estimates of population demographic parameters with major conservation implications.

Evolutionary Biology

Genetic surfing in human populations: from genes to genomes

Genetic surfing describes the spatial spread and increase in frequency of variants that are not lost by genetic drift and serial migrant sampling during a range expansion. Genetic surfing does not modify the total number of derived alleles in a population or in an individual genome, but it leads to a loss of heterozygosity along the expansion axis, implying that derived alleles are more often in homozygous state. Genetic surfing also affects selected variants on the wave front, making them behave almost like neutral variants during the expansion. In agreement with theoretical predictions, human genomic data reveals an increase in recessive mutation load with distance from Africa, an expansion load likely to have developed during the expansion of human populations out of Africa.

Genomics