Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

The design and analysis of binary variable traits in common garden genetic experiments of highly fecund species to assess heritability

Many biologically important traits are binomially distributed, with their key phenotypes being presence or absence. Despite their prevalence, estimating the heritability of binomial traits presents both experimental and statistical challenges. Here we develop both an empirical and computational methodology for estimating the narrow-sense heritability of binary traits for highly fecund species. Our experimental approach controls for undesirable culturing effects, while minimizing culture numbers, increasing feasibility in the field. Our statistical approach accounts for known issues with model-selection by using a permutation test to calculate significance values and includes both fitting and power calculation methods. We illustrate our methodology by estimating the narrow-sense heritability for larval settlement, a key life-history trait, in the reef-building coral Orbicella faveolata. The experimental, statistical and computational methods, along with all of the data from this study, were deployed in the R package multiDimBio.

Evolutionary Biology

Estimating K in Genetic Mixture Models

A key quantity in the analysis of structured populations is the parameter K, which describes the number of subpopulations that make up the total population. Inference of K ideally proceeds via the model evidence, which is equivalent to the likelihood of the model. However, the evidence in favour of a particular value of K cannot usually be computed exactly, and instead programs such as SO_SCPLOWTRUCTUREC_SCPLOW make use of simple heuristic estimators to approximate this quantity. We show - using simulated data sets small enough that the true evidence can be computed exactly - that these simple heuristics often fail to estimate the true evidence, and that this can lead to incorrect conclusions about K. Our proposed solution is to use thermodynamic integration (TI) to estimate the model evidence. After outlining the TI methodology we demonstrate the effectiveness of this approach using a range of simulated data sets. We find that TI can be used to obtain estimates of the model evidence that are orders of magnitude more accurate and precise than those based on simple heuristics. Furthermore, estimates of K based on these values are found to be more reliable than those based on a suite of model comparison statistics. Our solution is implemented for models both with and without admixture in the software TO_SCPLOWRUEC_SCPLOWK.

Evolutionary Biology

Genetic evidence challenges the native status of a threatened freshwater fish (Carassius carassius) in England

A fundamental consideration for the conservation of a species is the extent of its native range, however defining a native range is often challenging as changing environments drive shifts in species distributions over time. The crucian carp, Carassius carassius (L.) is a threatened freshwater fish native to much of Europe, however the extent of this range is ambiguous. One particularly contentious region is England, in which C. carassius is currently considered native on the basis of anecdotal evidence. Here, we use 13 microsatellite loci, population structure analyses and approximate bayesian computation (ABC), to empirically test the native status of C. carassius in England. Contrary to the current consensus, ABC yields strong support for introduced origins of C. carassius in England, with posterior distribution estimates placing their introduction in the 15th century, well after the loss of the doggerland landbridge. This result brings to light an interesting and timely debate surrounding our motivations for the conservation of species. We discuss this topic, and make arguments for the continued conservation of C. carassius in England, despite its non-native origins.

Evolutionary Biology

Stochastic Analysis Of An Incoherent Feedforward Genetic Motif

Gene products (RNAs, proteins) often occur at low molecular counts inside individual cells, and hence are subject to considerable random fluctuations (noise) in copy number over time. Not surprisingly, cells encode diverse regulatory mechanisms to buffer noise. One such mechanism is the incoherent feedforward circuit. We analyze a simplistic version of this circuit, where an upstream regulator X affects both the production and degradation of a protein Y. Thus, any random increase in Xs copy numbers would increase both production and degradation, keeping Y levels unchanged. To study its stochastic dynamics, we formulate this network into a mathematical model using the Chemical Master Equation formulation. We prove that if the functional dependence of Ys production and degradation on X is similar, then the steady-distribution of Ys copy numbers is independent of X. To investigate how fluctuations in Y propagate downstream, a protein Z whose production rate only depend on Y is introduced. Intriguingly, results show that the extent of noise in Z increases with noise in X, in spite of the fact that the magnitude of noise in Y is invariant of X. Such counter intuitive results arise because X enhances the time-scale of fluctuations in Y, which amplifies fluctuations in downstream processes. In summary, while feedforward systems can buffer a protein from noise in its upstream regulators, noise can propagate downstream due to changes in the time-scale of fluctuations.

Molecular Biology

foxc1a genetically interacts with ripply1 to regulate mesp-ba expression and somitogenesis in the zebrafish embryo

Somitogenesis is a fundamental segmentation process that forms the vertebrate body plan. A network of transcription factors is essential in establishing the spatial temporal order of this process. One such transcription factor is mesp-ba which has an important role in determining somite boundary formation. Its expression in somitogenesis is tightly regulated by the transcriptional activator Tbx6 and the repressor Ripply1 via a feedback regulatory network. Loss of foxc1a function in zebrafish leads to lack of anterior somite formation and reduced mesp-ba expression. Here we examine how foxc1a interacts with the tbx6-ripply1 network to regulate mesp-ba expression. In foxc1a morphants, anterior somites did not form at 12.5 hours post fertilization (hpf). At 22 hpf posterior somites formed, whereas anterior somites remained absent. In ripply1 morphants, no somites were observed at any time point. The expression of mesp-ba was reduced in the foxc1a morphants and expanded anteriorly in ripply1 morphants. The tbx6 expression domain was smaller and shifted anteriorly in the foxc1a morphants. Double knockdown of foxc1a and ripply1 resulted in absence of anterior somite formation while posterior somites did form, suggesting a partial rescue of the ripply1 phenotype. However, unlike the single foxc1a morphants, expression of mesp-ba was restored in the anterior PSM. Expression of tbx6 was expanded anteriorly in the double morphants. In conclusion, both foxc1a and ripply1 morphants displayed defects in somitogenesis, but their individual loss of function had opposing effects on mesp-ba expression. Loss of ripply1 appears to have rescued the mesp-ba expression in the foxc1a morphant, suggesting that intersection of these parallel regulatory mechanisms is required for normal mesp-ba expression and somite formation.

Developmental Biology

DIRC3 and close to NABP1 Genetics Polymorphisms correlated with Prognostic Survival in Patients with Laryngeal Squamous Cell Carcinoma

Laryngeal squamous cell carcinoma (LSCC) is one of the most common and aggressive malignancies in the upper digestive tract that has a high mortality rate and a poor prognosis. Prognostic factors were determined through multivariate Cox regression analysis. The overall survival rates were calculated by the Kaplan-Meier method. The SPSS statistical software package version 17.0 (SPSS Inc., Chicago, IL, USA) was used for all analyses. Median follow-up was 38 (range 3-122) months and the median survival time was 48 months. We adjusted to confounding factors (total laryngectomy, poor differentiation, T3-T4 stage, N1-N2 stage, III-IV TNM stage) into multivariate Cox proportional hazards model, we confirmed rs11903757 GT genotype (HR = 2.036; 95% CI, 1.071 - 3.872; p = 0.030) and rs966423 TT genotype (HR = 11.677; 95% CI, 3.901 - 34.950; p = 0.000) were significantly correlated with prognostic survival of patients with LSCC compared with rs11903757 TT genotype and rs966423 CC genotype, respectively. Our research provided new evidence for patients with LSCC, it seemed to be the first that demonstrated rs11903757 GT genotype on chromosome 2q32.3 close to NABP1 and rs966423 TT genotype in the intron region of DIRC3 on chromosome 2q35 predict poor prognostic survival in patients with LSCC.

Plant Biology

Genomic data reveal similar genetic differentiation between invertebrates living under and on a riverine floodplain

Little is known about the life histories, population connectivity, or dispersal mechanisms of shallow groundwater organisms. Here we used RAD-seq to analyze population structure in two aquifer species: Paraperla frontalis, a stonefly with groundwater larvae and aerial adults, and Stygobromus sp., a groundwater-obligate amphipod. We found similar levels of connectivity in each species between floodplains separated by ~70 river km in the Flathead River basin of NW Montana, USA. Given that Stygobromous lacks the aboveground life stage of P. frontalis, our findings suggest that aquifer-obligate species might have previously unrecognized dispersal capacity.

Genomics

Medical subject heading (MeSH) annotations illuminate maize genetics and evolution

High-density marker panels and/or whole-genome sequencing,coupled with advanced phenotyping pipelines and sophisticated statistical methods, have dramatically increased our ability to generate lists of candidate genes or regions that are putatively associated with phenotypes or processes of interest. However, the speed with which we can validate genes, or even make reasonable biological interpretations about the principles underlying them, has not kept pace. A promising approach that runs parallel to explicitly validating individual genes is analyzing a set of genes together and assessing the biological similarities among them. This is often achieved via gene ontology (GO) analysis, a powerful tool that involves evaluating publicly available gene annotations. However, additional tools such as Medical Subject Headings (MeSH terms) can also be used to evaluate sets of genes to make biological interpretations. In this manuscript, wedescribe utilizing MeSH terms to make biological interpretations in maize. MeSH terms are assigned to PubMed-indexed manuscripts by the National Library of Medicine, and can be directly mapped to genes to develop gene annotations. Once mapped, these terms can be evaluated for enrichment in sets of genes or similarity between gene sets to provide biological insights. Here, we implement MeSH analyses in five maize datasets to demonstrate how MeSH can be leveraged by the maize and broader crop-genomics community.

Genomics

Population Genetic Analysis of the DARC Locus (Duffy) Reveals Adaptation from Standing Variation Associated with Malaria Resistance in Humans

The human DARC (Duffy antigen receptor for chemokines) gene encodes a membrane-bound chemokine receptor crucial for the infection of red blood cells by Plasmodium vivax, a major causative agent of malaria. Of the three major allelic classes segregating in human populations, the FY*O allele has been shown to protect against P. vivax infection and is near fixation in sub-Saharan Africa, while FY*B and FY*A are common in Europe and Asia, respectively. Due to the combination of its strong geographic differentiation and association with malaria resistance, DARC is considered a canonical example of a locus under positive selection in humans.\n\nHere, we use sequencing data from over 1,000 individuals in twenty-one human populations, as well as ancient human and great ape genomes, to analyze the fine scale population structure of DARC. We estimate the time to most recent common ancestor (TMRCA) of the FY*O mutation to be 42 kya (95% CI: 34-49 kya). We infer the FY*O null mutation swept to fixation in Africa from standing variation with very low initial frequency (0.1%) and a selection coefficient of 0.043 (95% CI:0.011-0.18), which is among the strongest estimated in the genome. We estimate the TMRCA of the FY*A mutation to be 57 kya (95% CI: 48-65 kya) and infer that, prior to the sweep of FY*O, all three alleles were segregating in Africa, as highly diverged populations from Asia and =Khomani San hunter-gatherers share the same FY*A haplotypes. We test multiple models of admixture that may account for this observation and reject recent Asian or European admixture as the cause.\n\nAuthor SummaryInfectious diseases have undoubtedly played an important role in ancient and modern human history. Yet, there are relatively few regions of the genome involved in resistance to pathogens that have shown a strong selection signal. We revisit the evolutionary history of a gene associated with resistance to the most common malaria-causing parasite, Plasmodium vivax, and show that it is one of regions of the human genome that has been under strongest selective pressure in our evolutionary history (selection coefficient: 5%). Our results are consistent with a complex evolutionary history of the locus involving selection on a mutation that was at a very low frequency in the ancestral African population (standing variation) and a large differentiation between European, Asian and African populations.

Genomics

Continuous Genetic Recording with Self-Targeting CRISPR-Cas in Human Cells

The ability to longitudinally track and record molecular events in vivo would provide a unique opportunity to monitor signaling dynamics within cellular niches and to identify critical factors in orchestrating cellular behavior. We present a self-contained analog memory device that enables the recording of molecular stimuli in the form of DNA mutations in human cells. The memory unit consists of a self-targeting guide RNA (stgRNA) cassette that repeatedly directs Streptococcus pyogenes Cas9 nuclease activity towards the DNA that encodes the stgRNA, thereby enabling localized, continuous DNA mutagenesis as a function of stgRNA expression. We analyze the temporal sequence evolution dynamics of stgRNAs containing 20, 30 and 40 nucleotide SDSes (Specificity Determining Sequences) and create a population-based recording metric that conveys information about the duration and/or intensity of stgRNA activity. By expressing stgRNAs from engineered, inducible RNA polymerase (RNAP) III promoters, we demonstrate programmable and multiplexed memory storage in human cells triggered by doxycycline and isopropyl {beta}-D-1-thiogalactopyranoside (IPTG). Finally, we show that memory units encoded in human cells implanted in mice are able to record lipopolysaccharide (LPS)-induced acute inflammation over time. This tool, which we call Mammalian Synthetic Cellular Recorder Integrating Biological Events (mSCRIBE), provides a unique strategy for investigating cell biology in vivo and in situ and may drive further applications that leverage continuous evolution of targeted DNA sequences in mammalian cells.\n\nOne Sentence SummaryBy designing self-targeting guide RNAs that repeatedly direct Cas9 nuclease activity towards their own DNA, we created multiplexed analog memory operators that can record biologically relevant information in vitro and in vivo, such as the magnitude and duration of exposure to TNA.

Synthetic Biology

A reference dataset of 5.4 million phased human variants validated by genetic inheritance from sequencing a three-generation 17-member pedigree

Improvement of variant calling in next-generation sequence data requires a comprehensive, genome-wide catalogue of high-confidence variants called in a set of genomes for use as a benchmark. We generated deep, whole-genome sequence data of seventeen individuals in a three-generation pedigree and called variants in each genome using a range of currently available algorithms. We used haplotype transmission information to create a phased \"platinum\" variant catalogue of 4.7 million single nucleotide variants (SNVs) plus 0.7 million small (1-50bp) insertions and deletions (indels) that are consistent with the pattern of inheritance in the parents and eleven children of this pedigree. Platinum genotypes are highly concordant with the current catalogue of the National Institute of Standards and Technology for both SNVs (>99.99%) and indels (99.92%), and add a validated truth catalogue that has 26% more SNVs and 45% more indels. Analysis of 334,652 SNVs that were consistent between informatics pipelines yet inconsistent with haplotype transmission (\"non-platinum\") revealed that the majority of these variants are de novo and cell-line mutations or reside within previously unidentified duplications and deletions. The reference materials from this study are a resource for objective assessment of the accuracy of variant calls throughout genomes.

Bioinformatics

Coding translational rates: the hidden genetic code

In this paper we propose that translational rate is modulated by pairs of consecutive codons or bicodons. By a statistical analysis of coding sequences, associated with low or with high abundant proteins, we found some bicodons with significant preference usage for either of these sets. These usage preferences cannot be explained by the frequency usage of the single codons. We compute a pause propensity measure of all bicodons in nine organisms, which reveals that in many cases bicodon preference is shared between related organisms. We found that bicodons associated with sequences encoding low abundant proteins are involved in translational attenuation reported in SufIprotein in E. coli. Furthermore, we observe that the misfolding in the drug-transport protein, encoded by MDR1 gene, is better explained by a big change in the pause propensity due to the synonymous bicodon variant, rather than by a relatively small change in the codon usage. These findings suggest that bicodon usage can be a more powerful framework to understand translational speed, protein folding efficiency, and to improve protocols to optimize heterologous gene expression.

Genomics

Functional genetic characterization by CRISPR-Cas9 of two enhancers of FOXP2 in a child with speech and language impairment

Mutations in the coding region of FOXP2 are known to cause speech and language impairment. Microdeletions involving the region downstream the gene have been also associated to speech and cognitive deficits. We recently described a girl harbouring a complex chromosomal rearrangement with one breakpoint downstream the gene that might affect her speech and cognitive abilities via physical separation of distant regulatory DNA elements. In this study, we have used highly efficient targeted chromosomal deletions induced by the CRISPR/Cas9 genome editing tool to demonstrate the functionality of two enhancers, FOXP2-Eproximal and FOXP2-Edistal, located in the intergenic region between FOXP2 and its adjacent MDFIC gene. Deletion of any of these two functional enhancers in the neuroblastomic cell line SK-N-MC downregulates FOXP2 and decreases FOXP2 protein levels, conversely it upregulates MDFIC and increases MDFIC protein levels. This suggests that both regulatory elements may be shared between FOXP2 and MDFIC. We expect these findings contribute to a deeper understanding of how FOXP2 and MDFIC are regulated to pace neuronal development supporting speech and language.

Molecular Biology

A Rationally Designed Aminoacyl-tRNA Synthetase for Genetically Encoded Fluorescent Amino Acids

The incorporation of non-canonical amino acids into proteins has emerged as a promising strategy to manipulate and study protein structure-function relationships with superior precision in vitro and in vivo. To date, fluorescent non-canonical amino acids (f-ncAA) have been successfully incorporated in proteins expressed in bacterial systems, Xenopus oocytes, and HEK-293T cells. Here, we describe the rational generation of an orthogonal aminoacyltRNA synthetase based on the E. coli tyrosine synthetase that is capable of encoding the f-ncAA tyr-coumarin in HEK-293T cells.

Biochemistry

Targeted re-sequencing reveals the genetic signatures and the reticulate history of barley domestication

O_LIBarley (Hordeum vulgare L.) is an established model to study domestication of the Fertile Crescent cereals. Recent molecular data suggested that domesticated barley genomes consist of the ancestral blocks descending from multiple wild barley populations. However, the relationship between the mosaic ancestry patterns and the process of domestication itself remained unclear.\nC_LIO_LITo address this knowledge gap, we identified candidate domestication genes using selection scans based on targeted resequencing of 433 wild and domesticated barley accessions. We conducted phylogenetic, population structure, and ancestry analyses to investigate the origin of the domesticated barley haplotypes separately at the neutral and candidate domestication loci.\nC_LIO_LIWe discovered multiple selective sweeps that occurred on all barley chromosomes during domestication in the background of several ancestral wild populations. The ancestry analyses demonstrated that, although the ancestral blocks of the domesticated barley genomes descended from all over the Fertile Crescent, the candidate domestication loci originated specifically in its eastern and western parts.\nC_LIO_LIThese findings provided first molecular evidence in favor of multiple barley domestications in the Levantine and Zagros clusters of the origin of agriculture.\nC_LI

Plant Biology

Larger Numbers can Impede Adaptation in Microbial Populations despite Entailing Greater Genetic Variation

Periodic bottlenecks play a major role in shaping the adaptive dynamics of natural and laboratory populations of asexual microbes. Here we study how they affect the Extent of Adaptation (EoA), in such populations. EoA, the average fitness gain relative to the ancestor, is the quantity of interest in a large number of microbial experimental-evolution studies which assume that for any given bottleneck size (N0) and number of generations between bottlenecks (g), the harmonic mean size (HM=N0g) will predict the ensuing evolutionary dynamics. However, there are no theoretical or empirical validations for HM being a good predictor of EoA. Using experimental-evolution with Escherichia coli and individual-based simulations, we show that HM fails to predict EoA (i.e., higher N0g does not lead to higher EoA). This is because although higher g allows populations to arrive at superior benefits by entailing increased variation, it also reduces the efficacy of selection, which lowers EoA. We show that EoA can be maximized in evolution experiments by either maximizing N0 and/or minimizing g. We also conjecture that N0/g is a better predictor of EoA than N0g. Our results call for a re-evaluation of the role of population size in predicting fitness trajectories. They also aid in predicting adaptation in asexual populations, which has important evolutionary, epidemiological and economic implications.

Evolutionary Biology

Genetic,transcriptome, proteomic and epidemiological evidence for blood brain barrier disruption and polymicrobial brain invasion as determinant factors in Alzheimers disease.

Multiple pathogens have been detected in Alzheimers disease (AD) brains. A bioinformatics approach was used to assess relationships between pathogens and AD genes (GWAS), the AD hippocampal transcriptome and plaque or tangle proteins. Host/pathogen interactomes (C.albicans, C.Neoformans, Bornavirus, B.Burgdorferri, cytomegalovirus, Ebola virus, HSV-1, HERV-W, HIV-1, Epstein-Barr, hepatitis C, influenza, C.Pneumoniae, P.Gingivalis, H.Pylori, T.Gondii, T.Cruzi) significantly overlap with misregulated AD hippocampal genes, with plaque and tangle proteins and, except Bornavirus, Ebola and HERV-W, with AD genes. Upregulated AD hippocampal genes match those upregulated by multiple bacteria, viruses, fungi or protozoa in immunocompetent blood cells. AD genes are enriched in bone marrow and immune locations and in GWAS datasets reflecting pathogen diversity, suggesting selection for pathogen resistance. The age of AD patients implies resistance to infections afflicting the younger. APOE4 protects against malaria and hepatitis C, and immune/inflammatory gain of function applies to APOE4, CR1, TREM2 and presenilin variants. 30/78 AD genes are expressed in the blood brain barrier (BBB), which is disrupted by AD risk factors (ageing, alcohol, aluminium, concussion, cerebral hypoperfusion, diabetes, homocysteine, hypercholesterolaemia, hypertension, obesity, pesticides, pollution, physical inactivity, sleep disruption and smoking). The BBB and AD benefit from statins, NSAIDs, oestrogen, melatonin and the Mediterranean diet. Polymicrobial involvement is supported by the upregulation of pathogen sensors/defenders (bacterial, fungal, viral) in the AD brain, blood or CSF. Cerebral pathogen invasion permitted by BBB inadequacy, activating a hyper-efficient immune/inflammatory system, betaamyloid and other antimicrobial defence may be responsible for AD which may respond to antibiotic, antifungal or antiviral therapy.

Neuroscience

A comprehensive survey of genetic variation in 20,691 subjects from four large cohorts

The Nurses Health Study (NHS), Nurses Health Study II (NHSII), Health Professionals Follow Up Study (HPFS) and the Physicians Health Study (PHS) have collected detailed longitudinal data on multiple exposures and traits for approximately 310,000 study participants over the last 35 years. Over 160,000 study participants across the cohorts have donated a DNA sample and to date, 20,691 subjects have been genotyped as part of genome-wide association studies (GWAS) of twelve primary outcomes. However, these studies utilized six different GWAS arrays making it difficult to conduct analyses of secondary phenotypes or share controls across studies. To allow for secondary analyses of these data, we have created three new datasets merged by platform family and performed imputation using a common reference panel, the 1,000 Genomes Phase I release. Here, we describe the methodology behind the data merging and imputation and present imputation quality statistics and association results from two GWAS of secondary phenotypes (body mass index (BMI) and venous thromboembolism (VTE)).\n\nWe observed the strongest BMI association for the FTO SNP rs55872725 ({beta}=0.45, p=3.48x10-22), and using a significance level of p=0.05, we replicated 19 out of 32 known BMI SNPs. For VTE, we observed the strongest association for the rs2040445 SNP (OR=2.17, 95% CI: 1.79-2.63, p=2.70x10-15), located downstream of F5 and also observed significant associations for the known ABO and F11 regions. This pooled resource can be used to maximize power in GWAS of phenotypes collected across the cohorts and for studying gene-environment interactions as well as rare phenotypes and genotypes.

epidemiology