Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Population genetics of Paramecium mitochondrial genomes: recombination, mutation spectrum, and efficacy of selection.

The evolution of mitochondrial genomes and their population-genetic environment among unicellular eukaryotes are understudied. Ciliate mitochondrial genomes exhibit a unique combination of characteristics, including a linear organization and the presence of multiple genes with no known function or detectable homologs in other eukaryotes. Here we study the variation of ciliate mitochondrial genomes both within and across thirteen highly diverged Paramecium species, including multiple species from the P. aurelia species complex, with four outgroup species: P. caudatum, P. multimicronucleatum, and two strains that may represent novel related species. We observe extraordinary conservation of gene order and protein-coding content in Paramecium mitochondria across species. In contrast, significant differences are observed in tRNA content and copy number, which is highly conserved in species belonging to the P. aurelia complex but variable among and even within the other Paramecium species. There is an increase in GC content from ~20% to ~40% on the branch leading to the P. aurelia complex. Patterns of polymorphism in population-genomic data and mutation-accumulation experiments suggest that the increase in GC content is primarily due to changes in the mutation spectra in the P. aurelia species. Finally, we find no evidence of recombination in Paramecium mitochondria and find that the mitochondrial genome appears to experience either similar or stronger efficacy of purifying selection than the nucleus.

evolutionary biology

CLADES: A Classification-based Machine Learning Method for Species Delimitation from Population Genetic Data

Species are considered to be the basic unit of ecological and evolutionary studies. Since multi-locus genomic data are becoming increasingly available, there has been considerable interests in the use of DNA sequence data to delimit species. In this paper, we show that machine learning can be used for species delimitation. There exists no species delimitation methods that are based on machine learning. Our method treats the species delimitation problem as a classification problem. It is a problem of identifying the category of a new observation on the basis of training data. Extensive simulation is first conducted over a broad range of evolutionary parameters for training purpose. Each pair of known populations are combined to form training samples with a label of \"same species\" or \"different species\". We use Support Vector Machine (SVM) to train a classifier using a set of summary statistics computed from training samples as features. The trained classifier can classify a test sample to two outcomes: \"same species\" or \"different species\". Given multi-locus genomic data of multiple related organisms or populations, our method (called CLADES) performs species delimitation by first classifying pairs of populations. CLADES then delimits species by maximizing the likelihood of species assignment for multiple populations. CLADES is evaluated through extensive simulation and also tested on real genetic data. We show that CLADES is both accurate and efficient for species delimitation when compared with existing methods. CLADES can be useful especially when existing methods have difficulty in delimitation, e.g. with short species divergence time and gene flow.

evolutionary biology

Developmental genetics in a complex adaptive structure, the weevil rostrum

The rostrum of weevils (Curculionidae) is a novel, complex, adaptive structure that has enabled this huge beetle radiation to feed on and oviposit in a wide spectrum of plant hosts, correlated with diverse life histories and tremendous disparity in rostrum forms. In order to understand the development and evolution of this structure, transcriptomes were produced in de novo assemblies from the developing pre-pupal head tissues of two distantly related curculionids, the rice weevil (Sitophilus oryzae) and the mountain pine beetle (Dendroctonus ponderosae), which have highly divergent rostra. While there are challenges in assessing differences among transcriptomes and in relative gene expression from divergent taxa, tests for differential expression patterns of transcripts yielded lists of candidate genes to examine in future work. RNA interference was performed with S. oryzae for functional insight into the Hox gene Sex combs reduced. Scr has a conserved function in labial and prothoracic identities, but it also demonstrates a novel role in reduction of ventral head structures, namely the gula, submentum, and associated sulci, in weevils. Ultimately, this study makes strides towards elucidating how the weevil rostrum initially formed and the profound phenotypic diversity it has acquired throughout the curculionoid lineages. It furthermore initiates a better understanding of the genetic framework that permitted the diversification of such an immense lineage as the weevils.\n\nSummary statementThis study begins exploring the development of a novel, complex structure in one of the largest families of organisms, the weevils.

developmental biology

Astroplastic: A start-to-finish process for polyhydroxybutyrate production from solid human waste using genetically engineered bacteria to address the challenges for future manned Mars missions

Space exploration has long been a source of inspiration, challenging scientists and engineers to find innovative solutions to various problems. One of the current focuses in space exploration is to send humans to Mars. However, the challenge of transporting materials to Mars and the need for waste management processes are two major obstacles for these long-duration missions.\n\nTo address these two challenges a process called Astroplastic was developed that produces polyhydroxybutyrate (PHB) from solid human waste, which can be used to 3D print useful items for astronauts. PHB granules are naturally produced by bacteria such as Ralstonia eutropha and Pseudomonas aeruginosa for carbon and energy storage. The phaJ, phaC, and phaCBA genes were cloned from these native PHB-producing bacteria into Escherichia coli. These genes code for enzymes that aid in PHB production by converting products of glycolysis and {beta}-oxidation pathways, such as acetyl-CoA and enoyl-CoA, into PHB. To ensure a continuous PHB production system and to eliminate the need for cell lysis to extract PHB, recombinant E. coli was engineered to use the genes in its natural type I secretion system to secrete PHB. The C-terminal of the HlyA secretion tag was fused to phasin (PhaP), a protein originally from R. eutropha. Phasin-HlyA electrostatically binds PHB granules and transports them outside of the cell.\n\nIn addition to genetically engineering bacteria, a concept for start-to-finish PHB production process was designed. Integrating expert feedback and experimental results, conditions for each step of the process including the collection and storage of waste, volatile fatty acid (VFA) fermentation, VFA extraction, PHB fermentation, and PHB extraction were optimized. The optimized system will provide a sustainable and continuous PHB production system, which will address the problems of transportation costs and waste management for future space missions.\n\nFinancial DisclosureMindfuel Science Alberta Foundation Genome Alberta GenScript Polyferm Canada GeekStarter Alberta Integrated DNA Technologies University of Calgary University of Calgary Cumming School of Medicine University of Calgary Bachelor of Sciences University of Calgary Schulich School of Engineering University of Calgary OBrien Centre for the Bachelor of Health Sciences City of Calgary Alberta Innovates The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.\n\nCompeting InterestsThe authors have declared that no competing interests exist.\n\nEthics StatementN/A\n\nData AvailabilityAll data are freely available without restriction.

synthetic biology

Proof of concept for quantitative urine NMR metabolomics pipeline for large-scale epidemiology and genetics

BackgroundQuantitative molecular data from urine are rare in epidemiology and genetics. NMR spectroscopy could provide these data in high-throughput, and it has already been applied in epidemiological settings to analyse urine samples. However, quantitative protocols for large-scale applications are not available.\n\nMethodsWe describe in detail how to prepare urine samples and perform NMR experiments to obtain quantitative metabolic information. Semi-automated quantitative lineshape fitting analyses were set up for 43 metabolites and applied to data from various analytical test samples and from 1,004 individuals from a population-based epidemiological cohort. Novel analyses on how urine metabolites associate with quantitative serum NMR metabolomics data (61 metabolic measures; n=995) were performed. In addition, confirmatory genome-wide analyses of urine metabolites were conducted (n=578). The fully automated quantitative regression-based spectral analysis is demonstrated for creatinine and glucose (n= 4,548).\n\nResultsIntra-assay metabolite variations were mostly <5% indicating high robustness and accuracy of the urine NMR spectroscopy methodology per se. Intra-individual metabolite variations were large, ranging from 6% to 194%. However, population-based inter-individual metabolite variations were even larger (from 14% to 1655%), providing a sound base for epidemiological applications. Metabolic associations between urine and serum were found clearly weaker than those within serum and within urine, indicating that urinary metabolomics data provide independent metabolic information. Two previous genome-wide hits for formate and 2-hydroxyisobutyrate were replicated at genome-wide significance.\n\nConclusionsQuantitative urine metabolomics data suggest broad novelty for systems epidemiology. A roadmap for an open access methodology is provided.

epidemiology

Tissue-specific transcriptome for Poeciliopsis prolifica reveals evidence for genetic adaptation related to the evolution of a placental fish.

The evolution of the placenta is an excellent model to examine the evolutionary processes underlying adaptive complexity due to the recent, independent derivation of placentation in divergent animal lineages. In fishes, the family Poeciliidae offers the opportunity to study placental evolution with respect to variation in degree of post-fertilization maternal provisioning among closely related sister species. In this study, we present a detailed examination of a new reference transcriptome sequence for the live-bearing, matrotrophic fish, Poeciliopsis prolifica, from multiple-tissue RNA-seq data. We describe the genetic components active in liver, brain, late-stage embryo, and the maternal placental/ovarian complex, as well as associated patterns of positive selection in a suite of orthologous genes found in fishes. Results indicate the expression of many signaling transcripts, \"non-coding\" sequences and repetitive elements in the maternal placental/ovarian complex. Moreover, patterns of positive selection in protein sequence evolution were found associated with live-bearing fishes, generally, and the placental P. prolifica, specifically, that appear independent of the general live-bearer lifestyle. Much of the observed patterns of gene expression and positive selection are congruent with the evolution of placentation in fish functionally converging with mammalian placental evolution and with the patterns of rapid evolution facilitated by the teleost-specific whole genome duplication event.

genomics

Genomic characterization of six virus-associated cancers identifies changes in the tumor microenvironment and altered genetic programming

Viruses affect approximately 20% of all human cancers and express immunogenic proteins that make these tumor types potent targets for immune checkpoint inhibitors. In this study, we apply computational tools to The Cancer Genome Atlas and other datasets to define how virus infection shapes the tumor microenvironment and genetic architecture of 6 virus-associated tumor types. Across cancers, the cellular composition of the tumor microenvironment varied based on viral status, with infected tumors often exhibiting increased infiltration of cytolytic cell types. Analyses of the infiltrating T cell receptor repertoire revealed that Epstein-Barr virus was associated with decreased diversity in multiple cancers, suggesting an antigen-driven immune response. Tissue-specific gene expression signatures capturing these virus-induced transcriptomic changes successfully predicted virus status in independent datasets and were associated with both immune- and proliferation-related features that were predictive of prognosis. The analyses presented suggest viruses have distinct effects in different tumors with implications for immunotherapy.

cancer biology

Systematically investigating the key features of the nuclease deactivated Cpf1 for tunable multiplex genetic regulation

With a unique crRNA processing capability, the CRISPR associated Cpf1 protein holds great potential for multiplex gene regulation. Unlike the well-studied Cas9 protein, however, conversion of Cpf1 to a transcription regulator and its related properties have not been systematically explored yet. In this study, we investigated the mutation schemes and crRNA requirements for the nuclease deactivated Cpf1 (dCpf1). By shortening the direct repeat sequence, we obtained genetically stable crRNA co-transcripts and improved gene repression with multiplex targeting. A screen of diversity-enriched PAM library was designed to investigate the PAM-dependency of gene regulation by dCpf1 from Francisella novicida and Lachnospiraceae bacterium. We found novel PAM patterns that elicited strong or medium gene repressions. Using a computational algorithm, we predicted regulatory outputs for all possible PAM sequences, which spanned a large dynamic range that could be leveraged for regulatory purposes. These newly identified features will facilitate the efficient design of CRISPR-dCpf1 based systems for tunable multiplex gene regulation.

synthetic biology

Polygenic adaptation and convergent evolution across both growth and cardiac genetic pathways in African and Asian rainforest hunter-gatherers

Different human populations facing similar environmental challenges have sometimes evolved convergent biological adaptations, for example hypoxia resistance at high altitudes and depigmented skin in northern latitudes on separate continents. The pygmy phenotype (small adult body size), a characteristic of hunter-gatherer populations inhabiting both African and Asian tropical rainforests, is often highlighted as another case of convergent adaptation in humans. However, the degree to which phenotypic convergence in this polygenic trait is due to convergent vs. population-specific genetic changes is unknown. To address this question, we analyzed high-coverage sequence data from the protein-coding portion of the genomes (exomes) of two pairs of populations, Batwa rainforest hunter-gatherers and neighboring Bakiga agriculturalists from Uganda, and Andamanese rainforest hunter-gatherers (Jarawa and Onge) and Brahmin agriculturalists from India. We observed signatures of convergent positive selection between the Batwa and Andamanese rainforest hunter-gatherers across the set of genes with annotated growth factor binding functions (p < 0.001). Unexpectedly, for the rainforest groups we also observed convergent and population-specific signatures of positive selection in pathways related to cardiac development (e.g. cardiac muscle tissue development; p = 0.001). We hypothesize that the growth hormone sub-responsiveness likely underlying the pygmy phenotype may have led to compensatory changes in cardiac pathways, in which this hormone also plays an essential role. Importantly, in the agriculturalist populations we did not observe similar patterns of positive selection on sets of genes associated with either growth or cardiac development, indicating that our results most likely reflect a history of convergent adaptation to the similar ecology of rainforest hunter-gatherers rather than a more common or general evolutionary pattern for human populations.

genomics

Using genetic drug-target networks to develop new drug hypotheses for major depressive disorder

The major depressive disorder (MDD) working group of the Psychiatric Genomics Consortium (PGC) has published a genome-wide association study (GWAS) for MDD in 130,664 cases, identifying 44 risk variants. We used these results to investigate potential drug targets and repurposing opportunities. We built easily interpretable bipartite drug-target networks integrating interactions between drugs and their targets, genome-wide association statistics and genetically predicted expression levels in different tissues, using our online tool Drug Targetor (drugtargetor.com). We also investigated drug-target relationships and drug effects on gene expression that could be impacting MDD. MAGMA was used to perform pathway analyses and S-PrediXcan to investigate the directionality of tissue-specific expression levels in patients vs. controls. Outside the major histocompatibility complex (MHC) region, 25 druggable genes were significantly associated with MDD after multiple testing correction, and 19 were suggestively significant. Several drug classes were significantly enriched, including monoamine reuptake inhibitors, sex hormones, antipsychotics and antihistamines, indicating an effect on MDD and potential repurposing opportunities. These findings require validation in model systems and clinical examination, but also show that GWAS may become a rich source of new therapeutic hypotheses for MDD and other psychiatric disorders that need new - and better - treatment options.

genomics

Genetic regulatory mechanisms of smooth muscle cells map to coronary artery disease risk loci

Coronary artery disease (CAD) is the leading cause of death globally. Genome-wide association studies (GWAS) have identified more than 95 independent loci that influence CAD risk, most of which reside in non-coding regions of the genome. To interpret these loci, we generated transcriptome and whole-genome datasets using human coronary artery smooth muscle cells (HCASMC) from 52 unrelated donors, as well as epigenomic datasets using ATAC-seq on a subset of 8 donors. Through systematic comparison with publicly available datasets from GTEx and ENCODE projects, we identified transcriptomic, epigenetic, and genetic regulatory mechanisms specific to HCASMC. We assessed the relevance of HCASMC to CAD risk using transcriptomic and epigenomic level analyses. By jointly modeling eQTL and GWAS datasets, we identified five genes (SIPA1, TCF21, SMAD3, FES, and PDGFRA) that modulate CAD risk through HCASMC, all of which have relevant functional roles in vascular remodeling. Comparison with GTEx data suggests that SIPA1 and PDGFRA influence CAD risk predominantly through HCASMC, while other annotated genes may have multiple cell and tissue targets. Together, these results provide new tissue-specific and mechanistic insights into the regulation of a critical vascular cell type associated with CAD in human populations.

genomics

Human T cell receptor occurrence patterns encode immune history, genetic background, and receptor specificity

The T cell receptor (TCR) repertoire encodes immune exposure history through the dynamic formation of immunological memory. Statistical analysis of repertoire sequencing data has the potential to decode disease associations from large cohorts with measured phenotypes. However, the repertoire perturbation induced by a given immunological challenge is conditioned on genetic background via major histocompatibility complex (MHC) polymorphism. We explore associations between MHC alleles, immune exposures, and shared TCRs in a large human cohort. Using a previously published repertoire sequencing dataset augmented with high-resolution MHC genotyping, our analysis reveals rich structure: striking imprints of common pathogens, clusters of co-occurring TCRs that may represent markers of shared immune exposures, and substantial variations in TCR-MHC association strength across MHC loci. Guided by atomic contacts in solved TCR:peptide-MHC structures, we identify sequence covariation between TCR and MHC. These insights and our analysis framework lay the groundwork for further explorations into TCR diversity.

immunology

A conformational sensor based on genetic code expansion reveals an autocatalytic component in EGFR activation

Epidermal growth factor receptor (EGFR) activation by growth factors (GFs) relies on dimerization and allosteric activation of its intrinsic kinase activity, resulting in trans-phosphorylation of tyrosines on its C-terminal tail. While structural and biochemical studies identified this EGF-induced allosteric activation, imaging collective EGFR activation in cells and molecular dynamics simulations pointed at additional catalytic EGFR activation mechanisms. To gain more insight in EGFR activation mechanisms in living cells, we developed a Forster Resonance Energy Transfer (FRET) based conformational EGFR indicator (CONEGI) using genetic code expansion that reports on conformational transitions in the EGFR activation loop. Comparing conformational transitions, self-association and auto-phosphorylation of CONEGI and its Y845F mutant revealed that Y845 phosphorylation induces a catalytically active conformation in EGFR monomers. This conformational transition depends on EGFR kinase activity and auto-phosphorylation on its C-terminal tail, generating a looped causality that leads to autocatalytic amplification of EGFR phosphorylation at low EGF dose.

molecular biology

Genetic risk for schizophrenia and developmental delay is associated with shape and microstructure of midline white matter structures

Genomic copy number variants (CNVs) are amongst the most highly penetrant genetic risk factors for neuropsychiatric disorders. The scarcity of carriers of individual CNVs and their phenotypical heterogeneity limits investigations of the associated neural mechanisms and endophenotypes. We applied a novel design based on CNV penetrance for schizophrenia and developmental delay that allows us to identify structural sequelae that are most relevant to neuropsychiatric disorders. Our focus on brain structural abnormalities was based on the hypothesis that convergent mechanisms contributing to neurodevelopmental disorders would likely manifest in the macro- and microstructure of white matter and cortical and subcortical grey matter. 21 adult participants carrying neuropsychiatric risk CNVs (including those located at 22q11.2, 15q11.2, 1q21.1, 16p11.2, and 17q12) and 15 age- and gender matched controls underwent T1-weighted structural, diffusion and quantitative T1 relaxometry MRI.\n\nThe macro- and microstructural properties of the cingulum bundles were associated with penetrance for both developmental delay and schizophrenia, in particular curvature along the anterior-posterior axis (Sz: pcorr=0.026; DD: pcorr=0.035) and intracellular volume fraction (Sz: pcorr=0.019; DD: pcorr=0.064) Further principal component analysis showed alterations in the interrelationships between the volumes of several mid-line white matter structures (Sz: pcorr=0.055; DD, pcorr=0.027). In particular, the ratio of volumes in the splenium and body of the corpus callosum was significantly associated with both penetrance scores (Sz: p=0.037; DD; p=0.006). Our results are consistent with the notion that a significant alteration in developmental trajectories of mid-line white-matter structures constitutes a common neurodevelopmental aberration contributing to risk for schizophrenia and intellectual disability.

neuroscience

A re-inducible genetic cascade patterns the anterior-posterior axis of insects in a threshold-free fashion

Gap genes mediate the division of the anterior-posterior axis of insects into different fates through regulating downstream hox genes. Decades of tinkering the segmentation gene network of the long-germ fruit fly Drosophila melanogaster led to the conclusion that gap genes are regulated (at least initially) through a threshold-based French Flag model, guided by both anteriorly- and posteriorly-localized morphogen gradients. In this paper, we show that the expression patterns of gap genes in the intermediate-germ beetle Tribolium castaneum are mediated by a threshold-free Speed Regulation mechanism, in which the speed of a genetic cascade of gap genes is regulated by a posterior gradient of the transcription factor Caudal. We show this by re-inducing the leading gap gene (namely, hunchback) resulting in the re-induction of the gap gene cascade at arbitrary points in time. This demonstrates that the gap gene network is self-regulatory and is primarily under the control of a posterior speed regulator in Tribolium and possibly all insects.

developmental biology

A genetically encoded fluorescent sensor for in vivo imaging of GABA

Current techniques for monitoring GABA, the primary inhibitory neurotransmitter in vertebrates, cannot follow ephemeral transients in intact neural circuits. We applied the design principles used to create iGluSnFR, a fluorescent reporter of synaptic glutamate, to develop a GABA sensor using a protein derived from a previously unsequenced Pseudomonas fluorescens strain. Structure-guided mutagenesis and library screening led to a usable iGABASnFR ({Delta}F/Fmax ~ 2.5, Kd ~ 9 M, good specificity, adequate kinetics). iGABASnFR is genetically encoded, detects single action potential-evoked GABA release events in culture, and produces readily detectable fluorescence increases in vivo in mice and zebrafish. iGABASnFR enabled tracking of: (1) mitochondrial GABA content and its modulation by an anticonvulsant; (2) swimming-evoked GABAergic transmission in zebrafish cerebellum; (3) GABA release events during inter-ictal spikes and seizures in awake mice; and (4) GABAergic tone decreases during isoflurane anesthesia. iGABASnFR will permit high spatiotemporal resolution of GABA signaling in intact preparations.

neuroscience

A genetically-encoded fluorescent sensor enables rapid and specific detection of dopamine in flies, fish, and mice

Dopamine (DA) is a central monoamine neurotransmitter involved in many physiological and pathological processes. A longstanding yet largely unmet goal is to measure DA changes reliably and specifically with high spatiotemporal precision, particularly in animals executing complex behaviors. Here we report the development of novel genetically-encoded GPCR-Activation-Based-DA (GRABDA) sensors that enable these measurements. In response to extracellular DA rises, GRABDA sensors exhibit large fluorescence increases ({Delta}F/F0[~]90%) with sub-second kinetics, nanomolar to sub-micromolar affinities, and excellent molecular specificity. Importantly, GRABDA sensors can resolve a single-electrical-stimulus evoked DA release in mouse brain slices, and detect endogenous DA release in the intact brains of flies, fish, and mice. In freely-behaving mice, GRABDA sensors readily report optogenetically-elicited nigrostriatal DA release and depict dynamic mesoaccumbens DA changes during Pavlovian conditioning or during sexual behaviors. Thus, GRABDA sensors enable spatiotemporal precise measurements of DA dynamics in a variety of model organisms while exhibiting complex behaviors.

neuroscience

Characteristics and origins of non-functional Pm21 alleles in Dasypyrum villosum and wheat genetic stocks

Most Dasypyrum villosum resources are highly resistant to wheat powdery mildew that carries Pm21 alleles. However, in the previous studies, four D. villosum lines (DvSus-1 [~] DvSus-4) and two wheat-D. villosum addition lines (DA6V#1 and DA6V#3) were reported to be susceptible to powdery mildew. In the present study, the characteristics of non-functional Pm21 alleles in the above resources were analyzed after Sanger sequencing. The results showed that loss-of-functions of Pm21 alleles Pm21-NF1 [~] Pm21-NF3 isolated from DvSus-1, DvSus-2/DvSus-3 and DvSus-4 were caused by two potential point mutations, a 1-bp deletion and a 1281-bp insertion, respectively. The non-functional Pm21 alleles in DA6V#1 and DA6V#3 were same to that in DvSus-4 and DvSus-2/DvSus-3, respectively, indicating that the susceptibilities of the two wheat genetic stocks came from their D. villosum donors. The origins of non-functional Pm21 alleles were also investigated in this study. Except the target variants involved, the sequences of Pm21-NF2 and Pm21-NF3 were identical to that of Pm21-F2 and Pm21-F3 in the resistant D. villosum lines DvRes-2 and DvRes-3, derived from the accessions GRA961 and GRA1114, respectively. It was suggested that the non-functional alleles Pm21-NF2 and Pm21-NF3 originated from the wild-type alleles Pm21-F2 and Pm21-F3. In summary, this study gives an insight into the sequence characteristics of non-functional Pm21 alleles and their origins in natural population of D. villosum.

plant biology