Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Degradation of benzene by the heavy-metal resistant bacterium Cupriavidus metallidurans CH34 reveals its catabolic potential for aromatic compounds

Benzene, toluene, ethylbenzene and the three xylene isomers are monoaromatic contaminants widely distributed on polluted sites. Some microorganisms have developed mechanisms to degrade these compounds, but their aerobic and anaerobic degradation is inhibited in presence of heavy metals, such as mercury or lead. In this report, the degradation of benzene and other aromatic compounds catalyzed by the metal resistant bacterium Cupriavidus metallidurans CH34 was characterized. A metabolic reconstruction of aromatic catabolic pathways was performed based on bioinformatics analyses. Functionality of the predicted pathways was confirmed by growing strain CH34 on benzene, toluene, o-xylene, p-cymene, 3-hydroxybenzoate, 4-hydroxybenzoate, 3-hydroxyphenylacetate, 4-hydroxyphenylacetate, homogentisate, catechol, naphthalene, and 2-aminophenol as sole carbon and energy sources. Benzene catabolic pathway was further characterized. Results showed that firstly benzene is transformed into phenol and, thereafter, into catechol. Benzene is degraded under aerobic conditions via a combined pathway catalyzed by three Bacterial Multicomponent Monooxygenases: a toluene-2-monoxygenase (TomA012345), a toluene-4-monooxygenase (TmoABCDEF) and a phenol-2-hydroxylase (PhyZABCDE). A catechol-2,3-dioxygenase (TomB) expressed at early exponential phase cleaves the catechol ring in meta-position; an ortho-cleavage of catechol is accomplished by a catechol-1,2-dioxygenase (CatA) at late exponential phase instead. This study additionally shows that C. metallidurans CH34 is capable of degrading benzene in presence of heavy metals, such as Hg(II) or Pb(II). This capability of degrading aromatic compounds in presence of heavy metals is rather unusual among environmental bacteria; therefore, C. metallidurans CH34 seems to be a promising candidate for developing novel bioremediation process for multi-contaminated environments.\n\nHIGHLIGHTSO_LIThe strain Cupriavidus metallidurans CH34 is capable to degrade benzene aerobically\nC_LIO_LIBenzene oxydation is mediated by bacterial multicomponent monoxygenases\nC_LIO_LIStrain CH34 is able to grow using a broad range of aromatic compounds as sole carbon and energy source\nC_LIO_LIBenzene degradation occurs even in presence of heavy metals such as mercury and lead\nC_LI

microbiology

ssbio: A Python Framework for Structural Systems Biology

SummaryWorking with protein structures at the genome-scale has been challenging in a variety of ways. Here, we present ssbio, a Python package that provides a framework to easily work with structural information in the context of genome-scale network reconstructions, which can contain thousands of individual proteins. The ssbio package provides an automated pipeline to construct high quality genome-scale models with protein structures (GEM-PROs), wrappers to popular third-party programs to compute associated protein properties, and methods to visualize and annotate structures directly in Jupyter notebooks, thus lowering the barrier of linking 3D structural data with established systems workflows.\n\nAvailability and Implementationssbio is implemented in Python and available to download under the MIT license at http://github.com/SBRG/ssbio. Documentation and Jupyter notebook tutorials are available at http://ssbio.readthedocs.io/en/latest/. Interactive notebooks can be launched using Binder at https://mybinder.org/v2/gh/SBRG/ssbio/master?filepath=Binder.ipynb.\n\nContactnmih@ucsd.edu\n\nSupplementary InformationSupplementary data are available at Bioinformatics online.

systems biology

Trajectory-Based Parameterization of a Coarse-Grained Forcefield for High-Throughput Protein Simulation

The traditional trade-off in biomolecular simulation between accuracy and computational efficiency is predicated on the assumption that detailed forcefields are typically well-parameterized (i.e. obtaining a significant fraction of possible accuracy). We re-examine this trade-off in the more realistic regime in which parameterization is a greater source of bias than the level of detail in the forcefield. To address parameterization of coarse-grained forcefields, we use the contrastive divergence technique from machine learning to train directly from simulation trajectories on 450 proteins. In our scheme, the computational efficiency of the model enables high accuracy through precise tuning of the Boltzmann ensemble over a large collection of proteins. This method is applied to our recently developed Upside model [1], where the free energy for side chains are rapidly calculated at every time-step, allowing for a smooth energy landscape without steric rattling of the side chains. After our contrastive divergence training, the model is able to fold proteins up to approximately 100 residues de novo on a single core in CPU core-days. Additionally, the improved Upside model is a strong starting point both for investigation of folding dynamics and as an inexpensive Bayesian prior for protein physics that can be integrated with additional experimental or bioinformatic data.

biophysics

ciRS-7 exonic sequence is embedded in a long-noncoding RNA locus

ciRS-7 is an intensely studied, highly expressed and conserved circRNA. Essentially nothing is known about its biogenesis, including the location of its promoter. A prevailing assumption has been that ciRS-7 is an exceptional circRNA because it is transcribed from a locus lacking any mature linear RNA transcripts of the same sense. Our interest in the biogenesis of ciRS-7 led us to develop an algorithm to define its promoter. This approach predicted that the human ciRS-7 promoter coincides with that of the long non-coding RNA, LINC00632. We validated this prediction using multiple orthogonal experimental assays. We also used computational approaches and experimental validation to establish that ciRS-7 exonic sequence is embedded in linear transcripts that are flanked by cryptic exons in both human and mouse. Together, this experimental and computational evidence generate a new view of regulation in this locus: (a) ciRS-7 is like other circRNAs, as it is spliced into linear transcripts; (b) expression of ciRS-7 is primarily determined by the chromatin state of LINC00632 promoters; (c) transcription and splicing factors sufficient for ciRS-7 biogenesis are expressed in cells that lack detectable ciRS-7 expression. These findings have significant implications for the study of the regulation and function of ciRS-7, and the analytic framework we developed to jointly analyze RNA-seq and ChlP-seq data reveal the potential for genome-wide discovery of important biological regulation missed in current reference annotations.\n\nAuthor SummarycircRNAs were recently discovered to be a significant product of host gene expression programs but little is known about their transcriptional regulation. Here, we have studied the expression of a well-known circRNA named ciRS-7. ciRS-7 has an unusual function for a circRNA; it is believed to be a miRNA sponge. Previously, ciRS-7 was thought to be transcribed from a locus lacking any mature linear isoforms, unlike all other circular RNAs known to be expressed in human cells. However, we have found this to be false; using a combination of bioinformatic and experimental genetic approaches, in both human and mouse, we discovered that linear transcripts containing the ciRS-7 exonic sequence, linking it to upstream genes. This suggests the potential for additional functional roles of this important locus and provides critical information to begin study on the biogenesis of ciRS-7.

genetics

Very low depth whole genome sequencing in complex trait association studies

MotivationVery low depth sequencing has been proposed as a cost-effective approach to capture low-frequency and rare variation in complex trait association studies. However, a full characterisation of the genotype quality and association power for very low depth sequencing designs is still lacking.\n\nResultsWe perform cohort-wide whole genome sequencing (WGS) at low depth in 1,239 individuals (990 at 1x depth and 249 at 4x depth) from an isolated population, and establish a robust pipeline for calling and imputing very low depth WGS genotypes from standard bioinformatics tools. Using genotyping chip, whole-exome sequencing (WES, 75x depth) and high-depth (22x) WGS data in the same samples, we examine in detail the sensitivity of this approach, and show that imputed 1x WGS recapitulates 95.2% of variants found by imputed GWAS with an average minor allele concordance of 97% for common and low-frequency variants. In our study, 1x further allowed the discovery of 140,844 true low-frequency variants with 73% genotype concordance when compared to high-depth WGS data. Finally, using association results for 57 quantitative traits, we show that very low depth WGS is an efficient alternative to imputed GWAS chip designs, allowing the discovery of up to twice as many true association signals than the classical imputed GWAS design.\n\nSupplementary DataSupplementary Data are appended to this manuscript.

genetics

Estrogen represses Tgfbr1 and Bmpr1a expression via estrogen receptor beta in MC3T3-E1 cells

MC3T3-E1 is a clonal pre-osteoblastic cell line derived from newborn mouse calvaria, which is commonly used in osteoblast studies. To investigate the effects of estrogen on osteoblasts, we treated MC3T3-E1 cells with various concentrations of estrogen and assessed their proliferation. Next, we performed RNA deep sequencing to investigate the effects on estrogen target genes. Bmpr1a and Tgfbr1, important participants in the TGF-beta signaling pathway, were down-regulated in our deep sequencing results. Bioinformatics analysis revealed that estrogen receptor response elements (EREs) were present in the Bmpr1a and Tgfbr1 promoters. Culturing the cells with the estrogen receptor (ER) alpha or beta antagonists 1,3-bis(4-hydroxyphenyl)-4-methyl-5-[4-(2-piperidinylethoxy)phenol]-1H-pyrazole dihydrochloride (MPP) or 4-[2-phenyl-5,7-bis(trifluoromethyl) pyrazolo[1,5-alpha]pyrimidin-3-yl] phenol (PTHPP), respectively, demonstrated that ER beta is involved in the estrogen-mediated repression of Tgfbr1 and Bmpr1a.The chromatin immunoprecipitation (ChIP) results were consistent with the conclusion that E2 increased the binding of ER beta at the EREs located in the Tgfbr1 and Bmpr1a promoters. Our research provides new insight into the role of estrogen in bone metabolisms.

biochemistry

Fundamental limits on dynamic inference from single cell snapshots

Single cell expression profiling reveals the molecular states of individual cells with unprecedented detail. However, because these methods destroy cells in the process of analysis, they cannot measure how gene expression changes over time. But some information on dynamics is present in the data: the continuum of molecular states in the population can reflect the trajectory of a typical cell. Many methods for extracting single cell dynamics from population data have been proposed. However, all such attempts face a common limitation: for any measured distribution of cell states, there are multiple dynamics that could give rise to it, and by extension, multiple possibilities for underlying mechanisms of gene regulation. Here, we describe the aspects of gene expression dynamics that cannot be inferred from a static snapshot alone and identify assumptions necessary to constrain a unique solution for cell dynamics from static snapshots. We translate these constraints into a practical algorithmic approach, Population Balance Analysis (PBA), which makes use of a method from spectral graph theory to solve a class of high dimensional differential equations. We use simulations to show the strengths and limitations of PBA, and then apply it to single-cell profiles of hematopoietic progenitor cells (HPCs). Cell state predictions from this analysis agree with HPC fate assays reported in several papers over the past two decades. By highlighting the fundamental limits on dynamic inference faced by any method, our framework provides a rigorous basis for dynamic interpretation of a gene expression continuum and clarifies best experimental designs for trajectory reconstruction from static snapshot measurements.\n\nSignificanceSeeing a snapshot of individuals at different stages of a process can reveal what the process would look like for a single individual over time. Biologists apply this principle to infer temporal sequences of gene expression states in cells from measurements made at a single moment in time. However, these inferences are fundamentally under-determined. Using a conservation law, we enumerate reasons that there is no unique dynamics associated with a single snapshot, limiting our ability to infer gene regulatory mechanisms. We then propose a method for dynamic inference that provides a unique dynamic solution under defined approximations and apply it to data from bone marrow stem cells. Overall, this study introduces formal biophysical approaches to single cell bioinformatics.\n\nClassificationBIOLOGICAL SCIENCES / Systems Biology

systems biology

Euplotid: A Linux-based platform to physically edit the genome

1Life continues to shock and amaze us, reminding us that truth is far stranger than fiction. http://Euplotid.io is a quantized geometric model of the eukaryotic cell, an attempt at quantifying the incredible complexity that gives rise to a living cell by beginning from the smallest unit, a quanta. Starting from the very bottom we are able to build the pieces which when hierarchically and combinatorially combined produce the emergent complex behavior that even a single celled organism can show. Euplotid is composed of a set of quantized geometric 3D building blocks and constantly evolving dockerized bioinformatic pipelines enabling a user to build and interact with the local regulatory architecture of every gene starting from DNA-interactions, chromatin accessibility, and RNA-sequencing. Reads are quantified using the latest computational tools and the results are normalized, quality-checked, and stored. The local regulatory architecture of each gene is built using a Louvain based graph partitioning algorithm parameterized by the chromatin extrusion model and CTCF-CTCF interactions. Cis-Regulatory Elements are defined using chromatin accessibility peaks which are mapped to Transcriptional Start Sites based on inclusion within the same neighborhood. Deep Neural Networks are trained in order to provide a statistical model mimicking transcription factor binding, giving the ability to identify all Transcription Factors within a given chromatin accessibility peak. By in-silico mutating and re-applying the neural network we are able to gauge the impact of a transition mutation on the binding of any transcription factor. The annotated output can be visualized in a variety of 1D, 2D, 3D and 4D ways overlaid with existing bodies of knowledge such as GWAS results or PDB structures. Once a particular CRE of interest has been identified a Base Editor mediated transition mutation can then be performed in a relevant model for further study. O_FIG O_LINKSMALLFIG WIDTH=182 HEIGHT=200 SRC="FIGDIR/small/170159v15_ufig1.gif" ALT="Figure 1"> View larger version (39K): org.highwire.dtl.DTLVardef@1a55a1borg.highwire.dtl.DTLVardef@bebb0dorg.highwire.dtl.DTLVardef@1ea5b11org.highwire.dtl.DTLVardef@100e735_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 0.1:C_FLOATNO Graphical Abstract C_FIG

biophysics

Expanding primary cells from mucoepidermoid and other salivary gland neoplasms for genetic and chemosensitivity testing

Restricted availability of cell and animal models is a rate-limiting step for investigation of salivary gland neoplasm pathophysiology and therapeutic response. Conditionally reprogrammed cell (CRC) technology enables establishment of primary epithelial cell cultures from patient material. This study tested a translational workflow for acquisition, expansion and testing of CRC-derived primary cultures of salivary gland neoplasms from patients presenting to an academic surgical practice. Results showed cultured cells were sufficient for epithelial cell-specific transcriptome characterization to detect candidate therapeutic pathways and fusion genes in addition to screening for cancer-risk-associated single nucleotide polymorphisms (SNPs) and driver gene mutations through exome sequencing. Focused study of primary cultures of a low-grade mucoepidermoid carcinoma demonstrated Amphiregulin-Mechanistic Target of Rapamycin-AKT/Protein kinase B (AKT) pathway activation, identified through bioinformatics and subsequently confirmed as present in primary tissue and preserved through different secondary 2D and 3D culture media and xenografts. Candidate therapeutic testing showed that the allosteric AKT inhibitor MK2206 reproducibly inhibited cell survival across different culture formats. In contrast, the cells appeared resistant to the adenosine triphosphate competitive AKT inhibitor GSK690693. Procedures employed here illustrate an approach for reproducibly obtaining material for pathophysiological studies of salivary gland neoplasms, and other less common epithelial cancer types, that can be executed without compromising pathological examination of patient specimens. The approach permits combined genetic and cell-based physiological and therapeutic investigations in addition to more traditional pathologic studies and can be used to build sustainable bio-banks for future inquiries.

cancer biology

Structural basis of the Cope rearrangement and C-C bond-forming cascade in hapalindole/fischerindole biogenesis

STRUCTURESThe atomic coordinates and structure factors for:\n\nHpiC1 W73M/K132M SeMet (P212121) -1.7 [A]\n\nHpiC1 native (C2) -1.5 [A]\n\nHpiC1 native (P42) -2.1 [A]\n\nHpiC1 Y101F (C2) -1.4 [A]\n\nHpiC1 Y101S (C2) -1.4 [A]\n\nHpiC1 F138S (P21) -1.7 [A]\n\nHpiC1 Y101F/F138S (P21 -1.65 [A] have been deposited with the Research Collaboratory for Structural Bioinformatics as Protein Data Bank entries 5WPP, 5WPR, 6AL6, 5WPR, 5WPU, 6AL7, and 6AL8 (www.rcsb.org).\n\nGRANTSThis work was supported by: The authors thank the National Science Foundation under the CCI Center for Selective C-H Functionalization (CHE-1205646), the National Institutes of Health (CA70375 to RMW and DHS), R35 GM118101, R01 GM076477 and the Hans W. Vahlteich Professorship (to DHS) for financial support. M.G-B. thanks the Ramon Areces Foundation for a postdoctoral fellowship. J.N.S. acknowledges the support of the National Institute of General Medical Sciences of the National Institutes of Health under Award Number F32GM122218. Computational resources were provided by the UCLA Institute for Digital Research and Education (IDRE) and the Extreme Science and Engineering Discovery Environment (XSEDE), which is supported by the NSF (OCI-1053575). The content does not necessarily represent the official views of the National Institutes of Health.\n\nABSTRACTHapalindole alkaloids are a structurally diverse class of cyanobacterial natural products defined by their varied polycyclic ring systems and diverse biological activities. These polycyclic scaffolds are generated from a common biosynthetic intermediate by the Stig cyclases in three mechanistic steps, including a rare Cope-rearrangement, 6-exo-trig cyclization, and electrophilic aromatic substitution. Here we report the structure of HpiC1, a Stig cyclase that catalyzes the formation of 12-epi-hapalindole U in vitro. The 1.5 [A] structure reveals a dimeric assembly with two calcium ions per monomer and the active sites located at the distal ends of the protein dimer. Mutational analysis and computational methods uncovered key residues for an acid catalyzed [3,3]-sigmatropic rearrangement and specific determinants that control the position of terminal electrophilic aromatic substitution leading to a switch from hapalindole to fischerindole alkaloids.

biochemistry

Global accumulation of circRNAs during aging in Caenorhabditis elegans

Circular RNAs (CircRNAs) are a newly appreciated class of RNAs that lack free 5{acute} and 3{acute} ends, are expressed by the thousands in diverse forms of life, and are mostly of enigmatic function. Ostensibly due to their resistance to exonucleases, circRNAs are known to be exceptionally stable. Here, we examined the global profile of circRNAs in C. elegans during aging by performing ribo-depleted total RNA-seq from the fourth larval stage (L4) through 10-day old adults. Using stringent bioinformatic criteria and experimental validation, we annotated 1,166 circRNAs, including 575 newly discovered circRNAs. These circRNAs were derived from 797 genes with diverse functions, including genes involved in the determination of lifespan. A massive accumulation of circRNAs during aging was uncovered. Many hundreds of circRNAs were significantly increased among the aging time-points and increases of select circRNAs by over 40-fold during aging were quantified by qRT-PCR. The age-accumulation of circRNAs was not accompanied by increased expression of linear RNAs from the same host genes. We attribute the global scale of circRNA age-accumulation to the high composition of postmitotic cells in adult C. elegans, coupled with the high resistance of circRNAs to decay. These findings suggest that the exceptional stability of circRNAs might explain age-accumulation trends observed from neural tissues of other organisms, which also have a high composition of post-mitotic cells. Given the suitability of C. elegans for aging research, it is now poised as an excellent model system to determine if there are functional consequences of circRNA accumulation during aging.

genomics

Metagenomic Insights into The Diversity and Functions of Microbial Assemblages in Tasik Kenyir Ecosystem

Tropical freshwater lake such as Tasik Kenyir are underrepresented among the growing number of environmental metagenomic data sets. In Tasik Kenyir, water from two different sites, pristine and disturbed areas were sampled. After the filtration process, genomic DNA from both sites were extracted using Meta-G-nome DNA isolation kit and shotgun metagenomic sequencing was carried out on Illumina HiSeq2500 Desktop Sequencer (Illumina, Inc.). Raw data were then trimmed and assembled using Metagenomic Assembler program, MetaVelvet. Data analysis was carried out using software Blast2GO (BioBam Bioinformatic S.L). The total number of sequence reads was 189,158 from TKS1.5m (disturbed area) and 246,577 from TKS2.5m (pristine area).The results indicate that sequence reads of microbial species were presence at disturbed area near the aquaculture zone was lower than the sequence reads of microbial species were presence at pristine area. When compared to archaea, both samples were dominated by bacteria (more than 90%) suggesting that bacteria are absolutely dominant in the prokaryotic communities in the freshwater samples. The lake appears to contain a mixture of autotrophs and heterotrophs capable of performing main biogeochemical cycles like nitrogen fixation by Klebsiella sp for TKS1.5m and Pontibacter sp. for TKS2.5m. and carbon fixation by heterotrophic Alcaligenes sp. and Shewanella decolorationi in TKS1.5m, and by Pantoea sp. in TKS2.5m. Present study will advance our understanding of the importance of freshwater microbial communities for ecosystem and human health.

microbiology

Genome-wide Enhancer Maps Differ Significantly in Genomic Distribution, Evolution, and Function

Non-coding gene regulatory enhancers are essential to transcription in mammalian cells. As a result, a large variety of experimental and computational strategies have been developed to identify cis-regulatory enhancer sequences. In practice, most studies consider enhancers identified by only a single method, and the concordance of enhancers identified by different methods has not been comprehensively evaluated. Here, we assess the similarities of enhancer sets identified by ten representative strategies in four biological contexts and evaluate the robustness of downstream conclusions to the choice of identification strategy. All pairs of enhancer sets we evaluated overlap significantly more than expected by chance; however, we also found significant dissimilarity between enhancer sets in their genomic characteristics, evolutionary conservation, and association with functional loci within each context. We find most regions identified as enhancers are supported by only one method. The disagreement is sufficient to influence interpretation of GWAS SNPs and eQTL, and to lead to disparate conclusions about enhancer biology and disease mechanisms. We also find only limited evidence that regions identified by multiple enhancer identification methods are better candidates than those identified by a single method. Our results highlight the inherent complexity of enhancer biology and argue that current approaches have yet to adequately account for enhancer diversity. As a result, we cannot recommend the use of any single enhancer identification strategy in isolation. To facilitate assessment of enhancer diversity on studies conclusions, we developed creDB, a database of enhancer annotations designed to integrate into bioinformatics workflows. While our findings highlight a major challenge to mapping the genetic architecture of complex disease and interpreting regulatory variants found in patient genomes, a systematic understanding of similarities and differences in enhancer identification methodology will ultimately enable robust inferences about gene regulatory sequences.

genetics

Dynamic marine viral infections and major contribution to photosynthetic processes shown by regional and seasonal picoplankton metatranscriptomes

Viruses are an important top-down control on microbial communities, yet their direct study in natural environments has been hindered by culture limitations1-3. The advance of sequencing and bioinformatics over the last decade enabled the cultivation independent study of viruses. Many studies focus on assembling new viral genomes4-6 and studying viral diversity using marker genes amplified from free viruses7,8. We used cellular metatranscriptomics to study community-wide viral infections at three coastal California sites throughout a year. Generation of and recruitment to viral contigs (> 5kbp, N=66) allowed tracking of infection dynamics over time and space. Here we show that while these assemblies represent viral populations, they are likely biased towards clonal or low diversity assemblages. Furthermore, we demonstrate that published T4-like cyanophages (N=50) and pelagiphages (N=4), having genomic continuity between close relatives, are better tracked using marker genes. Additionally, we demonstrate determination of potential hosts by matching infection dynamics with microbial community composition. Finally, we quantify the relative contribution of various cyanobacteria and viruses to photosystem-II psbA expression in our study sites. We show sometimes >50% of all cyanobacterial+viral psbA expression we observed is of viral origin, which highlights the proportion of infected cells and makes viruses a remarkable contributor to photosynthesis and oxygen production.

ecology

MicrobiomeDB: a systems biology platform for integrating, mining and analyzing microbiome experiments.

MicrobiomeDB (http://microbiomeDB.org) is a data discovery and analysis platform that empowers researchers to fully leverage experimental variables to interrogate microbiome datasets. MicrobiomeDB was developed in collaboration with the Eukaryotic Pathogens Bioinformatics Resource Center (http://EuPathDB.org) and leverages the infrastructure and user interface of EuPathDB, which allows users to construct in silico experiments using an intuitive graphical strategy approach. The current release of the database integrates microbial census data with sample details for nearly 14,000 samples originating from human, animal and environmental sources, including over 9,000 samples from healthy human subjects in the Human Microbiome Project (http://portal.ihmpdcc.org/). Query results can be statistically analyzed and graphically visualized via interactive web applications launched directly in the browser, providing insight into microbial community diversity and allowing users to identify taxa associated with any experimental covariate.

microbiology

A genome resequencing-based genetic map reveals the recombination landscape of an outbred parasitic nematode in the presence of polyploidy and polyandry

The parasitic nematode Haemonchus contortus is an economically and clinically important pathogen of small ruminants, and a model system for understanding the mechanisms and evolution of traits such as anthelmintic resistance. Anthelmintic resistance is widespread and is a major threat to the sustainability of livestock agriculture globally; however, little is known about the genome architecture and parameters such as recombination that will ultimately influence the rate at which resistance may evolve and spread. Here we performed a genetic cross between two divergent strains of H. contortus, and subsequently used whole-genome re-sequencing of a female worm and her brood to identify the distribution of genome-wide variation that characterises these strains. Using a novel bioinformatic approach to identify variants that segregate as expected in a pseudo-testcross, we characterised linkage groups and estimated genetic distances between markers to generate a chromosome-scale F1 genetic map composed of 1,618 SNPs. We exploited this map to reveal the recombination landscape, the first for any parasitic helminth species, demonstrating extensive variation in recombination rate within and between chromosomes. Analyses of these data also revealed the extent of polyandry, whereby at least eight males were found to have contributed to the genetic variation of the progeny analysed. Triploid offspring were also identified, which we hypothesise are the result of nondisjunction during female meiosis or polyspermy. These results expand our knowledge of the genetics of parasitic helminths and the unusual life-history of H. contortus, and will enable more precise characterisation of the evolution and inheritance of genetic traits such as anthelmintic resistance. This study also demonstrates the feasibility of whole-genome resequencing data to directly construct a genetic map in a single generation cross from a non-inbred non-model organism with a complex lifecycle.\n\nAuthor summaryRecombination is a key genetic process, responsible for the generation of novel genotypes and subsequent phenotypic variation as a result of crossing over between homologous chromosomes. Populations of strongylid nematodes, such as the gastrointestinal parasites that infect livestock and humans, are genetically very diverse, but little is known about patterns of recombination across the genome and how this may contribute to the genetics and evolution of these pathogens. In this study, we performed a genetic cross to quantify recombination in the barbers pole worm, Haemonchus contortus, an important parasite of sheep and goats. The reproductive traits of this worm make standard genetic crosses challenging, but by generating whole-genome sequence data from a female worm and her offspring, we identified genetic variants that act as though they come from a single mating cross, allowing the use of standard statistical approaches to build a genetic map and explore the distribution and rates of recombination throughout the genome. A number of genetic signatures associated with H. contortus life history traits were revealed in this analysis: we extend our understanding of multiple paternity (polyandry) in this species, and provide evidence and explanation for sporadic increases in chromosome complements (polyploidy) among the progeny. The resulting genetic map will aid in population genomic studies in general and enhance ongoing efforts to understand the genetic basis of resistance to the drugs used to control these worms, as well as for related species that infect humans throughout the world.

genomics

Development and validation of an expanded carrier screen that optimizes sensitivity via full-exon sequencing and panel-wide copy-number-variant identification

PurposeBy identifying pathogenic variants across hundreds of genes, expanded carrier screening (ECS) enables prospective parents to assess risk of transmitting an autosomal recessive or X-linked condition. Detection of at-risk couples depends on the number of conditions tested, the diseases respective prevalences, and the screens sensitivity for identifying disease-causing variants. Here we present an analytical validation of a 235-gene sequencing-based ECS with full coverage across coding regions, targeted assessment of pathogenic noncoding variants, panel-wide copy-number-variant (CNV) calling, and customized assays for technically challenging genes.\n\nMethodsNext-generation sequencing, a customized bioinformatics pipeline, and expert manual call review were used to identify single-nucleotide variants, short insertions and deletions, and CNVs for all genes except FMR1 and those whose low disease incidence or high technical complexity precludes novel variant identification or interpretation. Variant calls were compared to reference and orthogonal data.\n\nResultsValidation of our ECS data demonstrated >99% analytical sensitivity and >99% specificity. A preliminary assessment of 15,177 patient samples reveals the substantial impact on fetal disease-risk detection attributable to novel CNV calling (13.9% of risk) and technically challenging conditions (15.5% of risk), such as congenital adrenal hyperplasia.\n\nConclusionValidated, high-fidelity identification of different variant types--especially in diseases with complicated molecular genetics--maximizes at-risk couple detection.

genetics

CXCR4 involvement in neurodegenerative diseases

Neurodegenerative diseases likely share common underlying pathobiology. Although prior work has identified susceptibility loci associated with various dementias, few, if any, studies have systematically evaluated shared genetic risk across several neurodegenerative diseases. Using genome-wide association data from large studies (total n = 82,337 cases and controls), we utilized a previously validated approach to identify genetic overlap and reveal common pathways between progressive supranuclear palsy (PSP), frontotemporal dementia (FTD), Parkinsons disease (PD) and Alzheimers disease (AD). In addition to the MAPT H1 haplotype, we identified a variant near the chemokine receptor CXCR4 that was jointly associated with increased risk for PSP and PD. Using bioinformatics tools, we found strong physical interactions between CXCR4 and four microglia related genes, namely CXCL12, TLR2, RALB and CCR5. Evaluating gene expression from post-mortem brain tissue, we found that expression of CXCR4 and microglial genes functionally related to CXCR4 was dysregulated across a number of neurodegenerative diseases. Furthermore, in a mouse model of tauopathy, expression of CXCR4 and functionally associated genes was significantly altered in regions of the mouse brain that accumulate neurofibrillary tangles most robustly. Beyond MAPT, we show dysregulation of CXCR4 expression in PSP, PD, and FTD brains, and mouse models of tau pathology. Our multi-modal findings suggest that abnormal signaling across a network of microglial genes may contribute to neurodegeneration and may have potential implications for clinical trials targeting immune dysfunction in patients with neurodegenerative diseases.

genetics