Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Rosetta Antibody Design (RAbD): A General Framework for Computational Antibody Design

A structural-bioinformatics-based computational methodology and framework have been developed for the design of antibodies to targets of interest. RosettaAntibodyDesign (RAbD) samples the diverse sequence, structure, and binding space of an antibody to an antigen in highly customizable protocols for the design of antibodies in a broad range of applications. The program samples antibody sequences and structures by grafting structures from a widely accepted set of the canonical clusters of CDRs (North et al., J. Mol. Biol., 406:228-256, 2011). It then performs sequence design according to amino acid sequence profiles of each cluster, and samples CDR backbones using a flexible-backbone design protocol incorporating cluster-based CDR constraints. Starting from an existing experimental or computationally modeled antigen-antibody structure, RAbD can be used to redesign a single CDR or multiple CDRs with loops of different length, conformation, and sequence. We rigorously benchmarked RAbD on a set of 60 diverse antibody-antigen complexes, using two design strategies - optimizing total Rosetta energy and optimizing interface energy alone. We utilized two novel metrics for measuring success in computational protein design. The design risk ratio (DRR) is equal to the frequency of recovery of native CDR lengths and clusters divided by the frequency of sampling of those features during the Monte Carlo design procedure. Ratios greater than 1.0 indicate that the design process is picking out the native more frequently than expected from their sampled rate. We achieved DRRs for the non-H3 CDRs of between 2.4 and 4.0. The antigen risk ratio (ARR) is the ratio of frequencies of the native amino acid types, CDR lengths, and clusters in the output decoys for simulations performed in the presence and absence of the antigen. For CDRs, we achieved cluster ARRs as high as 2.5 for L1 and 1.5 for H2. For sequence design simulations without CDR grafting, the overall recovery for the native amino acid types for residues that contact the antigen in the native structures was 72% in simulations performed in the presence of the antigen and 48% in simulations performed without the antigen, for an ARR of 1.5. For the non-contacting residues, the ARR was 1.08. This shows that the sequence profiles are able to maintain the amino acid types of these conserved, buried sites, while recovery of the exposed, contacting residues requires the presence of the antigen-antibody interface. We tested RAbD experimentally on both a lambda and kappa antibody-antigen complex, successfully improving their affinities 10 to 50 fold by replacing individual CDRs of the native antibody with new CDR lengths and clusters.\n\nAuthor SummaryAntibodies are proteins produced by the immune system to attack infections and cancer and are also used as drugs to treat cancer and autoimmune diseases. The mechanism that has evolved to produce them is able to make 10s of millions of different antibodies, each with a different surface used to bind the foreign or mutated molecule. We have developed a method to design antibodies computationally, based on the 1000s of experimentally determined three-dimensional structures of antibodies available. The method works by treating pieces of these structures as a collection of parts that can be combined in new ways to make better antibodies. Our method has been implemented in the protein modeling program Rosetta, and is called RosettaAntibodyDesign (RAbD). We tested RAbD both computationally and experimentally. The experimental test shows that we can improve existing antibodies by 10 to 50 fold, paving the way for design of entirely new antibodies in the future.

bioengineering

Precision genome-editing with CRISPR/Cas9 in human induced pluripotent stem cells

Genome engineering in human induced pluripotent stem cells (iPSCs) represent an opportunity to examine the contribution of pathogenic and disease modifying alleles to molecular and cellular phenotypes. However, the practical application of genome-editing approaches in human iPSCs has been challenging. We have developed a precise and efficient genome-editing platform that relies on allele-specific guideRNAs (gRNAs) paired with a robust method for culturing and screening the modified iPSC clones. By applying an allele-specific gRNA design strategy, we have demonstrated greatly improved editing efficiency without the introduction of additional modifications of unknown consequence in the genome. Using this approach, we have modified nine independent iPSC lines at five loci associated with neurodegeneration. This genome-editing platform allows for efficient and precise production of isogenic cell lines for disease modeling. Because the impact of CRISPR/Cas9 on off-target sites remains poorly understood, we went on to perform thorough off-target profiling by comparing the mutational burden in edited iPSC lines using whole genome sequencing. The bioinformatically predicted off-target sites were unmodified in all edited iPSC lines. We also found that the numbers of de novo genetic variants detected in the edited and unedited iPSC lines were similar. Thus, our CRISPR/Cas9 strategy does not specifically increase the mutational burden. Furthermore, our analyses of the de novo genetic variants that occur during iPSC culture and genome-editing indicate an enrichment of de novo variants at sites identified in dbSNP. Taken together, we propose that this enrichment represents regions of the genome more susceptible to mutation. Herein, we present an efficient and precise method for allele-specific genome-editing in iPSC and an analyses pipeline to distinguish off-target events from de novo mutations occurring with culture.

genetics

A potential link between tuberculosis and lung cancer through non-coding RNAs

Pulmonary tuberculosis caused by Mycobacterium and lung cancer are two major causes of deaths worldwide and the former increases the risk of developing lung cancer. However, the precise molecular mechanism of Mycobacterium associated increased risk of lung cancer is not entirely understood. Here, using in silico approaches, we show that hsa-mir-21 and M. tuberculosis sRNA_1096 and sRNA_1414 could play important roles in the pathogenesis of both these diseases. Further, we postulated a \"Genetic remittance\" hypothesis where these sRNAs may play important roles. The sRNA_1096 could be involved in tuberculosis through multiple infectious processes, and if transferred to the host, it may activate the TLR8 mediated pro-metastatic inflammatory pathway by acting as a ligand to TLR8 similar to the mir-21 leading to lung tumorigenesis and chemo-resistance. Analogous to SH3GL1, it may also regulate cell cycle. On the other hand, sRNA_1414 is probably involved in survivability and drug response of the pathogen. However, it may be a metastatic factor for lung cancer providing EPS8L1 and SORBS1 like functions upon remittance. Further, all these three non-coding RNAs are predicted to act in rifampicin resistance in Mycobacterium. Currently, we are applying robust bioinformatics strategies and conducting experimental validations to confirm our in-silico findings and hypothesis.

cancer biology

The Early Diagnosis in Lung Cancer by the Detection of Circulating Tumor DNA

BackgroundRemarkable advances for clinical diagnosis and treatment in cancers including lung cancer involve cell-free circulating tumor DNA (ctDNA) detection through next generation sequencing. However, before the sensitivity and specificity of ctDNA detection can be widely recognized, the consistency of mutations in tumor tissue and ctDNA should be evaluated. The urgency of this consistency is extremely obvious in lung cancer to which great attention has been paid to in liquid biopsy field.\n\nMethodsWe have developed an approach named systematic error correction sequencing (Sec-Seq) to improve the evaluation of sequence alterations in circulating cell-free DNA. Averagely 10 ml preoperative blood samples were collected from 30 patients containing pulmonary space occupying pathological changes by traditional clinic diagnosis. cfDNA from plasma, genomic DNA from white blood cells, and genomic DNA from solid tumor of above patients were extracted and constructed as libraries for each sample before subjected to sequencing by a panel contains 50 cancer-associated genes encompassing 29 kb by custom probe hybridization capture with average depth >40000, 7000, or 6300 folds respectively.\n\nResultsDetection limit for mutant allele frequency in our study was 0.1%. The sequencing results were analyzed by bioinformatic expertise based on our previous studies on the baseline mutation profiling of circulating cell-free DNA and the clinicopathological data of these patients. Among all the lung cancer patients, 78% patients were predicted as positive by ctDNA sequencing when the shreshold was defined as at least one of the hotspot mutations detected in the blood (ctDNA) was also detected in tumor tissue. Pneumonia and pulmonary tuberculosis were detected as negative according to the above standard. When evaluating all hotspots in driver genes in the panel, 24% mutations detected in tumor tissue (tDNA) were also detected in patients blood (ctDNA). When evaluating all genetic variations in the panel, including all the driver genes and passenger genes, 28% detected in tumor tissue (tDNA) were also detected in patients blood (ctDNA). Positive detection rates of plasma ctDNA in stage I lung cancer patients is 85%, compared with 17% of tumor biomarkers.\n\nConclusionWe demonstrated the importance of sequencing both circulating cell-free DNA and genomic DNA in tumor tissue for ctDNA detection in lung cancer currently. We also determined and confirmed the consistency of ctDNA and tumor tissue through NGS according to the criteria explored in our studies. Our strategy can initially distinguish the lung cancer from benign lesions of lung. Our work shows that the consistency will be benefited from the optimization in sensitivity and specificity in ctDNA detection.

cancer biology

plot2DO: a tool to assess the quality and distribution of genomic data

SummaryMicrococcal nuclease digestion followed by deep sequencing (MNase-seq) is the most used method to investigate nucleosome organization on a genome-wide scale. We present plot2DO, a software package for creating 2D occupancy plots, which allows biologists to evaluate the quality of MNase-seq data and to visualize the distribution of nucleosomes near the functional regions of the genome (e.g. gene promoters, origins of replication, etc.).\n\nAvailability And ImplementationThe plot2DO open source package is freely available on GitHub at https://github.com/rchereji/plot2DO under the MIT license.\n\nContactrazvan.chereji@nih.gov\n\nSupplementary InformationSupplementary data are available at Bioinformatics online.

genomics

Mutational burden of giant synaptic genes may be the cause of Alzheimer’s disease

All drug trials of the Alzheimers disease (AD) have failed to slow the progression of dementia in phase III studies, and the most effective therapeutic strategy remains controversial due to the poorly understood disease mechanisms. For AD drug design, amyloid beta (A{beta}) and its cascade have been the primary focus since decades ago, but mounting evidence indicates that the underpinning molecular pathways of AD are more complex than the classical reductionist models.\n\nSeveral genome-wide association studies (GWAS) have recently shed light on dark aspects of AD from a hypothesis-free perspective. Here, I use this novel insight to suggest that the amyloid cascade hypothesis may be a wrong model for AD therapeutic design. I review 23 novel genetic risk loci and show that, as a common theme, they code for receptor proteins and signal transducers of cell adhesion pathways, with clear implications in synaptic development, maintenance, and function. Contrary to the A{beta}-based interpretation, but further reinforcing the unbiased genome-wide insight, the classical hallmark genes of AD including the amyloid precursor protein (APP), presenilins (PSEN), and APOE also take part in similar pathways of growth cone adhesion and contact-guidance during brain development. On this basis, I propose that a disrupted synaptic adhesion signaling nexus, rather than a protein aggregation process, may be the central point of convergence in AD mechanisms. By an exploratory bioinformatics analysis, I show that synaptic adhesion proteins are encoded by largest known human genes, and these extremely large genes may be vulnerable to DNA damage accumulation in aging due to their mutational fragility. As a prototypic example and an immediately testable hypothesis based on this argument, I suggest that mutational instability of the large Lrp1b tumor suppressor gene may be the primary etiological trigger for APOE/dab1 signaling disruption in late-onset AD.\n\nIn conclusion, the large gene instability hypothesis suggests that evolutionary forces of brain complexity have led to emergence of large and fragile synaptic genes, and these unstable genes are the bottleneck etiology of aging disorders including senile dementias. A paradigm shift is warranted in AD prevention and therapeutic design.\n\nGlossary

neuroscience

Enabling Precision Medicine via standard communication of NGS provenance, analysis, and results

A personalized approach based on a patients or pathogens unique genomic sequence is the foundation of precision medicine. Genomic findings must be robust and reproducible, and experimental data capture should adhere to FAIR guiding principles. Moreover, effective precision medicine requires standardized reporting that extends beyond wet lab procedures to computational methods. The BioCompute framework (https://osf.io/zm97b/) enables standardized reporting of genomic sequence data provenance, including provenance domain, usability domain, execution domain, verification kit, and error domain. This framework facilitates communication and promotes interoperability. Bioinformatics computation instances that employ the BioCompute framework are easily relayed, repeated if needed and compared by scientists, regulators, test developers, and clinicians. Easing the burden of performing the aforementioned tasks greatly extends the range of practical application. Large clinical trials, precision medicine, and regulatory submissions require a set of agreed upon standards that ensures efficient communication and documentation of genomic analyses. The BioCompute paradigm and the resulting BioCompute Objects (BCO) offer that standard, and are freely accessible as a GitHub organization (https://github.com/biocompute-objects) following the \"Open-Stand.org principles for collaborative open standards development\". By communication of high-throughput sequencing studies using a BCO, regulatory agencies (e.g., FDA), diagnostic test developers, researchers, and clinicians can expand collaboration to drive innovation in precision medicine, potentially decreasing the time and cost associated with next generation sequencing workflow exchange, reporting, and regulatory reviews.

scientific communication and education

Desert Tortoises in the Genomic Age: Population Genetics and the Landscape

The California Department of Fish and Wildlife (CDFW) provided research funds to study the conservation genomics and landscape genomics of the Mojave desert tortoise, Gopherus agassizii, in response to the Desert Renewable Energy Conservation Plan (DRECP). To do this, we consolidated tissue samples of the desert tortoise from across the species range within California and southern Nevada, generated a DNA dataset consisting of full genomes of 270 tortoises, and analyzed the way in which the environment of the desert tortoise has determined modern patterns of relatedness and genetic diversity across the landscape. Here we present the implications of these results for the conservation and landscape genomics of the desert tortoise. Our work strongly indicates that several well-defined genetic groups exist within the species, including a primary north-south genetic discontinuity at the Ivanpah Valley and another separating western from eastern Mojave samples. We also use existing desert tortoise habitat modeling data with a novel extension of genetic \"resistance distance\" using geographic maps of continuous space to predict the relative impacts of five proposed development alternatives within the DRECP and rank them with respect to their likely impacts on desert tortoise gene flow and connectivity in the Mojave. Finally, we analyzed the impacts of each of the 214 distinct proposed development area \"chunks,\" derived from the proposed development polygons, and ranked each chunk in terms of its range-wide impacts on desert tortoise gene flow.\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=150 SRC=\"FIGDIR/small/195743_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (147K):\norg.highwire.dtl.DTLVardef@b60416org.highwire.dtl.DTLVardef@1c649f2org.highwire.dtl.DTLVardef@120e625org.highwire.dtl.DTLVardef@e59b34_HPS_FORMAT_FIGEXP M_FIG C_FIG PrefaceO_ST_ABSContextC_ST_ABSThe following document is a report that was submitted to the California Department of Fish and Wildlife, describing a series of analyses to help understand the impacts of several alternative spatial configurations of renewable energy development on gene flow of the federally threatened Mojave desert tortoise. These development alternatives were the centerpiece of the Desert Renewable Energy Conservation Plan (DRECP), a landscape-level land use planning initiative undertaken by the Bureau of Land Management (BLM), U.S. Fish and Wildlife Service (USFWS), California Energy Commission (CEC), and the California Department of Fish and Wildlife (CDFW). We were tasked by the California Department of Fish and Wildlife with providing a detailed analysis of these alternative plans on desert tortoise gene flow, and submitted the report for the public comment period for the initial implementation of the DRECP.\n\nFuture PlansO_ST_ABSCurrent state of landscape-level planning for the Mojave desert tortoiseC_ST_ABSThe five proposed land use configuration alternatives analyzed in the subsequent report include public and private lands spread across several counties in California. Shortly after the end of the DRECPs public comment period, the government agencies that developed the DRECP announced that they would be splitting its implementation into two phases: one that deals with land use decisions on BLM-controlled lands and one that deals with non-BLM areas (Sahagun 2015).\n\nPhase I of the DRECP was approved by the Bureau of Land Management on September 14, 2016 (U.S. Bureau of Land Management 2016). This phase includes land use planning decisions for BLM-administered lands. Specifically, 388,000 acres of public lands were designated as development focus areas (DFAs). In applications for leasing lands for renewable energy development, DFAs will not require the same degree of environmental evaluation prior to permitting, as theyve already been evaluated in the context of the DRECP. The application process for renewable energy development within DFAs will be streamlined to encourage development in these areas. Phase I also designated a total of 6,527,000 acres for natural resource conservation. This includes California Desert National Conservation Lands, Areas of Critical Environmental Concern, and Wildlife Allocations. A further 2,691,000 acres were designated for recreation under Phase I. Phase II of the DRECP is currently under development in conjunction with county-level governments to extend this landscape-level planning beyond BLM-administered lands.\n\nAuthor ContributionsThis was a collaborative report. Evan McCartney-Melstad performed the simulations of the low-coverage full genome approach (see Figure 2); conducted all of the laboratory work to generate the genome sequences; performed all of the bioinformatic analyses to bring the raw sequence data to the various stages required for different analyses; wrote the software to quickly estimate pairwise genetic relationships between individuals using read count data in low coverage sequence data (www.github.com/atcg/cPWP); performed some of the population genetic analyses; and wrote and edited several sections of the report.\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=122 SRC=\"FIGDIR/small/195743_fig2.gif\" ALT=\"Figure 2\">\nView larger version (16K):\norg.highwire.dtl.DTLVardef@30a49forg.highwire.dtl.DTLVardef@187c3d8org.highwire.dtl.DTLVardef@4abe98org.highwire.dtl.DTLVardef@126febe_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 2.C_FLOATNO Comparison of two different sequencing approaches in their ability to differentiate very slightly differentiated populations (Fst=0.001)\n\nC_FIG Peter Ralph (in collaboration with Gideon Bradburd and Erik Lundgren) invented and implemented the random walk-based gene flow model that we used to estimate reductions in gene flow due to development, and also developed the theory behind the read-based pairwise pi and genetic covariance estimation used here, in addition to writing and editing several sections of the report. Gideon Bradburd also performed some of the population genetic analyses and wrote and edited several sections of the report. Jannet Vu collected and curated the spatial environmental data and generated the maps that are included in the report (Figures 10, A10-A13), and also wrote Appendices I and IV. Bridgette Hagerty, Fran Sandmeier, Chava Weitzman, and C. Richard Tracy contributed approximately 1,000 desert tortoise blood samples that they collected (at great effort), in addition to knowledge of tortoise ecology and conservation, as well as the results of previous microsatellite-based genetic analyses and editing of the report. H. Bradley Shaffer wrote and edited several sections of the report, and is listed as the lead author for his role in conceiving of and obtaining funding support for the project.\n\nO_FIG O_LINKSMALLFIG WIDTH=154 HEIGHT=200 SRC=\"FIGDIR/small/195743_fig10.gif\" ALT=\"Figure 10\">\nView larger version (82K):\norg.highwire.dtl.DTLVardef@11e99acorg.highwire.dtl.DTLVardef@1faf2f5org.highwire.dtl.DTLVardef@64c440org.highwire.dtl.DTLVardef@1907837_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 10.C_FLOATNO Spatial configuration of the proposed development chunks (see Appendix 4).\n\nC_FIG

evolutionary biology

Understanding Functional Roles of Native Pentose-Specific Transporters for Activating Dormant Pentose Metabolism in Yarrowia lipolytica

Pentoses including xylose and arabinose are the second-most prevalent sugars of lignocellulosic biomass that can be harnessed for biological conversion. Although Yarrowia lipolytica has emerged as a promising industrial microorganism for production of high-value chemicals and biofuels, its native pentose metabolism is poorly understood. Our previous study demonstrated that Y. lipolytica (ATCC MYA-2613) has endogenous enzymes for D-xylose assimilation, but inefficient xylitol dehydrogenase causes Y. lipolytica to assimilate xylose poorly. In this study, we investigated the functional roles of native sugar-specific transporters for activating the dormant pentose metabolism in Y. lipolytica. By screening a comprehensive set of 16 putative pentose-specific transporters, we identified two candidates, YALI0C04730p and YALI0B00396p, that enhanced xylose assimilation. The engineered mutants YlSR207 and YlSR223, overexpressing YALI0C04730p and YALI0B00396p, respectively, improved xylose assimilation approximately 23% and 50% in comparison to YlSR102, a parent engineered strain overexpressing solely the native xylitol dehydrogenase gene. Further, we activated and elucidated a widely unknown, native L-arabinose-assimilating pathway in Y. lipolytica through transcriptomic and metabolic analyses. We discovered that Y. lipolytica can co-consume xylose and arabinose, where arabinose utilization shares transporters and metabolic enzymes of some intermediate steps of the xylose-assimilating pathway. Arabinose assimilation was synergistically enhanced in the presence of xylose while xylose assimilation was competitively inhibited by arabinose. L-arabitol dehydrogenase is the rate-limiting step responsible for poor arabinose utilization in Y. lipolytica. Overall, this study sheds light on the cryptic pentose metabolism of Y. lipolytica and further helps guide strain engineering of Y. lipolytica for enhanced assimilation of pentose sugars.\n\nIMPORTANCEThe oleaginous yeast Yarrowia lipolytica is a promising industrial platform microorganism for production of high-value chemicals and fuels. For decades since its isolation, Y. lipolytica has often been known to be incapable of assimilating pentose sugars, xylose and arabinose, that are dominantly present in lignocellulosic biomass. Through bioinformatic, transcriptomic and enzymatic studies, we have uncovered the dormant pentose metabolism of Y. lipolytica. Remarkably, unlike most yeast strains that share the same transporters for importing hexose and pentose sugars, we discovered that Y. lipolytica possess the native pentose-specific transporters. By overexpressing these transporters together with the rate-limiting D-xylitol and L-arabitol dehydrogenases, we activated the dormant pentose metabolism of Y. lipolytica. Overall, this study provides a fundamental understanding of the dormant pentose metabolism of Y. lipolytica and guides future metabolic engineering of Y. lipolytica for enhanced conversion of pentose sugars to high-value chemicals and fuels.

physiology

Ixodes scapularis does not harbor a stable midgut microbiome

Hard ticks of the order Ixodidae serve as vectors for numerous human pathogens, including the causative agent of Lyme Disease Borrelia burgdorferi. Tick-associated microbes can influence pathogen colonization, offering the potential to inhibit disease transmission through engineering of the tick microbiota. Here, we investigate whether B. burgdorferi encounters abundant bacteria within the midgut of wild adult Ixodes scapularis, its primary vector. Through the use of controlled sequencing methods and confocal microscopy, we find that the majority of field-collected adult I. scapularis harbor limited internal microbial communities that are dominated by endosymbionts. A minority of I. scapularis ticks harbor abundant midgut bacteria and lack B. burgdorferi. We find that the lack of a stable resident midgut microbiota is not restricted to I. scapularis since extension of our studies to I. pacificus, Amblyomma maculatum, and Dermacentor spp showed similar patterns. Finally, bioinformatic examination of the B. burgdorferi genome revealed the absence of genes encoding known interbacterial interaction pathways, a feature unique to the Borrelia genus within the phylum Spirochaetes. Our results suggest that reduced selective pressure from limited microbial populations within ticks may have facilitated the evolutionary loss of genes encoding interbacterial competition pathways from Borrelia.

microbiology

Phylostratigraphic analysis of tumor and developmental transcriptomes reveals relationship between oncogenesis, phylogenesis and ontogenesis

The question of the existence of cancer is inadequately answered by invoking somatic mutations or the disruptions of cellular and tissue control mechanisms. As such uniformly random events alone cannot account for the almost inevitable occurrence of an extremely complex process such as cancer. In the different epistemic realm, an ultimate explanation of cancer is that cancer is a reversion of a cell to an ancestral pre-Metazoan state, i.e. a cellular form of atavism. Several studies have suggested that genes involved in cancer have evolved at particular evolutionary time linked to the unicellular-multicellular transition. Here we used a refined phylostratigraphic analysis of evolutionary ages of the known genes/pathways associated with cancer and the genes differentially expressed between normal and cancer tissue as well as between embryonic and mature (differentiated) cells. We found that cancer-specific transcriptomes and cancer-related pathways were enriched for genes that evolved in the pre-Metazoan era and depleted of genes that evolved in the post-Metazoan era. By contrast an opposite relation was found for cell maturation: the age distribution frequency of the genes expressed in differentiated epithelial cells were enriched for post-Metazoan genes and depleted of pre-Metazoan ones. These findings support the atavism theory that cancer cells manifest the reactivation of an ancient ancestral state featuring unicellular modalities. Thus our bioinformatics analyses suggest that not only does oncogenesis recapitulate ontogenesis, and ontogenesis recapitulates phylogenesis, but also oncogenesis recapitulates phylogenesis. This more encompassing perspective may offer a natural organizing framework for genetic alterations in cancers and point to new treatment options that target the genes controlling the atavism transition.\n\nOne Sentence SummaryTracing cancer gene evolutionary ages revealed that cancer reverts to a pre-existing early Metazoan state.

cancer biology

Short linear motifs in intrinsically disordered regions modulate HOG signaling capacity

The effort to characterize intrinsically disordered regions of signaling proteins is rapidly expanding. An important class of disordered interaction modules are ubiquitous and functionally diverse elements known as short linear motifs (SLiMs). To further examine the role of SLiMs in signal transduction, we used a previously devised bioinformatics method to predict evolutionarily conserved SLiMs within a well-characterized pathway in S. cerevisiae. Using a single cell, reporter-based flow cytometry assay in conjunction with a fluorescent reporter driven by a pathway-specific promoter, we quantitatively assessed pathway output via systematic deletions of individual motifs. We found that, when deleted, 34% (10/29) of predicted SLiMs displayed a significant decrease in pathway output, providing evidence that these motifs play a role in signal transduction. In addition, we show that perturbations of parameters in a previously published stochastic model of HOG signaling could reproduce the quantitative effects of 4 out of 7 mutations in previously unknown SLiMs. Our study suggests that, even in well-characterized pathways, large numbers of functional elements remain undiscovered, and that challenges remain for application of systems biology models to interpret the effects of mutations in signalling pathways.\n\nOne-sentence SummaryMutations of short conserved elements in disordered regions have quantitative effects on a model signaling pathway.

systems biology

Multivariate genome-wide association study of rapid automatized naming and rapid alternating stimulus in Hispanic and African American youth.

Reading disability is a complex neurodevelopmental disorder that is characterized by difficulties in reading despite educational opportunity and normal intelligence. Performance on rapid automatized naming (RAN) and rapid alternating stimulus (RAS) tests gives a reliable predictor of reading outcome. These tasks involve the integration of different neural and cognitive processes required in a mature reading brain. Most studies examining the genetic factors that contribute to RAN and RAS performance have focused on pedigree-based analyses in samples of European descent, with limited representation of groups with Hispanic or African ancestry. In the present study, we conducted a multivariate genome-wide association analysis to identify shared genetic factors that contribute to performance across RAN Objects, RAN Letters, and RAS Letters/Numbers in a sample of Hispanic and African American youth (n=1,331). We then tested whether these factors also contribute to variance in reading fluency and word reading. Genome-wide significant, pleiotropic, effects across RAN Objects, RAN Letters, and RAS Letters/Numbers were observed for SNPs located on chromosome 10q23.31 (rs1555839, multivariate association, p=2.23 x 10-8), which also showed significant association with reading fluency and word reading performance (p <0.001). Bioinformatic analysis of this region using epigenetic data from the NIH Roadmap Epigenomics Mapping Consortium indicates active transcription of the gene RNLS in the brain. Neuroimaging genetic analysis of fourteen cortical regions in an independent sample of typically developing children across multiple ethnicities (n=690) showed that rs1555839 was associated with variation in volume of the right inferior parietal cortex--a region of the brain that processes numerical information and has been implicated in reading disability. This study provides support for a novel locus on chromosome 10q23.31 associated with RAN, RAS, and reading-related performance.\n\nAUTHOR SUMMARYReading disability has a strong genetic component that is explained by multiple genes and genetic factors. The complex genetic architecture along with diverse cognitive impairments associated with reading disability, poses challenges in identifying novel genes and variants that confer risk. One method to begin parsing genetic and neurobiological mechanisms that contribute to reading disability is to take advantage of the high correlation among reading-related cognitive traits like rapid automatized naming (RAN) and rapid alternating stimulus (RAS) to identify shared genetic factors that contribute to common biological mechanisms. In the present study, we used a multivariate genome-wide analysis approach that identified a region of chromosome 10q23.31 associated with variation in RAN Objects, RAN Letters, and RAS Letters/Numbers performance in a sample of 1,331 Hispanic and African American youth in the Genes, Reading, and Dyslexia (GRaD) Study. Genetic variants in this region were also associated with reading fluency in GRaD, and differences in brain structures implicated in reading disability in a separate sample of 690 children. The gene, RNLS, is located within the implicated region of chromosome 10q23.31 and plays a role in breaking down a class of chemical messengers known to affect attention, learning, and memory in the brain. These findings provide a basis to inform our understanding of the biological basis of reading disability.

genetics

A most wanted list of conserved protein families with no known domains

The number and proportion of genes with no known function are growing rapidly. To quantify this phenomenon and provide criteria for prioritizing genes for functional characterization, we developed a bioinformatics pipeline that identifies robustly defined protein families with no annotated domains, ranks these with respect to phylogenetic breadth, and identifies them in metagenomics data. We applied this approach to 271 965 protein families from the SFams database and discovered many with no functional annotation, including >118 000 families lacking any known protein domain. From these, we prioritized 6 668 conserved protein families with at least three sequences from organisms in at least two distinct classes. These Function Unknown Families (FUnkFams) are present in Tara Oceans Expedition and Human Microbiome Project metagenomes, with distributions associated with sampling environment. Our findings highlight the extent of functional novelty in sequence databases and establish an approach for creating a \"most wanted\" list of genes to characterize.

genomics

Two C++ Libraries for Counting Trees on a Phylogenetic Terrace

MotivationThe presence of terraces in phylogenetic tree space, that is, a potentially large number of distinct tree topologies that have exactly the same analytical likelihood score, was first described by Sanderson et al, (2011). However, popular software tools for maximum likelihood and Bayesian phylogenetic inference do not yet routinely report, if inferred phylogenies reside on a terrace, or not. We believe, this is due to the unavailability of an efficient library implementation to (i) determine if a tree resides on a terrace, (ii) calculate how many trees reside on a terrace, and (iii) enumerate all trees on a terrace.\n\nResultsIn our bioinformatics programming practical we developed two efficient and independent C++ implementations of the SUPERB algorithm by Constantinescu and Sankoff (1995) for counting and enumerating the trees on a terrace. Both implementations yield exactly the same results and are more than one order of magnitude faster and require one order of magnitude less memory than a previous 3rd party python implementation.\n\nAvailabilityThe source codes are available under GNU GPL at https://github.com/terraphast\n\nContactAlexandros.Stamatakis@h-its.org

evolutionary biology

Mmp10 is required for post-translational methylation of arginine at the active site of methyl-coenzyme M reductase

Catalyzing the key step for anaerobic methane production and oxidation, methyl-coenzyme M reductase or Mcr plays a key role in the global methane cycle. The McrA subunit possesses up to five post-translational modifications (PTM) at its active site. Bioinformatic analyses had previously suggested that methanogenesis marker protein 10 (Mmp10) could play an important role in methanogenesis. To examine its role, MMP1554, the gene encoding Mmp10 in Methanococcus maripaludis, was deleted with a new genetic tool, resulting in the specific loss of the 5-(S)-methylarginine PTM of residue 275 in the McrA subunit and a 40~60 % reduction in the maximal rates of methane formation by whole cells. Methylation was restored by complementations with the wild-type gene. However, the rates of methane formation of the complemented strains were not always restored to the wild type level. This study demonstrates the importance of Mmp10 and the methyl-Arg PTM on Mcr activity.

microbiology

Impact of sequence variant detection and bacterial DNA extraction methods on the measurement of microbial community composition in human stool

BackgroundThe human gut microbiome has been widely studied in the context of human health and metabolism, however the question of how to analyze this community remains contentious. This study compares new and previously well established methods aimed at reducing bias in bioinformatics analysis (QIIME 1 and DADA2) and bacterial DNA extraction of human fecal samples in 16S rRNA marker gene surveys.\n\nResultsAnalysis of a mock DNA community using DADA2 identified more chimeras (QIIME 1: 0.70% of total reads vs DADA2: 1.96%), fewer sequence variants, (QIIME 1: 1297.4 + 98.88 vs. DADA2: 136.27 + 11.35, mean + SD) and correct taxa at a higher resolution of classification (i.e. genus-level) than open reference OTU picking in QIIME 1. Additionally, the extraction of whole cell mock community bacterial DNA using four commercially available kits resulted in varying DNA yield, quality and bacterial community composition. Of the four kits compared, ZymoBIOMICS DNA Miniprep Kit provided the greatest yield, with a slight enrichment of Enterococcus. However, QIAamp Fast DNA Stool Mini Kit resulted in the highest DNA quality. Mo Bio PowerFecal DNA Kit had the most dramatic effect on the mock community composition, resulting in an increased proportion of members of the family Enterobacteriaceae and genus Eshcerichia as well as members of genera Lactobacillus and Pseudomonas. The presence of a sterile fecal matrix had a slight, but inconsistent effect on the yield, quality and taxa identified after extraction with all four DNA extraction kits. Extraction of bacterial DNA from native stool samples revealed a distinct effect of the DNA stabilization reagent DNA/RNA Shield on community composition, causing an increase in the detected abundance of members of orders Bifidobacteriales, Bacteroidales, Turicibacterales, Clostridiales and Enterobacteriales.\n\nConclusionThese results confirm that the DADA2 algorithm is superior to sequence clustering by similarity to determine microbial community structure. Additionally, commercially available kits used for bacterial DNA extraction from fecal samples have some effect on the proportion of high abundance members detected in a microbial community, but it is less significant than the effect of using DNA stabilization reagent, DNA/RNA Shield.

molecular biology

OMGene: Mutual improvement of gene models through optimisation of evolutionary conservation

BackgroundThe accurate determination of the genomic coordinates for a given gene - its gene model - is of vital importance to the utility of its annotation, and the accuracy of bioinformatic analyses derived from it. Currently-available methods of computational gene prediction, while on the whole successful, often disagree on the model for a given predicted gene, with some or all of the variant gene models failing to match the biologically observed structure. Many prediction methods can be bolstered by using experimental data such as RNA-seq and mass spectrometry. However, these resources are not always available, and rarely give a comprehensive portrait of an organisms transcriptome due to temporal and tissue-specific expression profiles.\n\nResultsOrthology between genes provides evolutionary evidence to guide the construction of gene models. OMGene (Optimise My Gene) aims to optimise gene models in the absence of experimental data by optimising the derived amino acid alignments for gene models within orthogroups. Using RNA-seq data sets from plants and fungi, considering intron/exon junction representation and exon coverage, and assessing the intra-orthogroup consistency of subcellular localisation predictions, we demonstrate the utility of OMGene for improving gene models in annotated genomes.\n\nConclusionsWe show that significant improvements in the accuracy of gene model annotations can be made in both established and de novo annotated genomes by leveraging information from multiple species.

genomics