Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Global analysis of the RpaB regulon based on the positional distribution of HLR1 sequences and comparative differential RNA-Seq data

The transcription factor RpaB regulates the expression of genes encoding photosynthesis-associated proteins during light acclimation. The binding site of RpaB is the HLR1 motif, a pair of imperfect octameric direct repeats, separated by two random nucleotides. Here, we used high-resolution mapping data of transcriptional start sites (TSSs) in the model Synechocystis sp. PCC 6803 in conjunction with the positional distribution of HLR1 sites for the global prediction of the RpaB regulon. The results demonstrate that RpaB regulates the expression of more than 150 promoters, driving the transcription of protein-coding and non-coding genes and antisense transcripts under low light and upon the shift to high light when DNA binding activity is lost. Transcriptional activation by RpaB is achieved when the HLR1 motif is located 66 to 45 nt upstream, repression occurs when it is close to or overlapping the TSS. Selected examples were validated by multiple experimental approaches, including chromatin affinity purification, reporter gene, northern hybridization and electrophoretic mobility shift assays. We found that RpaB controls ssr2016/pgr5, which is involved in cyclic electron flow and state transitions; six out of nine ferredoxins; three of four FtsH proteases; gcvP/slr0293, encoding a crucial photorespiratory protein; and nirA and isiA for which we suggest cross-regulation with the transcription factors NtcA or FurA, respectively. In addition to photosynthetic gene functions, RpaB contributes to the control of genes affiliated with nitrogen assimilation, cofactor biosyntheses, the CRISPR system and the circadian clock, making it one of the most versatile regulators in cyanobacteria.\n\nSignificance StatementRpaB is a transcription factor in cyanobacteria and in the chloroplasts of several lineages of eukaryotic algae. Like other important transcription factors, the gene encoding RpaB cannot be deleted, making the study of deletion mutants impossible. Based on a bioinformatic approach, we increased the number of known genes controlled by RpaB by a factor of 5. Depending on the distance to the TSS, RpaB mediates transcriptional activation or repression. The high number and functional diversity among its target genes and co-regulation with other transcriptional regulators characterize RpaB as a regulatory hub.

microbiology

PP4-dependent HDAC3 dephosphorylation discriminates between axonal regeneration and regenerative failure

The molecular mechanisms discriminating between regenerative failure and success remain elusive. While a regeneration-competent peripheral nerve injury mounts a regenerative gene expression response in bipolar dorsal root ganglia (DRG) sensory neurons, a regeneration-incompetent central spinal cord injury does not. This dichotomic response offers a unique opportunity to investigate the fundamental biological mechanisms underpinning regenerative ability. Following a pharmacological screen with small molecule inhibitors targeting key epigenetic enzymes in DRG neurons we identified HDAC3 signalling as a novel candidate brake to axonal regenerative growth. In vivo, we determined that only a regenerative peripheral but not a central spinal injury induces an increase in calcium, which activates protein phosphatase 4 that in turn dephosphorylates HDAC3 thus impairing its activity and enhancing histone acetylation. Bioinformatics analysis of ex vivo H3K9ac ChIPseq and RNAseq from DRG followed by promoter acetylation and protein expression studies implicated HDAC3 in the regulation of multiple regenerative pathways. Finally, genetic or pharmacological HDAC3 inhibition overcame regenerative failure of sensory axons following spinal cord injury. Together, these data indicate that PP4-dependent HDAC3 dephosphorylation discriminates between axonal regeneration and regenerative failure.\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=122 SRC=\"FIGDIR/small/446963_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (40K):\norg.highwire.dtl.DTLVardef@1bd577corg.highwire.dtl.DTLVardef@1ba990aorg.highwire.dtl.DTLVardef@195813corg.highwire.dtl.DTLVardef@57abc7_HPS_FORMAT_FIGEXP M_FIG Graphical AbstractFollowing central nervous system (CNS) spinal injury, protein phosphatase 4/2 activity is not induced since calcium levels remain unchanged compared to uninjured conditions. HDAC3 remains phosphorylated and occupies deacetylated chromatin contributing to its compaction inhibiting gene expression. Following peripheral nervous system (PNS) sciatic injury, protein phosphatase 4/2 activity is induced by calcium. HDAC3 is dephosphorylated leading to its inhibition and release from chromatin sites contributing to increase in histone acetylation and in the expression of regeneration associated genes (RAGs).\n\nC_FIG

neuroscience

MicroRNA-200c suppresses epithelial-mesenchymal transition of ovarian cancer by targeting cofilin-2

This study investigated the effects of microRNA-200c (miR-200c) and cofilin-2 (CFL2) in regulating epithelial-mesenchymal transition (EMT) in ovarian cancer. The level of miR-200c was lower in invasive SKOV3 cells than that in non-invasive OVCAR3 cells, whereas CFL2 showed the opposite trend. Bioinformatics analysis and dual-luciferase reporter gene assays indicated that CFL2 was a direct target of miR-200c. Furthermore, SKOV3 and OVCAR3 cells were transfected with miR-200c mimic or inhibitor, pCDH-CFL2 (CFL2 overexpression), or CFL2 shRNA (CFL2 silencing). MiR-200c inhibition and CFL2 overexpression resulted in elevated levels of both CFL2 and vimentin while reducing E-cadherin expression. They also increased ovarian cancer cell invasion and migration in vitro and in vivo and increased the tumor volumes. Conversely, miR-200c mimic and CFL2 shRNA exerted the opposite effects as those aforementioned. In addition, the effects of pCDH-CFL2 and CFL2 shRNA were reversed by the miR-200c mimic and inhibitor, respectively. This finding suggested that miR-200c could be a potential tumor suppressor by targeting CFL2 in the EMT process.

cancer biology

RERconverge: an R package for associating evolutionary rates with convergent traits

Motivation: When different lineages of organisms independently adapt to similar environments, selection often acts repeatedly upon the same genes, leading to signatures of convergent evolutionary rate shifts at these genes. With the increasing availability of genome sequences for organisms displaying a variety of convergent traits, the ability to identify genes with such convergent rate signatures would enable new insights into the molecular basis of these traits.\n\nResults: Here we present the R package RERconverge, which tests for association between relative evolutionary rates of genes and the evolution of traits across a phylogeny. RERconverge can perform associations with binary and continuous traits, and it contains tools for visualization and enrichment analyses of association results.\n\nAvailability: RERconverge source code, documentation, and a detailed usage walk-through are freely available at https://github.com/nclark-lab/RERconverge. Datasets for mammals, Drosophila, and yeast are available at https://bit.ly/2J2QBnj.\n\nContact: mchikina@pitt.edu\n\nSupplementary information: Supplementary information, containing detailed vignettes for usage of RERconverge, are available at Bioinformatics online.

evolutionary biology

Metagenomic characterization of the viral community of the South Scotia Ridge

Viruses are the most abundant biological entities in aquatic ecosystems and harbor an enormous genetic diversity. While their great influence on the marine ecosystems is widely acknowledged, current information about their diversity remains scarce. Aviral metagenomic analysis of two surfaces and one bottom water sample was conducted from sites on the South Scotia Ridge (SSR) near the Antarctic Peninsula, during the austral summer 2016. The taxonomic composition and diversity of the viral communities were investigated and a functional assessment of the sequences was determined. Phylotypic analysis showed that most viruses belonging to the order Caudovirales, in particular, the family Podoviridae (41.92-48.7%), which is similar to the viral communities from the Pacific Ocean. Functional analysis revealed a relatively high frequency of phage-associated and metabolism genes. Phylogenetic analyses of phage TerL and Capsid_NCLDV (nucleocytoplasmic large DNA viruses) marker genes indicated that many of the sequences associated with Caudovirales and NCLDV were novel and distinct from known complete phage genomes. High Phaeocystis globosa virus virophage (Pgvv) signatures were found in SSR area and complete and partial Pgvv-like were obtained which may have an influence on host-virus interactions in the area during summer. Our study expands the existing knowledge of viral communities and their diversities from the Antarctic region and provides basic data for further exploring polar microbiomes.\n\nImportanceIn this study, we used high-throughput sequencing and bioinformatics analysis to analyze the viral community structure and biodiversity of SSR in the open sea near the Antarctic Peninsula. The results showed that the SSR viromes are novel, oceanic-related viromes and a high proportion of sequence reads was classified as unknown. Among known virus counterparts, members of the order Caudovirales were most abundant which is consistent with viromes from the Pacific Ocean. In addition, phylogenetic analyses based on the viral marker genes (TerL and MCP) illustrate the high diversity among Caudovirales and NCLDV. Combining deep sequencing and a random subsampling assembly approach, a new Pgvv-like group was also found in this region, which may a signification factor regulating virus-host interactions.

microbiology

Differential gene expression, including Sjfs800, in Schistosoma japonicum females before, during, and after male-female pairing

Schistosomiasis is a prevalent but neglected tropical disease caused by parasitic trematodes of the genus Schistosoma, with the primary disease-causing species being S. haematobium, S. mansoni, and S. japonicum. Male-female pairing of schistosomes is necessary for sexual maturity and the production of a large number of eggs, which are primarily responsible for schistosomiasis dissemination and pathology. Here, we used microarray hybridization, bioinformatics, quantitative PCR, in situ hybridization, and gene silencing assays to identify genes that play critical roles in S. japonicum reproduction biology, particularly in vitellarium development, a process that affects male-female pairing, sexual maturation, and subsequent egg production. Microarray hybridization analyses generated a comprehensive set of genes differentially transcribed before and after male-female pairing. Although the transcript profiles of females were similar 16 and 18 days after host infection, marked gene expression changes were observed at 24 days. The 30 most abundantly transcribed genes on day 24 included those associated with vitellarium development. Among these, genes for female-specific 800 (fs800), eggshell precursor protein, and superoxide dismutase (cu-zn-SOD) were substantially upregulated. Our in situ hybridization results in female S. japonicum indicated that cu-zn-SOD mRNA was highest in the ovary and vitellarium, eggshell precursor protein mRNA was expressed in the ovary, ootype, and vitellarium, and Sjfs800 mRNA was observed only in the vitellarium, localized in mature vitelline cells. Knocking down the Sjfs800 gene in female S. japonicum by approximately 60% reduced the number of mature vitelline cells, decreased rates of pairing and oviposition, and decreased the number of eggs produced in each male-female pairing by about 50%. These results indicate that Sjfs800 is essential for vitellarium development and egg production in S. japonicum and suggest that Sjfs800 regulation may provide a novel approach for the prevention or treatment of schistosomiasis.\n\nAuthor SummarySchistosomiasis is a common but largely unstudied tropical disease caused by parasitic trematodes of the genus Schistosoma. The eggs of schistosomes are responsible for schistosomiasis transmission and pathology, and the production of these eggs is dependent on the pairing of females and males. In this study, we determined which genes in Schistosoma japonicum females were differentially expressed before and after pairing with males, identifying the 30 most abundantly expressed of these genes. Among these 30 genes, we further characterized those in female S. japonicum that were upregulated after pairing and that were related to reproduction and vitellarium development, a process that affects male-female pairing, sexual maturation, and subsequent egg production. We identified three such genes, S. japonicum female-specific 800 (Sjfs800), eggshell precursor protein, and superoxide dismutase, and confirmed that the mRNAs for these genes were primarily localized in reproductive structures. By using gene silencing techniques to reduce the amount of Sjfs800 mRNA in females by about 60%, we determined that Sjfs800 plays a key role in development of the vitellarium and egg production. This finding suggests that regulation of Sjfs800 may provide a novel approach to reduce egg counts and thus aid in the prevention or treatment of schistosomiasis.

genomics

Functional Annotation Signatures of Disease Susceptibility Loci Improve SNP Association Analysis

We describe the development and application of a Bayesian statistical model for the prior probability of phenotype-genotype association that incorporates data from past association studies and publicly available functional annotation data regarding the susceptibility variants under study. The model takes the form of a binary regression of association status on a set of annotation variables whose coefficients were estimated through an analysis of associated SNPs housed in the GWAS Catalog (GC). The set of functional predictors we examined includes measures that have been demonstrated to correlate with the association status of SNPs in the GC and some whose utility in this regard is speculative: summaries of the UCSC Human Genome Browser ENCODE super-track data, dbSNP function class, sequence conservation summaries, proximity to genomic variants included in the Database of Genomic Variants (DGV) and known regulatory elements included in the Open Regulatory Annotation database (ORegAnno), PolyPhen-2 probabilities and RegulomeDB categories. Because we expected that only a fraction of the annotation variables would contribute to predicting association, we employed a penalized likelihood method to reduce the impact of non-informative predictors and evaluated the models ability to predict GC SNPs not used to construct the model. We show that the functional data alone are predictive of a SNPs presence in the GC. Further, using data from a genome-wide study of ovarian cancer, we demonstrate that their use as prior data when testing for association is practical at the genome-wide scale and improves power to detect associations.

Bioinformatics

On the optimal trimming of high-throughput mRNAseq data

The widespread and rapid adoption of high-throughput sequencing technologies has afforded researchers the opportunity to gain a deep understanding of genome level processes that underlie evolutionary change, and perhaps more importantly, the links between genotype and phenotype. In particular, researchers interested in functional biology and adaptation have used these technologies to sequence mRNA transcriptomes of specific tissues, which in turn are often compared to other tissues, or other individuals with different phenotypes. While these techniques are extremely powerful, careful attention to data quality is required. In particular, because high-throughput sequencing is more error-prone than traditional Sanger sequencing, quality trimming of sequence reads should be an important step in all data processing pipelines. While several software packages for quality trimming exist, no general guidelines for the specifics of trimming have been developed. Here, using empirically derived sequence data, I provide general recommendations regarding the optimal strength of trimming, specifically in mRNA-Seq studies. Although very aggressive quality trimming is common, this study suggests that a more gentle trimming, specifically of those nucleotides whose PO_SCPLOWHREDC_SCPLOW score <2 or <5, is optimal for most studies across a wide variety of metrics.

Bioinformatics

Gappy TotalReCaller for RNASeq Base-Calling and Mapping

Understanding complex mammalian biology depends crucially on our ability to define a precise map of all the transcripts encoded in a genome, and to measure their relative abundances. A promising assay depends on RNASeq approaches, which builds on next generation sequencing pipelines capable of interrogating cDNAs extracted from a cell. The underlying pipeline starts with base-calling, collect the sequence reads and interpret the raw-read in terms of transcripts that are grouped with respect to different splice-variant isoforms of a messenger RNA. We address a very basic problem involved in all of these pipeline, namely accurate Bayesian base-calling, which could combine the analog intensity data with suitable underlying priors on base-composition in the transcripts. In the context of sequencing genomic DNA, a powerful approach for base-calling has been developed in the TotalReCaller pipeline. For these purposes, it uses a suitable reference whole-genome sequence in a compressed self-indexed format to derive its priors. However, TotalReCaller faces many new challenges in the transcriptomic domain, especially since we still lack a fully annotated library of all possible transcripts, and hence a sufficiently good prior. There are many possible solutions, similar to the ones developed for TotalReCaller, in applications addressing de novo sequencing and assembly, where partial contigs or string-graphs could be used to boot-strap the Bayesian priors on base-composition. A similar approach would be applicable here too, partial assembly of transcripts can be used to characterize the splicing junctions or organize them in incompatibility graphs and then used as priors for TotalReCaller. The key algorithmic techniques for this purpose have been addressed in a forthcoming paper on Stringomics. Here, we address a related but fundamental problem, by assuming that we only have a reference genome, with certain intervals marked as candidate regions for ORF (Open Reading Frames), but not necessarily complete annotations regarding the 5 or 3 termini of a gene or its exon-intron structure. The algorithms we describe find the most accurate base-calls of a cDNA with the best possible segmentation, all mapped to the genome appropriately.

Bioinformatics

Unexpected links reflect the noise in networks

Gene regulatory networks are commonly used for modeling biological processes and revealing underlying molecular mechanisms. The reconstruction of gene regulatory networks from observational data is a challenging task, especially, considering the large number of involved players (e.g. genes) and much fewer biological replicates available for analysis. Herein, we proposed a new statistical method of estimating the number of erroneous edges that strongly enhances the commonly used inference approaches. This method is based on special relationship between correlation and causality, and allows to identify and to remove approximately half of erroneous edges. Using the mathematical model of Bayesian networks and positive correlation inequalities we established a mathematical foundation for our method. Analyzing real biological datasets, we found a strong correlation between the results of our method and the commonly used false discovery rate (FDR) technique. Furthermore, the simulation analysis demonstrates that in large networks, our new method provides a more precise estimation of the proportion of erroneous links than FDR.

Bioinformatics

Comment on “TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions” by Kim et al.

In the recent paper [1] (thereafter referred to as \"TopHat2paper\") the accuracy of TopHat2 was compared to other RNA-seq aligners. In this comment we re-examine most important analyses from the TopHat2paper and identify several deficiencies that significantly diminished performance of some of the aligners, including incorrect choice of mapping parameters, unfair comparison metrics, and unrealistic simulated data. Using STAR [2] as an exemplar, we demonstrate that correcting these deficiencies makes its accuracy equal or better than that of TopHat2. Furthermore, this exercise highlighted some serious issues with the TopHat2 algorithms, such as poor recall of alignments with a moderate (>3) number of mismatches, low sensitivity and high false discovery rate for splice junction detection, loss of precision for the realignment algorithm, and large number of false chimeric alignments. ...

Bioinformatics

Inferring tree causal models of cancer progression with probability raising

Existing techniques to reconstruct tree models of progression for accumulative processes, such as cancer, seek to estimate causation by combining correlation and a frequentist notion of temporal priority. In this paper, we define a novel theoretical framework called CAPRESE (CAncer PRogression Extraction with Single Edges) to reconstruct such models based on the notion of probabilistic causation defined by Suppes. We consider a general reconstruction setting complicated by the presence of noise in the data due to biological variation, as well as experimental or measurement errors. To improve tolerance to noise we define and use a shrinkage-like estimator. We prove the correctness of our algorithm by showing asymptotic convergence to the correct tree under mild constraints on the level of noise. Moreover, on synthetic data, we show that our approach outperforms the state-of-the-art, that it is efficient even with a relatively small number of samples and that its performance quickly converges to its asymptote as the number of samples increases. For real cancer datasets obtained with different technologies, we highlight biologically significant differences in the progressions inferred with respect to other competing techniques and we also show how to validate conjectured biological relations with progression models.

Bioinformatics

A Bayesian Method to Incorporate Hundreds of Functional Characteristics with Association Evidence to Improve Variant Prioritization

The increasing quantity and quality of functional genomic information motivate the assessment and integration of these data with association data, including data originating from genome-wide association studies (GWAS). We used previously described GWAS signals (\"hits\") to train a regularized logistic model in order to predict SNP causality on the basis of a large multivariate functional dataset. We show how this model can be used to derive Bayes factors for integrating functional and association data into a combined Bayesian analysis. Functional characteristics were obtained from the Encyclopedia of DNA Elements (ENCODE), from published expression quantitative trait loci (eQTL), and from other sources of genome-wide characteristics. We trained the model using all GWAS signals combined, and also using phenotype specific signals for autoimmune, brain-related, cancer, and cardiovascular disorders. The non-phenotype specific and the autoimmune GWAS signals gave the most reliable results. We found SNPs with higher probabilities of causality from functional characteristics showed an enrichment of more significant p-values compared to all GWAS SNPs in three large GWAS studies of complex traits. We investigated the ability of our Bayesian method to improve the identification of true causal signals in a psoriasis GWAS dataset and found that combining functional data with association data improves the ability to prioritise novel hits. We used the predictions from the penalized logistic regression model to calculate Bayes factors relating to functional characteristics and supply these online alongside resources to integrate these data with association data.\n\nAuthor SummaryLarge-scale genetic studies have had success identifying genes that play a role in complex traits. Advanced statistical procedures suggest that there are still genetic variants to be discovered, but these variants are difficult to detect. Incorporating biological information that affect the amount of protein or other product produced can be used to prioritise the genetic variants in order to identify which are likely to be causal. The method proposed here uses such biological characteristics to predict which genetic variants are most likely to be causal for complex traits.

Bioinformatics

Exploring DNA structures in real-time polymerase kinetics using Pacific Biosciences sequencer data

Pausing of DNA polymerase can indicate the presence of a DNA structure that differs from the canonical double-helix. Here we detail a method to investigate how polymerase pausing in the Pacific Biosciences sequencer reads can be related to DNA structure. The Pacific Biosciences sequencer uses optics to view a polymerase and its interaction with a single DNA molecule in real-time, offering a unique way to detect potential alternative DNA structures. We have developed a new way to examine polymerase kinetics and relate it to the DNA sequence by using a wavelet transform of read information from the sequencer. We use this method to examine how polymerase kinetics are related to nucleotide base composition. We then examine tandem repeat sequences known for their ability to form different DNA structures: (CGG)n and (CG)n repeats which can, respectively, form G-quadruplex DNA and Z-DNA. We find pausing around the (CGG)n repeat that may indicate the presence of G-quadruplexes in some of the sequencer reads. The (CG)n repeat does not appear to cause polymerase pausing, but its kinetics signature nevertheless suggests the possibility that alternative nucleotide conformations may sometimes be present. We discuss the implications of using our method to discover DNA sequences capable of forming alternative structures. The analyses presented here can be reproduced on any Pacific Biosciences kinetics data for any DNA pattern of interest using an R package that we have made publicly available.\n\nAuthor SummaryDNA can be found in various forms that differ from the double-helix first discovered by Watson and Crick in 1953. These alternative DNA structures depend on the DNA sequence, and researchers continue to explore which sequences have the potential to form alternative structures. Here we advance the use of Pacific Biosciences sequencer data to explore potential alternative DNA structures. The Pacific Bio-sciences sequencer provides an unprecedented way to examine the interaction between DNA polymerase and DNA by following a single polymerase in real time as it copies a DNA molecule. The pausing of DNA polymerase is a common method for exploring the DNA sequences that have the potential to form alternative DNA structures, and Pacific Biosciences data has previously been used to measure polymerase pausing at a slipped strand structure. DNA polymerase is known to pause at some of these alternative structures, such as the structure known as the G-quadruplex, a DNA structure that has potentially importing regulatory significance. We examine polymerase kinetics around a G-quadruplex, and find evidence of polymerase pausing in the Pacific Biosciences kinetics. We provide a method, with publicly available code, so that others can examine these polymerase kinetics for any sequence of interest.

Bioinformatics

A null model for Pearson coexpression networks

Gene coexpression networks inferred by correlation from high-throughput profiling such as microarray data represent a simple but effective technique for discovering and interpreting linear gene relationships. In the last years several approach have been proposed to tackle the problem of deciding when the resulting correlation values are statistically significant. This is mostly crucial when the number of samples is small, yielding a non negligible chance that even high correlation values are due to random effects. Here we introduce a novel hard thresholding solution based on the assumption that a coexpression network inferred by randomly generated data is expected to be empty. The theoretical derivation of the new bound by geometrical methods is shown together with applications in onco- and neurogenomics.

Bioinformatics

libRoadRunner: A High Performance SBML Compliant Simulator

SummaryWe describe libRoadRunner, a cross-platform, open-source, high performance C++ library for running and analyzing SBML-compliant models. libRoadRunner was created primarily to achieve high performance, ease of use, portability and an extensible architecture. libRoadRunner includes a comprehensive API, Plugin support, Python scripting and additional functionality such as stoichio-metric and metabolic control analysis.\n\nAccessibility and ImplementationTo maximize collaboration, we made libRoadRunner open source and released it under the Apache License, Version 2.0. To facilitate reuse, we have developed comprehensive Python bindings using SWIG (swig.org) and a C API. Li-bRoadRunner uses a number of statically linked third party libraries including: LLVM [4], libSBML [1], CVODE, NLEQ2, LAPACK and Poco. LibRoadRunner is supported on Windows, Mac OS X and Linux.\n\nSupplementary informationOnline documentation, build instructions and git source repository are available at http://www.libroadrunner.org

Bioinformatics

Accurate detection of de novo and transmitted INDELs within exome-capture data using micro-assembly

We present a new open-source algorithm, Scalpel, for sensitive and specific discovery of INDELs in exome-capture data. By combining the power of mapping and assembly, Scalpel searches the de Bruijn graph for sequence paths (contigs) that span each exon. The algorithm creates a single path for exons with no INDEL, two paths for an exon with a heterozygous mutation, and multiple paths for more exotic variations. A detailed repeat composition analysis coupled with a self-tuning k-mer strategy allows Scalpel to outperform other state-of-the-art approaches for INDEL discovery. We extensively compared Scalpel with a battery of >10000 simulated and >1000 experimentally validated INDELs between 1 and 100bp against two recent algorithms for INDEL discovery: GATK HaplotypeCaller and SOAPindel. We report anomalies for these tools in their ability to detect INDELs, especially in regions containing near-perfect repeats which contribute to high false positive rates. In contrast, Scalpel demonstrates superior specificity while maintaining high sensitivity. We also present a large-scale application of Scalpel for detecting de novo and transmitted INDELs in 593 families with autistic children from the Simons Simplex Collection. Scalpel demonstrates enhanced power to detect long ([&ge;]20bp) transmitted events, and strengthens previous reports of enrichment for de novo likely gene-disrupting INDEL mutations in children with autism with many new candidate genes. The source code and documentation for the algorithm is available at http://scalpel.sourceforge.net.

Bioinformatics

A coarse-grained elastic network atom contact model and its use in the simulation of protein dynamics and the prediction of the effect of mutations

Normal mode analysis (NMA) methods are widely used to study dynamic aspects of protein structures. Two critical components of NMA methods are coarse-graining in the level of simplification used to represent protein structures and the choice of potential energy functional form. There is a trade-off between speed and accuracy in different choices. In one extreme one finds accurate but slow molecular-dynamics based methods with all-atom representations and detailed atom potentials. On the other extreme, fast elastic network model (ENM) methods with Conly representations and simplified potentials that based on geometry alone, thus oblivious to protein sequence. Here we present ENCoM, an Elastic Network Contact Model that employs a potential energy function that includes a pairwise atom-type non-bonded interaction term and thus makes it possible to consider the effect of the specific nature of amino-acids on dynamics within the context of NMA. ENCoM is as fast as existing ENM methods and outperforms such methods in the generation of conformational ensembles. Here we introduce a new application for NMA methods with the use of ENCoM in the prediction of the effect of mutations on protein stability. While existing methods are based on machine learning or enthalpic considerations, the use of ENCoM, based on vibrational normal modes, is based on entropic considerations. This represents a novel area of application for NMA methods and a novel approach for the prediction of the effect of mutations. We compare ENCoM to a large number of methods in terms of accuracy and self-consistency. We show that the accuracy of ENCoM is comparable to that of the best existing methods. We show that existing methods are biased towards the prediction of destabilizing mutations and that ENCoM is less biased at predicting stabilizing mutations.

Bioinformatics