Search bioRxivSearch

Biology subjects

Noble, W.

Publications and source records attributed to Noble, W..

9 recordsLinked to original sources

An averaging strategy to reduce variability in target-decoy estimates of false discovery rate

Decoy database search with target-decoy competition (TDC) provides an intuitive, easy-to-implement method for estimating the false discovery rate (FDR) associated with spectrum identifications from shotgun proteomics data. However, the procedure can yield different results for a fixed dataset analyzed with different decoy databases, and this decoy-induced variability is particularly problematic for smaller FDR thresholds, datasets or databases. In such cases, the nominal FDR might be 1% but the true proportion of false discoveries might be 10%. The averaged TDC protocol combats this problem by exploiting multiple independently shuffled decoy databases to provide an FDR estimate with reduced variability. We provide a tutorial introduction to aTDC, describe an improved variant of the protocol that offers increased statistical power, and discuss how to deploy aTDC in practice using the Crux software toolkit.

bioinformatics

Predicting gene expression in the human malaria parasite Plasmodium falciparum

Empirical evidence suggests that the malaria parasite Plasmodium falciparum employs a broad range of mechanisms to regulate gene transcription throughout the organisms complex life cycle. To better understand this regulatory machinery, we assembled a rich collection of genomic and epigenomic data sets, including information about transcription factor (TF) binding motifs, patterns of covalent histone modifications, nucleosome occupancy, GC content, and global 3D genome architecture. We used these data to train machine learning models to discriminate between high-expression and low-expression genes, focusing on three distinct stages of the red blood cell phase of the Plasmodium life cycle. Our results highlight the importance of histone modifications and 3D chromatin architecture and suggest a relatively small role for TF binding in Plasmodium transcriptional regulation.

genomics

Combining high resolution and exact calibration to boost statistical power: A well-calibrated score function for high-resolution MS2 data

To achieve accurate assignment of peptide sequences to observed fragmentation spectra, a shotgun proteomics database search tool must make good use of the very high resolution information produced by state-of-the-art mass spectrometers. However, making use of this information while also ensuring that the search engines scores are well calibrated--i.e., that the score assigned to one spectrum can be meaningfully compared to the score assigned to a different spectrum--has proven to be challenging. Here, we describe a database search score function, the \"residue evidence\" (res-ev) score, that achieves both of these goals simultaneously. We also demonstrate how to combine calibrated res-ev scores with calibrated XCorr scores to produce a \"combined p-value\" score function. We provide a benchmark consisting of four mass spectrometry data sets, which we use to compare the combined p-value to the score functions used by several existing search engines. Our results suggest that the combined p-value achieves state-of-the-art performance, generally outperforming MS Amanda and Morpheus and performing comparably to MS-GF+. The res-ev and combined p-value score functions are freely available as part of the Tide search engine in the Crux mass spectrometry toolkit (http://crux.ms).

bioinformatics

Unsupervised embedding of single-cell Hi-C data

Single-cell Hi-C (scHi-C) data promises to enable scientists to interrogate the 3D architecture of DNA in the nucleus of the cell, studying how this structure varies stochastically or along developmental or cell cycle axes. However, Hi-C data analysis requires methods that take into account the unique characteristics of this type of data. In this work, we explore whether methods that have been developed previously for the analysis of bulk Hi-C data can be applied to scHi-C data. In this work, we apply methods designed for analysis of bulk Hi-C data to scHi-C data in conjunction with unsupervised embedding. We find that one of these methods, HiCRep, when used in conjunction with multidimensional scaling (MDS), strongly outperforms three other methods, including a technique that has been used previously for scHi-C analysis. We also provide evidence that the HiCRep/MDS method is robust to extremely low per-cell sequencing depth, that this robustness is improved even further when high-coverage and low-coverage cells are projected together, and that the method can be used to jointly embed cells from multiple published datasets.

bioinformatics

Changes in genome organization of parasite-specific gene families during the Plasmodium transmission stages

The development of malaria parasites throughout their various life cycle stages is controlled by coordinated changes in gene expression. We previously showed that the three-dimensional organization of the P. falciparum genome is strongly associated with gene expression during its replication cycle inside red blood cells. Here, we analyzed genome organization in the P. falciparum and P. vivax transmission stages. Major changes occurred in the localization and interactions of genes involved in pathogenesis and immune evasion, erythrocyte and liver cell invasion, sexual differentiation and master regulation of gene expression. In addition, we observed reorganization of subtelomeric heterochromatin around genes involved in host cell remodeling. Depletion of heterochromatin protein 1 (PfHP1) resulted in loss of interactions between virulence genes, confirming that PfHP1 is essential for maintenance of the repressive center. Overall, our results suggest that the three-dimensional genome structure is strongly connected with transcriptional activity of specific gene families throughout the life cycle of human malaria parasites.

systems biology

MoMo: Discovery of post-translational modification motifs

MotivationPost-translational modifications (PTMs) of proteins are associated with many significant biological functions and can be identified in high throughput using tandem mass spectrometry. Many PTMs are associated with short sequence patterns called \"motifs\" that help localize the modifying enzyme. Accordingly, many algorithms have been designed to identify these motifs from mass spectrometry data.\n\nResultsMoMo is a software tool for identifying motifs among sets of PTMs. The program re-implements two previously described algorithms, Motif-X and MoDL, packaging them in a web-accessible user interface. In addition to reading sequence files in FASTA format, MoMo is capable of directly parsing output files produced by commonly used mass spectrometry search engines. The resulting motifs are presented to the user in an HTML summary with motif logos and linked text files in MEME motif format.\n\nAvailabilitySource code and web server available at http://meme-suite.org\n\nContacttimothybailey@unr.edu and william-noble@uw.edu\n\nSupplementary informationSupplementary figures are available at Bioinformatics online.

bioinformatics

Comprehensive statistical inference of the clonal structure of cancer from multiple biopsies

A comprehensive characterization of tumor genetic heterogeneity is critical for understanding how cancers evolve and escape treatment. Although many algorithms have been developed for capturing tumor heterogeneity, they are designed for analyzing either a single type of genomic aberration or individual biopsies. Here we present THEMIS (Tumor Heterogeneity Extensible Modeling via an Integrative System), which allows for the joint analysis of different types of genomic aberrations from multiple biopsies taken from the same patient, using a dynamic graphical model. Simulation experiments demonstrate higher accuracy of THEMIS over its ancestor, TITAN. The heterogeneity analysis results from THEMIS are validated with single cell DNA sequencing from a clinical tumor biopsy. When THEMIS is used to analyze tumor heterogeneity among multiple biopsies from the same patient, it helps to reveal the mutation accumulation history, track cancer progression, and identify the mutations related to treatment resistance. We implement our model via an extensible modeling platform, which makes our approach open, reproducible, and easy for others to extend.

bioinformatics

Nucleotide sequence and DNaseI sensitivity are predictive of 3D chromatin architecture

Recently, Hi-C has been used to probe the 3D chromatin architecture of multiple organisms and cell types. The resulting collections of pairwise contacts across the genome have connected chromatin architecture to many cellular phenomena, including replication timing and gene regulation. However, high resolution (10 kb or finer) contact maps remain scarce due to the expense and time required for collection. A computational method for predicting pairwise contacts without the need to run a Hi-C experiment would be invaluable in understanding the role that 3D chromatin architecture plays in genome biology. We describe Rambutan, a deep convolutional neural network that predicts Hi-C contacts at 1 kb resolution using nucleotide sequence and DNaseI assay signal as inputs. Specifically, Rambutan identifies locus pairs that engage in high confidence contacts according to Fit-Hi-C, a previously described method for assigning statistical confidence estimates to Hi-C contacts. We first demonstrate Rambutans performance across chromosomes at 1 kb resolution in the GM12878 cell line. Subsequently, we measure Rambutans performance across six cell types. In this setting, the model achieves an area under the receiver operating characteristic curve between 0.7662 and 0.8246 and an area under the precision-recall curve between 0.3737 and 0.9008. We further demonstrate that the predicted contacts exhibit expected trends relative to histone modification ChlP-seq data, replication timing measurements, and annotations of functional elements such as promoters and enhancers. Finally, we predict Hi-C contacts for 53 human cell types and show that the predictions cluster by cellular function. [NOTE: After our original submission we discovered an error in our calling of statistically significant contacts. Briefly, when calculating the prior probability of a contact, we used the number of contacts at a certain genomic distance in a chromosome but divided by the total number of bins in the full genome. While we investigate what impact this had on our results, we ask that readers treat this manuscript skeptically.]

bioinformatics

Cis-Compound Mutations are Prevalent in Triple Negative Breast Cancer and Can Drive Tumor Progression

About 16% of breast cancers fall into a clinically aggressive category designated triple negative (TNBC) due to a lack of ERBB2, estrogen receptor and progesterone receptor expression1-3. The mutational spectrum of TNBC has been characterized as part of The Cancer Genome Atlas (TCGA)4; however, snapshots of primary tumors cannot reveal the mechanisms by which TNBCs progress and spread. To address this limitation we initiated the Intensive Trial of OMics in Cancer (ITOMIC)-001, in which patients with metastatic TNBC undergo multiple biopsies over space and time5. Whole exome sequencing (WES) of 67 samples from 11 patients identified 426 genes containing multiple distinct single nucleotide variants (SNVs) within the same sample, instances we term Multiple SNVs affecting the Same Gene and Sample (MSSGS). We find that >90% of MSSGS result from cis-compound mutations (in which both SNVs affect the same allele), that MSSGS comprised of SNVs affecting adjacent nucleotides arise from single mutational events, and that most other MSSGS result from the sequential acquisition of SNVs. Some MSSGS drive cancer progression, as exemplified by a TNBC driven by FGFR2(S252W;Y375C). MSSGS are more prevalent in TNBC than other breast cancer subtypes and occur at higher-than-expected frequencies across TNBC samples within TCGA. MSSGS may denote genes that play as yet unrecognized roles in cancer progression.

clinical trials