Search bioRxivSearch

Biology subjects

von Mering, C.

Publications and source records attributed to von Mering, C..

6 recordsLinked to original sources

Tree reconciliation combined with subsampling improves large scale inference of orthologous group hierarchies

BackgroundAn orthologous group (OG) comprises a set of orthologous and paralogous genes that share a last common ancestor (LCA). OGs are defined with respect to a chosen taxonomic level, which delimits the position of the LCA in time to a specified speciation event. A hierarchy of OGs expands on this notion, connecting more general OGs, distant in time, to more recent, fine-grained OGs, thereby spanning multiple levels of the tree of life. Large scale inference of OG hierarchies with independently computed taxonomic levels can suffer from inconsistencies between successive levels, such as the position in time of a duplication event. This can be due to confounding genetic signal or algorithmic limitations. Importantly, inconsistencies limit the potential use of OGs for functional annotation and third-party applications.\n\nResultsHere we present a new methodology to ensure hierarchical consistency of OGs across taxonomic levels. To resolve an inconsistency, we subsample the protein space of the OG members and perform gene tree-species tree reconciliation for each sampling. Differently from previous approaches, by subsampling the protein space, we avoid the notoriously diffcult task of accurately building and reconciling very large phylogenies. We implement the method into a high-throughput pipeline and apply it to the eggNOG database. We use independent protein domain definitions to validate its performance.\n\nConclusionThe presented consistency pipeline shows that, contrary to previous limitations, tree reconciliation can be a useful instrument for the construction of OG hierarchies. The key lies in the combination of sampling smaller trees and aggregating their reconciliations for robustness. Results show comparable or greater performance to previous pipelines. The code is available on Github at: https://github.com/meringlab/og_consistency_pipeline

bioinformatics

Viruses.STRING: A virus-host protein-protein interaction database

As viruses continue to pose risks to global health, having a better un-derstanding of virus-host protein-protein interactions aids in the development of treatments and vaccines. Here, we introduce Viruses.STRING, a protein-protein interaction database specifically catering to virus-virus and virus-host interactions. This database combines evidence from experimental and text-mining channels to provide combined probabilities for interactions between viral and host proteins. The database contains 177,425 interactions between 239 viruses and 319 hosts. The database is publicly available at viruses.string-db.org, and the interaction data can also be accessed through the latest version of the Cytoscape STRING app.

bioinformatics

Rapid inference of direct interactions in large-scale ecological networks from heterogeneous microbial sequencing data

The recent explosion of metagenomic sequencing data opens the door towards the modeling of microbial ecosystems in unprecedented detail. In particular, co-occurrence based prediction of ecological interactions could strongly benefit from this development. However, current methods fall short on several fronts: univariate tools do not distinguish between direct and indirect interactions, resulting in excessive false positives, while approaches with better resolution are so far computationally highly limited. Furthermore, confounding variables typical for cross-study data sets are rarely addressed. We present FlashWeave, a new approach based on a flexible Probabilistic Graphical Models framework to infer highly resolved direct microbial interactions from massive heterogeneous microbial abundance data sets with seamless integration of metadata. On a variety of benchmarks, FlashWeave outperforms state-of-the-art methods by several orders of magnitude in terms of speed while generally providing increased accuracy. We apply FlashWeave to a cross-study data set of 69 818 publicly available human gut samples, resulting in one of the largest and most diverse models of microbial interactions in the human gut to date.

systems biology

Pathway and network analysis of more than 2,500 whole cancer genomes

The catalog of cancer driver mutations in protein-coding genes has greatly expanded in the past decade. However, non-coding cancer driver mutations are less well-characterized and only a handful of recurrent non-coding mutations, most notably TERT promoter mutations, have been reported. Motivated by the success of pathway and network analyses in prioritizing rare mutations in protein-coding genes, we performed multi-faceted pathway and network analyses of non-coding mutations across 2,583 whole cancer genomes from 27 tumor types compiled by the ICGC/TCGA PCAWG project. While few non-coding genomic elements were recurrently mutated in this cohort, we identified 93 genes harboring non-coding mutations that cluster into several modules of interacting proteins. Among these are promoter mutations associated with reduced mRNA expression in TP53, TLE4, and TCF4. We found that biological processes had variable proportions of coding and non-coding mutations, with chromatin remodeling and proliferation pathways altered primarily by coding mutations, while developmental pathways, including Wnt and Notch, altered by both coding and non-coding mutations. RNA splicing was primarily targeted by non-coding mutations in this cohort, with samples containing non-coding mutations exhibiting similar gene expression signatures as coding mutations in well-known RNA splicing factors. These analyses contribute a new repertoire of possible cancer genes and mechanisms that are altered by non-coding mutations and offer insights into additional cancer vulnerabilities that can be investigated for potential therapeutic treatments.

cancer biology

Analysis of the human kinome and phosphatome reveals diseased signaling networks induced by overexpression

Kinase and phosphatase overexpression drives tumorigenesis and drug resistance in many cancer types. Signaling networks reprogrammed by protein overexpression remain largely uncharacterized, hindering discovery of paths to therapeutic intervention. We previously developed a single cell proteomics approach based on mass cytometry that enables quantitative assessment of overexpression effects on the signaling network. Here we applied this approach in a human kinome- and phosphatome-wide study to assess how 649 individually overexpressed proteins modulate the cancer-related signaling network in HEK293T cells. Based on these data we expanded the functional classification of human kinases and phosphatases and detected 208 novel signaling relationships. In the signaling dynamics analysis, we showed that increased ERK-specific phosphatases sustained proliferative signaling, and using a novel combinatorial overexpression approach, we confirmed this phosphatase-driven mechanism of cancer progression. Finally, we identified 54 proteins that caused ligand-independent ERK activation with potential as biomarkers for drug resistance in cells carrying BRAF mutations.

systems biology

MAPseq: Improved Speed, Accuracy And Consistency In Ribosomal RNA Sequence Analysis

Metagenomic sequencing has become crucial to studying microbial communities, but meaningful taxonomic analysis and inter-comparison of such data are still hampered by technical limitations, between-study design variability and inconsistencies between taxonomies used. Here we present MAPseq, a framework for reference-based rRNA metagenomic analysis that is up to 30% more accurate (F1/2 score) and up to one hundred times faster than existing solutions, providing in a single run multiple taxonomy classifications and hierarchical OTU mappings, for both amplicon and shotgun sequencing strategies, and for datasets of virtually any size. Availability: Source code and binaries are freely available at http://meringlab.org/software/mapseq/

bioinformatics