Search bioRxivSearch

Biology subjects

Soeding, J.

Publications and source records attributed to Soeding, J..

4 recordsLinked to original sources

Synthetic protein alignments by CCMgen quantify noise in residue-residue contact prediction

Compensatory mutations between protein residues that are in physical contact with each other can manifest themselves as statistical couplings between the corresponding columns in a multiple sequence alignment (MSA) of the protein family. Conversely, high coupling coefficients predict residues contacts. Methods for de-novo protein structure prediction based on this approach are becoming increasingly reliable. Their main limitation is the strong systematic and statistical noise in the estimation of coupling coefficients, which has so far limited their application to very large protein families. While most research has focused on boosting contact prediction quality by adding external information, little progress has been made to improve the statistical procedure at the core. In that regard, our lack of understanding of the sources of noise poses a major obstacle. We have developed CCMgen, the first method for simulating protein evolution by providing full control over the generation of realistic synthetic MSAs with pairwise statistical couplings between residue positions. This procedure requires an exact statistical model that reliably reproduces observed alignment statistics. With CCMpredPy we also provide an implementation of persistent contrastive divergence (PCD), a precise inference technique that enables us to learn the required high-quality statistical models. We demonstrate how CCMgen can facilitate the development and testing of contact prediction methods by analysing the systematic noise contributions from phylogeny and entropy. For that purpose we propose a simple entropy correction (EC) strategy which disentangles the correction for both sources of noise. We find that entropy contributes typically roughly twice as much noise as phylogeny.

bioinformatics

Reconstructing complex lineage trees from scRNA-seq data using MERLoT

Advances in single-cell transcriptomics techniques are revolutionizing studies of cellular differentiation and heterogeneity. Consequently, it becomes possible to track the trajectory of thousands of genes across the cellular lineage trees that represent the temporal emergence of cell types during dynamic processes. However, reconstruction of cellular lineage trees with more than a few cell fates has proved challenging. We present MERLoT (https://github.com/soedinglab/merlot), a flexible and user-friendly tool to reconstruct complex lineage trees from single-cell transcriptomics data and further impute temporal gene expression profiles along the reconstructed tree structures. We demonstrate MERLoTs capabilities on various real cases and hundreds of simulated datasets.

bioinformatics

PROSSTT: probabilistic simulation of single-cell RNA-seq data for complex differentiation processes

BackgroundSingle-cell RNA sequencing (scRNA-seq) is an enabling technology for the study of cellular differentiation and heterogeneity. From snapshots of the transcriptomic profiles of differentiating single cells, the cellular lineage tree that leads from a progenitor population to multiple types of differentiated cells can be derived. The underlying lineage trees of most published datasets are linear or have a single branchpoint, but many studies with more complex lineage trees will soon become available. To test and further develop tools for lineage tree reconstruction, we need test datasets with known trees.\n\nResultsPROSSTT can simulate scRNA-seq datasets for differentiation processes with lineage trees of any desired complexity, noise level, noise model, and size. PROSSTT also provides scripts to quantify the quality of predicted lineage trees.\n\nAvailabilityhttps://github.com/soedinglab/prosstt\n\nContactsoeding@mpibpc.mpg.de

bioinformatics

Bayesian multiple logistic regression for meta-analyses of GWAS

Genetic variants in genome-wide association studies (GWAS) are tested for disease association mostly using simple regression, one variant at a time. Standard approaches to improve power in detecting disease-associated SNPs use multiple regression with Bayesian variable selection in which a sparsity-enforcing prior on effect sizes is used to avoid overtraining and all effect sizes are integrated out for posterior inference. For binary traits, the logistic model has not yielded clear improvements over the linear model. For multi-SNP analysis, the logistic model required costly and technically challenging MCMC sampling to perform the integration.\n\nHere, we introduce the quasi-Laplace approximation to solve the integral and avoid MCMC sampling. We expect the logistic model to perform much better than multiple linear regression except when predicted disease risks are spread closely around 0.5, because only close to its inflection point can the logistic function be well approximated by a linear function. Indeed, in extensive benchmarks with simulated phenotypes and real genotypes, our Bayesian multiple LOgistic REgression method (B-LORE) showed considerable improvements (1) when regressing on many variants in multiple loci at heritabilities [≥] 0.4 and (2) for unbalanced case-control ratios.\n\nB-LORE also enables meta-analysis by approximating the likelihood functions of individual studies by multivariate normal distributions, using their means and covariance matrices as summary statistics. Our work should make sparse multiple logistic regression attractive also for other applications with binary target variables. B-LORE is freely available from: https://github.com/soedinglab/b-lore.

genetics