Search bioRxiv⌕ Search

Biology subjects

Popov, N.

Publications and source records attributed to Popov, N..

2 recordsLinked to original sources

A ML-framework for the discovery of next-generation IBD targets using a harmonized single-cell atlas of patient tissue

Target discovery for IBD has traditionally relied on genetic associations, which lack the cellular resolution needed to identify novel, actionable, cell type-specific disease pathways. Here, we describe an integrated analytical and experimental framework that leverages harmonized single-cell data to systematically discover novel therapeutic strategies for IBD. We used AMICA DBTM, Immunais harmonized database of single-cell RNA datasets to construct a harmonized 1 million single-cell atlas of the human intestine. We applied a machine learning framework (Immune Patient Representation, IPR) to identify disease-associated transcriptional programs and cell type-specific gene targets. Candidate targets were prioritized using atlas-derived metrics, refined using custom criteria emphasizing translational actionability, and validated across independent clinical cohorts. Select candidates were evaluated in human primary-cell models reflecting the targets cell-type context. The IPR framework identified 85 disease-associated transcriptional programs and ranked 400 cell type-specific target genes across immune and stromal lineages. Disease-associated programs were interpreted using a structured AI-assisted reasoning framework for structured biological reasoning, linking them to IBD-relevant pathways and guiding the identification of novel, promising gene targets. Functional validation of two cell-type-specific candidates, PTGIR in myeloid cells and IL6ST in fibroblasts, confirmed the reduction of inflammatory and fibrotic pathways linked to IBD pathology. Multi-omic profiling and projection of in vitro phenotypes to patient datasets demonstrated the reversal of disease-associated programs via mechanisms distinct from those of existing biologics. Our single-cell anchored, machine-learning framework integrates in silico discovery with experimental validation, revealing new cell type-specific therapeutic opportunities and supporting a scalable approach for precision target discovery in IBD and other immune-mediated diseases.

bioinformatics↗

AliMarko: A Novel Tool for Eukaryotic Virus Identification Using Expert-Guided Approach

Metagenomic sequencing is a valuable tool for studying viral diversity in biological samples. Analyzing this data is complex due to the high variability of viral genomes and their low representation in databases. We present the Alimarko pipeline, designed to streamline virus identification in metagenomic data. A key feature of our tool is the focus on the interpretability of findings: results are provided with tabular and visual information to help determine the confidence level in the identified viral sequences. The pipeline employs two approaches for identifying viral sequences: mapping to reference genomes and de novo assembly followed by the application of Hidden Markov Models (HMM). Additionally, it includes a step for phylogenetic analysis, which constructs a phylogenetic tree to determine the evolutionary relationships with reference sequences. We also emphasize reducing false-positive results. Reads related to cellular organisms are computationally depleted, and the identified viral sequences are checked against a list of potential contaminants. The output is an HTML document containing visualizations and tabular information designed to assist researchers in making informed decisions about the presence of viruses. Using our pipeline for total RNA sequencing of bat feces, we identified a range of viruses and rapidly determined the validity and phylogenetic relationships of the findings to known sequences with the aid of reports generated by AliMarko.

bioinformatics↗