Search bioRxiv⌕ Search

Biology subjects

Mizikovsky, D.

Publications and source records attributed to Mizikovsky, D..

6 recordsLinked to original sources

Unsupervised Variant Clustering Identifies Genetic Subtypes of Disease

Current approaches to identify disease subtypes rely on prior knowledge or external phenotypic references that limit their generalisability across traits. This study presents an unsupervised, phenotype- free framework that analyses the genetic co-occurrence of variants across individuals to identify variant clusters underpinning disease heterogeneity. The approach is developed and validated using simulated combinations of real phenotypes then applied to dissect disease subtypes of type 2 diabetes and asthma. The approach captures clinically established mechanistic profiles in type 2 diabetes and asthma without requiring any reference phenotype data and enables identification of plasma proteins for novel subtype- specific biomarkers and drug targets. Lastly, we demonstrate that variants with conflicting effects on disease-relevant traits are resolved into distinct clusters, identifying subtypes with opposing phenotypic profiles that are lost by the aggregated genetic risk. Co-occurrence-based clustering provides a scalable strategy for understanding disease heterogeneity across diverse populations by identifying biologically meaningful subtype structure from genotype data alone.

genomics↗

Glycaemic variability underlies myocyte dysfunction and myocardial injury risk in diabetes

Heart disease is the leading cause of morbidity and mortality in individuals with diabetes, due largely to risks associated with ischaemic injuries such as myocardial infarction (MI). We use human population genetic data to demonstrate that current biomarkers of hyperglycaemia do not account for risk of post-MI mortality in diabetes patients. This study therefore systematically evaluates glycaemic stress underpinning cardiovascular risk in diabetes. Using in vivo and in vitro models, we demonstrate that glycaemic variability rather than hyperglycaemia alone is a dominant risk factor for heart muscle dysfunction and myocardial injury sensitivity in diabetes. These findings provide new preclinical models for mechanistic and drug discovery studies and inform strategies for managing cardiovascular outcomes in patients with diabetes.

cell biology↗

Epigenetic constraint of cellular genomes evolutionarily links genetic variation to function

Cellular diversity is a product of evolution acting to drive divergent regulatory programs from a common genome. Here, we use cross-cell-type epigenetic conservation to gain insight into the impact of selective constraints on genome function and phenotypic variation. By comparing chromatin accessibility across hundreds of diverse cell-types, we identify 1.4% of the human genome safeguarded by conserved domains of facultative heterochromatin, which we term regions under "cellular constraint". We calculate single-base resolution cellular constraint scores and demonstrate robust prediction of functionally important coding and non-coding loci in a cell-type-, trait-, and disease-agnostic manner. Cellular constraint annotation enhances causal variant identification, drug discovery, and clinical diagnostic predictions. Furthermore, cell-constrained sequences share paradoxical evolutionary signals of positive and negative selection, suggesting a dynamic role in driving human adaptation. Overall, this study demonstrates that evolutionary chromatin dynamics can be leveraged to inform the translation of genetic discoveries into effective biological, therapeutic, and clinical outcomes.

genetics↗

A robust unsupervised clustering approach for high-dimensional biological imaging data reveals shared drug-induced morphological signatures

Modern biology increasingly relies on large-scale screening to generate high dimensional datasets with potential to accelerate discovery. However, analysing these complex datasets remains challenging, particularly in applications where the underlying structure and groupings are unknown, and high dimensionality introduces noise and artifacts that make follow up studies difficult to prioritise. Here, we present an unsupervised consensus clustering tool that quantifies biologically meaningful patterns based on multi-scale data organisation to guide decision-making in high-throughput screening. Using large-scale drug screening data in cancer cell lines and bacterium model, we demonstrate its ability to use diverse data inputs to prioritize robust drug clusters with shared biological mechanisms and conserved drug responses. This method addresses key limitations associated with prioritising robust, actionable hits from scalable screening data.

bioinformatics↗

HOPX governs a molecular and physiological switch between cardiomyocyte progenitor and maturation gene programs

This study establishes the homeodomain only protein, HOPX, as a determinant controlling the molecular switch between cardiomyocyte progenitor and maturation gene programs. Time-course single-cell gene expression with genome-wide footprinting reveal that HOPX interacts with and controls core cardiac networks by regulating the activity of mutually exclusive developmental gene programs. Upstream hypertrophy and proliferation pathways compete to regulate HOPX transcription. Mitogenic signals override hypertrophic growth signals to suppress HOPX and maintain cardiomyocyte progenitor gene programs. Physiological studies show HOPX directly governs genetic control of cardiomyocyte cell stress responses, electro-mechanical coupling, proliferation, and contractility. We use human genome-wide association studies (GWAS) to show that genetic variation in the HOPX-regulome is significantly associated with complex traits affecting cardiac structure and function. Collectively, this study provides a mechanistic link situating HOPX between competing upstream pathways where HOPX acts as a molecular switch controlling gene regulatory programs underpinning metabolic, signaling, and functional maturation of cardiomyocytes.

developmental biology↗

Organisation of gene programs revealed by unsupervised analysis of diverse gene-trait associations

Genome wide association studies provide statistical measures of gene-trait associations that reveal how genetic variation influences phenotypes. This study develops an unsupervised dimensionality reduction method called UnTANGLeD (Unsupervised Trait Analysis of Networks from Gene Level Data) which organises 16,849 genes into discrete gene programs by measuring the statistical association between genetic variants and 1,393 diverse complex traits. UnTANGLeD reveals 173 gene clusters enriched for protein-protein interactions and highly distinct biological processes governing development, signalling, disease, and homeostasis. We identify diverse gene networks with robust interactions but not associated with known biological processes. Analysis of independent disease traits shows that UnTANGLeD gene clusters are conserved across all complex traits, providing a simple and powerful framework to predict novel gene candidates and programs influencing orthogonal disease phenotypes. Collectively, this study demonstrates that gene programs co-ordinately orchestrating cell functions can be identified without reliance on prior knowledge, providing a method for use in functional annotation, hypothesis generation, machine learning and prediction algorithms, and the interpretation of diverse genomic data.

genomics↗