Search bioRxiv⌕ Search

Biology subjects

Waters, M.

Publications and source records attributed to Waters, M..

5 recordsLinked to original sources

SYNGAP1 haploinsufficiency disrupts early neurodevelopment and accelerates intrinsic neuronal maturation in human patient-derived models

SYNGAP1 developmental and epileptic encephalopathy (DEE) is a severe neurodevelopmental disorder characterised by intellectual disability, developmental delay, and refractory epilepsy caused by heterozygous variants in SYNGAP1, which encodes Synaptic Ras GTPase-activating protein 1. While SYNGAP1 is best known for its role at the postsynaptic density, increasing evidence indicates that haploinsufficiency also disrupts early neurodevelopment. Here, we used patient-derived induced pluripotent stem cell (iPSC) models to investigate early neurodevelopmental and neuronal phenotypes associated with SYNGAP1 haploinsufficiency. iPSCs derived from a female patient carrying the frameshift variant p.Leu150Valfs*6 were differentiated into two complementary models: micropatterned neural rosettes representing early neuroepithelial organisation and NGN2-induced excitatory neurons representing postmitotic functional development. Patient-derived neural rosettes displayed enlarged, dysmorphic lumens, indicating disrupted neuroepithelial organisation at the earliest stages of brain development. Transcriptomic profiling revealed widespread dysregulation of genes involved in neurodevelopment, cell adhesion and ion channel regulation, including coordinated downregulation of protocadherin family members. Whole-cell patch-clamp electrophysiology demonstrated reduced input resistance, larger action potential amplitudes, and increased inward and outward current densities, consistent with accelerated intrinsic neuronal maturation rather than generalized hyperexcitability. Together, these complementary findings demonstrate that SYNGAP1 haploinsufficiency disrupts early human brain development and accelerates intrinsic neuronal maturation, with pathogenic mechanisms emerging before synaptogenesis and extending beyond SYNGAP1s established synaptic role.

neuroscience↗

MetaHarmonizer: robust biomedical metadata harmonization and a contamination control for inflated LLM performance on public benchmarks

Public biomedical repositories hold substantial reuse potential, but inconsistent metadata routinely blocks integration across studies. Recent LLM-based harmonization approaches address scale but suffer from non-determinism, hallucinated ontology terms, and, in their highest-accuracy configurations, dependence on proprietary APIs or labeled fine-tuning data. A more fundamental concern is that LLM accuracies on widely-used public benchmarks may substantially inflate transferable capability: under a contamination-controlled evaluation protocol we developed, the apparent LLM-only advantage on the GDC schema-mapping benchmark is inverted and three out of five LLMs recovers 80-100% of GDC identifiers from zero-schema context, suggesting direct memorization. Building on this insight, we present MetaHarmonizer, an automated metadata harmonization system designed to be robust by construction: SchemaMapper aligns attribute names across schemas, and OntologyMapper standardizes values to controlled vocabularies. Both modules implement a multi-stage cascade that escalates to more resource-intensive methods only when earlier stages fall short, with all candidates grounded in pre-defined controlled vocabularies to preclude hallucinated outputs and LLMs used only as bounded preprocessing components rather than inference-time dependencies. On the GDC schema-matching benchmark, SchemaMapper with the deployment-optimized LLM-generated alias dictionary achieved 71.6% Top-1 accuracy and the higher Recall@GT than Magneto bipartite variants, recovering significantly more ground-truth mappings; with the best performing alias dictionary, it reached the highest Top-1/Top-5/Recall@GT, and also matched the best Magneto reranker (fine-tuned LLM-reranker) on MRR; and it also outperforms LLM-only performance under contamination-controlled conditions. On four EFO benchmarks, OntologyMapper achieved 77.9-95.5% Top-1 accuracy, outperforming text2term by up to 16.4 pp and direct LLM inference (against the smaller corpus) by 19.2 pp because memorization is not a viable shortcut for this task. Across both modules, calibrated confidence scores separate correct from incorrect predictions (AUC 0.73-0.94), enabling principled human-in-the-loop triage. Inference is fully local, deterministic, and computationally efficient - seconds on schema mapping and under a minute for ontology mapping of up to [~]7,000 terms against the pre-indexed 33,230-term corpus. Released as a Python package with a domain-agnostic architecture, MetaHarmonizer provides a scalable foundation for improving the FAIRness of biomedical data and enabling cross-study integration, alongside an evaluation methodology applicable to any LLM-augmented bioinformatics benchmark built on public benchmarks.

bioinformatics↗

Generalist large language models complement tailor-made predictors for tumor genomics interpretation

General-purpose large language models (LLMs) are trained on large corpora to acquire broad knowledge, but whether LLMs can replace, or augment, task-specific models is unclear. We evaluated LLMs on three real-world, clinically important tumor genomic interpretation tasks, in order of increasing difficulty: (i) distinguishing tumor from non-tumor mutations (n=34,415 variants), (ii) distinguishing driver from passenger mutations (n=13,469 variants), and (iii) inferring cancer type from tumor sequencing reports across multiple assays and institutions (n=102,791 samples). The best general-purpose LLMs performed as well as the benchmark tailor-made predictor for task (i). Ensembling tailor-made models with zero-shot LLMs improved their performance for tasks (i) and (ii). For task (iii), LLMs outperformed or supplemented tailor-made models on out-of-distribution data. Without fine-tuning, current LLMs already can be useful in clinical genomic interpretation by adding complementary expertise to tailor-made, state-of-the-art predictors.

genomics↗

Integrated histopathologic modeling of detailed tumor subtypes and actionable biomarkers

Accurate cancer subtyping with accompanying molecular characterization is critical for precision oncology. While machine learning approaches have been applied to both digital pathology and cancer genomics, previous work has been limited in sample size and has typically aggregated granular cancer subtypes into coarse groupings, likely obfuscating informative molecular and prognostic associations and phenotypic variation of more detailed tumor subtypes. Accordingly, we collated 378,123 hematoxylin and eosin (H&E)-stained whole-slide images (WSIs) with matched targeted DNA clinical sequencing results and OncoTree detailed cancer subtypes from a real-world cohort of 71,142 patients. Using this scaled, granular dataset and a cancer subtype knowledge graph, we developed Mosaic: a family of calibrated machine learning models using H&E WSI embeddings to classify tumors and identify molecular phenotypes across 163 detailed subtypes. The cancer subtyping module (Aeon) achieved an area under the receiver operating characteristic curve (AUROC) of 0.992 overall, with 161/163 subtypes reaching an AUROC [≥] 0.90 and improved performance over a state-of-the-art genomics-based classifier. The genomic inference module (Paladin) achieved an AUROC [≥] 0.80 for 167 pairs of detailed subtypes and genomic targets. We further used the learned histopathologic representations to i) identify key associations of the histopathologic embeddings with clinical biomarkers; ii) identify unsupervised sub-clusters of tumors with genomic determinants of tumor phenotype; iii) specify granular diagnoses for cancers of unknown primary, evaluated by genomic associations and expected clinical outcome distributions; iv) annotate functional significance for variants of uncertain significance (VUS); and v) identify cases that mimic the phenotypic effect of known DNA variants on H&E in the absence of detectable DNA alterations. Taken together, this work advances our understanding of phenotypic variation of granular tumor subtypes, their relevance to enhanced diagnostics, and their potential utility in risk stratification with multimodal machine learning in cancer.

cancer biology↗

Machine learning predictions improve identification of real-world cancer driver mutations

Characterizing and validating which mutations influence development of cancer is challenging. Machine learning has delivered significant advances in protein structure prediction, but its utility for identifying cancer drivers is less explored. We evaluated multiple computational methods for identifying cancer driver alterations. For identifying known drivers, methods incorporating protein structure or functional genomic data outperformed methods trained only on evolutionary data. We further validated VUSs annotated as pathogenic by testing their association with overall survival in two cohorts of patients with non-small cell lung cancer (N=7,965 and 977). "Pathogenic" VUSs in KEAP1 and SMARCA4 identified by several methods were associated with worse survival, unlike "benign" VUSs. "Pathogenic" VUSs exhibited mutual exclusivity with known oncogenic alterations at the pathway level, further suggesting biological validity. Despite training primarily on germline, rather than somatic, mutation data, computational predictions contribute to a more comprehensive understanding of tumor genetics as validated by real-world data.

genomics↗