Search bioRxiv⌕ Search

Biology subjects

Ko, Y. K.

Publications and source records attributed to Ko, Y. K..

3 recordsLinked to original sources

Capsular K-antigen Coats Outer Membrane Vesicles of Porphyromonas gingivalis

Periodontal disease is an inflammatory disorder that arises from dysbiosis of the subgingival microbiota, with Porphyromonas gingivalis acting as a keystone pathogen in the shift from health to disease. P. gingivalis employs multiple strategies to subvert host immune defenses, and its capsular K-antigen serves as a key virulence determinant. Here, a pre-adsorbed antiserum (pAds106) was generated by removing nonspecific antibodies using cells from a K-antigen-null mutant (W83{Delta}PG0106), resulting in exceptional specificity for the P. gingivalis K1-antigen. Immunofluorescence analysis revealed that the K-antigen preferentially coats outer membrane vesicles (OMVs), rather than attaching to the bacterial cell surface. This localization was further confirmed by ELISAs of density gradient ultracentrifuge-purified OMVs, with background signal detected in OMVs derived from K-antigen-deficient strains, non-K1-strains, and other oral Bacteroidetes. K-antigen-coated OMVs exhibited higher hydrophilicity and elicited weaker inflammatory responses compared to K-antigen-deficient OMVs, consistent with previously reported properties of encapsulated strains. Importantly, the antiserum detected K-antigen-coated OMVs in subgingival plaque from periodontal patients, suggesting that K-antigen is actively produced at diseased sites. These findings revise the prevailing view that K-antigen solely encapsulates the bacterial cell body and suggest that K-antigen-coated OMVs produced by P. gingivalis play distinct roles in immune evasion during periodontal disease.

microbiology↗

CAMP: Coreset Accelerated Metacell Partitioning enables scalable analysis of single-cell data

Scaling metacell inference to atlas-level single-cell datasets demands algorithms that are both computationally efficient and geometrically faithful. We introduce CAMP (Coreset Accelerated Metacell Partitioning), a metacell framework that preserves the intrinsic structure of the cellular manifold while enabling scalable analysis of millions of cells. CAMP leverages coreset-based sampling to construct a small, weighted subset of representative cells that approximates the full dataset with provable geometric guarantees. This formulation transforms metacell construction into a coreset inference problem, reducing runtime and memory complexity by up to an order of magnitude without loss of accuracy. Through extensive experiments, we show that CAMP produces metacells that are compact, well-separated, and biologically coherent, achieving performance on par with or exceeding existing methods including MetaCell, SuperCell, SEACells, and MetaQ. By combining theoretical efficiency with empirical robustness, CAMP establishes coreset acceleration as a principled foundation for scalable, high-fidelity metacell inference in single-cell transcriptomics.

bioinformatics↗

Coreset-based logistic regression for atlas-scale cell type annotation

Logistic regression provides a versatile and interpretable framework for analyzing single-cell and spatial omics data. It has been shown to outperform more complex models on key association tasks, such as marker gene identification, and on predictive tasks, such as cell type annotation. Logistic regression has also been used effectively for quantifying spatial proximity between cell types and has demonstrated competitive performance in predicting tissue phenotypes. As increasingly large single-cell atlases become publicly available, training logistic regression models on these datasets offers opportunities for building accurate and robust reference models but also poses major computational challenges. Here, we adapt and extend the theory of coresets for logistic regression. Specifically, we compute a previously introduced classification complexity measure using linear programming to identify omics datasets that admit coresets, i.e. random subsets of cells that preserve logistic loss. Our main theoretical finding is that this complexity measure is close to one for PCA-transformed single-cell datasets. Moreover, we prove that logistic loss is preserved under this linear transformation, suggesting a universal sampling scheme that requires only a constant number of cells per cell type to obtain accurate representations of the data. Experiments across several cell atlases demonstrate that these theoretical guarantees translate to accurate and scalable transfer of cell types and spatial niches, outperforming more complex models, including pre-trained language models and graph isomorphism networks.

bioinformatics↗