Search bioRxiv⌕ Search

Biology subjects

Kodikara, S.

Publications and source records attributed to Kodikara, S..

3 recordsLinked to original sources

Semisynthetic Simulation for Microbiome Data Analysis

High-throughput sequencing data lie at the heart of modern microbiome research. Effective analysis of these data requires careful preprocessing, modeling, and interpretation to detect subtle signals and avoid spurious associations. In this review, we discuss how simulation can serve as a sandbox to test candidate approaches, creating a setting that mimics real data while providing ground truth. This is particularly valuable for power analysis, methods benchmarking, and reliability analysis. We explain the probability, multivariate analysis, and regression concepts behind modern simulators and how different implementations make trade-offs between generality, faithfulness, and controllability. Recognizing that all simulators only approximate reality, we review methods to evaluate how accurately they reflect key properties. We also present case studies demonstrating the value of simulation in differential abundance testing, dimensionality reduction, network analysis, and data integration. Code for these examples is available in an online tutorial (https://go.wisc.edu/8994yz) that can be easily adapted to new problem settings.

bioinformatics↗

Protection of the Telomeric Junction by the Shelterin Complex

Shelterin serves critical roles in suppressing superfluous DNA damage repair pathways on telomeres. The junction between double-stranded telomeric tracts (dsTEL) and single-stranded telomeric overhang (ssTEL) is the most accessible region of the telomeric DNA. The shelterin complex contains dsTEL and ssTEL binding proteins and can protect this junction by bridging between the ssTEL and dsTEL tracts. To test this possibility, we monitored shelterin binding to telomeric DNA substrates with varying ssTEL and dsTEL lengths and quantified its impact on telomere accessibility using single-molecule fluorescence microscopy methods in vitro. We identified the first dsTEL repeat nearest to the junction as the preferred binding site for creating the shelterin bridge. Shelterin requires at least two ssTEL repeats while the POT1 subunit of shelterin that binds to ssTEL requires longer ssTEL tracts for stable binding to telomeres and effective protection of the junction region. The ability of POT1 to protect the junction is significantly enhanced by the 5-phosphate at the junction. Collectively, our results show that shelterin enhances the binding stability of POT1 to ssTEL and provides more effective protection compared to POT1 alone by bridging single- and double-stranded telomeric tracts. Table of Content Graphic O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=109 SRC="FIGDIR/small/608453v1_ufig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@a58e79org.highwire.dtl.DTLVardef@12cb172org.highwire.dtl.DTLVardef@135bf40org.highwire.dtl.DTLVardef@19f3fe4_HPS_FORMAT_FIGEXP M_FIG C_FIG

biophysics↗

Microbial network inference for longitudinal microbiome studies with LUPINE

The microbiome is a complex ecosystem of interdependent taxa that has traditionally been studied through cross-sectional studies. However, longitudinal microbiome studies are becoming increasingly popular. These studies enable researchers to infer taxa associations towards the understanding of coexistence, competition, and collaboration between microbes across time. Traditional metrics for association analysis, such as correlation, are limited due to the data characteristics of microbiome data (sparse, compositional, multivariate). Several network inference methods have been proposed, but have been largely unexplored in a longitudinal setting. We introduce LUPINE (LongitUdinal modelling with Partial least squares regression for NEtwork inference), a novel approach that leverages on conditional independence and low-dimensional data representation. This method is specifically designed to handle scenarios with small sample sizes and small number of time points. LUPINE is the first method of its kind to infer microbial networks across time, while considering information from all past time points and is thus able to capture dynamic microbial interactions that evolve over time. We validate LUPINE and its variant, LUPINE single (for single time point analysis) in simulated data and four case studies, where we highlight LUPINEs ability to identify relevant taxa in each study context, across different experimental designs (mouse and human studies, with or without interventions, as short or long time courses). We propose different metrics to compare the inferred networks and detect changes in the networks across time, groups or in response to external disturbances. LUPINE is a simple yet innovative network inference methodology that is suitable for, but not limited to, analysing longitudinal microbiome data. The R code and data are publicly available for readers interested in applying these new methods to their studies.

bioinformatics↗