Search bioRxiv⌕ Search

Biology subjects

Taada, A.

Publications and source records attributed to Taada, A..

2 recordsLinked to original sources

MurineCyto-Det: A High-Resolution Murine BALF Cytology Dataset for Leukocyte Segmentation and Detection

Automated analysis of murine bronchoalveolar lavage fluid (BALF) cytology is important for preclinical respiratory research, yet progress has been limited by the lack of publicly available, well-annotated mouse BALF image datasets. We present MurineCyto-Det, a high-resolution murine BALF cytology dataset comprising 333 image tiles of size 1024x1024 pixels, annotated across five cytological categories with both pixel-level segmentation masks and one-to-one matched bounding boxes. The dataset contains 14,551 annotated cell instances and supports two complementary analysis tasks: morphology-oriented cell segmentation and object-level cell detection. To establish reproducible benchmark baselines, we evaluated representative segmentation and detection models. The results demonstrate the practical utility of MurineCyto-Det while highlighting realistic challenges arising from class imbalance, small object size, irregular cell morphology, and ambiguous debris-like structures. MurineCyto-Det provides a standardized resource for developing, evaluating, and comparing automated methods for murine BALF cytology analysis. The dataset is publicly available at https://doi.org/10.5281/zenodo.17608677.

bioinformatics↗

A Systematic Approach Toward Implementing Machine Learning Techniques to Analyze Gut Microbiome Data

This study investigates the relationship between the gut microbiota and specific diseases. Data was collected from the Human Gut Microbiome Atlas, which examines regional variations across 20 countries on five continents, categorizing microbial species by taxonomy, from genus to species. The Atlas provides color-coded phylum classifications, numerical species counts within the same genus, and an analysis of dysbiosis-related associations with 23 diseases, as well as region-enriched species. The data stratified samples into distinct categories such as westernized, non-westernized, cancerous, and non-cancerous. The findings demonstrate that tree-based ensemble methods, such as Bagging and Boosting prediction methods, achieved the highest accuracies across all categories due to their robustness in handling the complex, high-dimensional data. The XGBoost model yielded the strongest predictive performance, achieving 91% accuracy for westernized cancer-associated samples, 84% accuracy for non-westernized cancer-associated samples, 92% accuracy for westernized samples, and 78% for non-westernized samples. Additionally, advanced topological data analysis was used to assess the global structure and underlying patterns within the dataset. ImportanceThis research aims to connect gut microbiome composition to diseases using global datasets from the Human Gut Microbiome Atlas. The goal was to evaluate how accurately different machine learning algorithms could classify microbiota species and diseases and predict disease associations by comparing westernized and non-westernized populations, including both cancerous and noncancerous groups. These findings can contribute to the future creation of population-specific and disease-specific microbial models.

bioinformatics↗