Search bioRxivSearch

FIND YOUR NEXT DISCOVERY

Results for “bioinformatics”

Original records, connected by a shared subject.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

33 records · Page 2Linked to original sources

TomatoPGFM: A graph-conditioned foundation model for tomato pangenomes

Most genomic foundation models are pretrained on independent linear assemblies and therefore do not explicitly represent population-level segment sharing or local graph connectivity. We developed TomatoPGFM, a graph-conditioned model pretrained on 54.65 Gb of sequence from 66 tomato (Solanum spp.) accessions. Sequence tokens were conditioned on pangenome node attributes and local adjacency, and the model was optimised using masked language modelling and graph-feature reconstruction. To evaluate model responses to graph-conditioned input, we compared aligned, shuffled and disabled graph inputs in 25,000 windows from the training panel. Sequence-aligned graph input produced lower masked language modelling loss than graph-off at all five curriculum stages in both training-panel strata, while the shuffled perturbation generally yielded intermediate losses. We then assessed sequence-only transfer in Solanum sitiens LA1974 and S. lycopersicum MicroTom, neither of which was used for graph construction or pretraining. Frozen-probe AUROC values for gene-versus-intergenic and coding-sequence-versus-intergenic classification ranged from 0.8489 to 0.9593. TomatoPGFM produced higher AUROC point estimates than DNABERT-2 in all four comparisons. Enabling the zero-feature GraphAdapter pathway with adjacency messaging disabled changed throughput by less than 1% at 512-2,048 positions under the tested configuration. Together, these results show that TomatoPGFM responds consistently to sequence-aligned pangenome context in training-panel sequences and provides informative sequence representations for genic-region classification in accessions excluded from graph construction and pretraining.

bioinformatics

The first OpenBind release: An open experimental structure-affinity dataset and benchmark for structure-based AI

High-quality experimental datasets that link protein-ligand structures with binding affinity data are essential for developing and evaluating structure-based machine learning methods. To help address this need, we established OpenBind as an open-science initiative to generate large-scale experimental datasets for structure-based AI and molecular discovery. Here, we describe the first public OpenBind release, which, to the best of our knowledge, is the largest public single-target experimental structure-affinity dataset. The dataset focuses on enteroviral 2A protease, comprising 925 crystallographic binding events from 699 compounds and associated affinity measurements for 601 compounds. It combines structures from an initial fragment screen and follow-on molecules, together with affinity data, linking experimentally determined protein-ligand binding modes to biophysical measurements within a coherent antiviral discovery campaign. We used this dataset to evaluate protein-ligand structure prediction, binding-affinity prediction, and virtual screening using representative structure-based methods, including docking and cofolding. This exposed several challenges that are central to practical structure-based modelling: docking performance depends strongly on binding-pocket conformation, poses are difficult to rank, and structure-based affinity prediction remains challenging. Fine-tuning OpenFold3-p2 on the fragment-screen structures substantially improved pose prediction and virtual screening for related follow-on compounds, demonstrating how early-stage experimental structures can support target-specific model adaptation.

bioinformatics

The DYNAM-O Toolbox: Characterizing Individualized Neural Signatures in Sleep EEG

Conventional sleep electroencephalography (EEG) measures often rely on predefined bands, thresholds, and averages that incompletely capture transient oscillatory dynamics across an entire night. Here, we introduce the Dynamic Oscillation (DYNAM-O) Toolbox, an open-source, cross-platform (MATLAB, Python, and Rust) software package for data-driven characterization of individualized neural dynamics in sleep EEG. DYNAM-O identifies transient oscillations as time-frequency peaks on multitaper spectrograms using a novel multi-resolution procedure, computes intrinsic and sleep-state-dependent extrinsic features for each event, and represents the overnight distributions of tens of thousands of TF-peaks as feature histograms spanning oscillation frequency, slow oscillation power, and slow oscillation phase. This distributional representation preserves continuous brain-state variation that could be obscured by averaging within conventional sleep stages. The toolbox further provides Gaussian and spline basis-based dimensionality reduction, visualization, and whole-histogram statistical testing tools to support both exploratory and hypothesis-driven analyses. To demonstrate its use for group-level inference, we analyzed overnight C3-channel EEG from 133 adults (71 females, 72 males; ages 20-35 years) in the Cleveland Family Study. Whole-histogram and parameterized-mode analyses reproduced the established higher center frequency of fast-spindle activity in females and additionally revealed greater low-alpha transient oscillatory activity in females, a pattern outside the conventional sleep spindle range. By completing the analysis cycle from TF-peak extraction to statistical inference, DYNAM-O provides an accessible and interpretable framework for studying individualized sleep physiology and identifying subtle, reproducible electrophysiological patterns.

bioinformatics

Sphingolipid metabolism-related genes as key regulatory hubs in white smoke inhalation induced lung injury

Objective White smoke inhalation injury (WSI) causes severe acute lung damage with no specific therapy currently available. Sphingolipid metabolism is implicated in pulmonary inflammation, but its transcriptional regulatory landscape in WSI remains unexplored. This study aimed to identify key sphingolipid metabolism related genes and evaluate their regulatory roles and therapeutic potential in WSI. Methods We established a rat model of WSI and performed integrated bulk RNA sequencing, weighted gene coexpression network analysis (WGCNA), and single-cell RNA sequencing (scRNAseq) to screen for differentially expressed sphingolipid metabolism-related genes (DESRGs). Protein-protein interaction (PPI) network with four centrality algorithms was used to prioritize hub genes. In silico gene knockout and molecular docking were conducted to assess regulatory functions and identify potential drug candidates. Results We identified 22 DESRGs that were predominantly enriched in DNA replication and cell cycle pathways rather than canonical sphingolipid metabolic processes. PPI consensus prioritized three hub genes--Top2a, Ttk, and Ccna2--with Top2a exhibiting the highest expression in epithelial cells and significant downregulation after smoke exposure. ScRNAseq revealed immune cell infiltration and epithelial differentiation trajectories. Virtual knockout showed that Top2a depletion affected the largest transcriptomic fraction (~0.4%) and was enriched in lysosome biogenesis, innate immunity, phagocytosis, and lipid catabolism. Molecular docking identified thalidomide as a high affinity ligand for Top2a (Vina score: -8.5 kcal/mol). Conclusion Our multiomics integrative framework identifies Top2a as a central regulatory hub linking sphingolipid associated inflammation to epithelial responses in WSI, and nominates thalidomide as a potential drug repurposing candidate. These findings provide prioritized targets for future translational investigation.

bioinformatics

From Public Archive to Reusable Resource: Characterizing Gut Microbiome Metadata in the NCBI SRA

Public sequencing repositories contain large amounts of gut microbiome data that could support cross-study comparison, reproducibility analysis, and microbiome foundation model development. However, the extent to which these data are structured, harmonized, and reusable at archive scale remains unclear. Here, we characterized publicly available gut microbiome sequencing metadata from the NCBI Sequence Read Archive using Google BigQuery, focusing on human gut metagenome, mouse gut metagenome, and broadly annotated gut metagenome records. We evaluated temporal growth, sequencing depth, BioSample and BioProject structure, platform and instrument use, metadata completeness, host attribution, publication linkage, and research themes from linked literature. Public gut microbiome data increased substantially over time and were dominated by human-associated datasets and Illumina sequencing platforms. Core technical metadata fields were highly complete, but biological context needed for reuse, including host identity, phenotype, study design, and disease status, was often inconsistently encoded or required recovery from BioSample attributes and linked publications. In the generic "gut metagenome" cohort, host identity could be assigned for only 13.00% of BioSamples, highlighting the limitations of broad organism annotations for automated cohort construction. Publication linkage was also incomplete at the archive level, although usable text was recovered for most linked publications. Topic modeling of SRA-linked literature showed persistent emphasis on core gut microbiota composition and increasing representation of human cohort and infant microbiome studies. Overall, these findings show that public gut microbiome data are extensive and technically rich but not uniformly analysis ready. Improved metadata harmonization, publication linkage, and biological context recovery will be necessary to support reliable large-scale reuse and AI-ready microbiome data resources.

bioinformatics

Constructing microbiome co-occurrence networks with confidence: A conditional, nonparametric, inference-based approach

Constructing microbial association networks is a common strategy for exploring relationships among taxa in microbiome studies. Although marginal correlation methods are easy to implement and allow formal inference, they can produce spurious edges driven by indirect associations through other taxa. Conditional graphical-modeling methods aim to recover direct associations, but many rely on Gaussian or linear assumptions and often provide limited uncertainty quantification. We propose a conditional, nonparametric approach based on the scaled expected conditional covariance (SEcov). SEcov measures population-level conditional association by residualizing each taxon with respect to the remaining taxa and scaling the resulting expected conditional covariance. The resulting estimator can incorporate flexible machine-learning methods for conditional-mean estimation and admits asymptotic normal inference, enabling p-values and confidence intervals for taxon-pair associations. We demonstrate through simulation studies that our proposed approach improves network recovery relative to other methods, and we illustrate the new method via construction of a co-occurrence network for the vaginal microbiome during pregnancy. IMPORTANCEHigh-throughput sequencing has made it possible to characterize microbial communities at large scale, and network analysis is widely used to summarize relationships among taxa. However, networks based on marginal correlations may include indirect associations, whereas many conditional graphical models rely on assumptions that may be difficult to justify for sparse, zero-inflated, compositional microbiome data. SEcov offers a practical alternative by estimating conditional associations nonparametrically and attaching inferential uncertainty to individual edges. This allows investigators to construct microbiome networks using statistically interpretable evidence for taxon-pair associations, rather than relying solely on arbitrary correlation cutoffs or regularization tuning parameters.

bioinformatics

An M-learner approach for heterogeneous mediation analysis with high-dimensional omics mediators

Causal mediation analysis is widely used to identify biological pathways linking exposures to outcomes, but most methods assume homogeneous mediation effects across individuals. In high-dimensional omics settings, this assumption can mask important heterogeneity driven by demographic, genetic, or environmental factors. We propose the M-high-learner, a flexible framework for detecting heterogeneous mediation effects with high-dimensional mediators. The method identifies mediators with subgroup-specific indirect effects while distinguishing them from null or homogeneous signals and controlling the type I error rate. It is computationally efficient, scalable, and yields interpretable sub-types. Simulation studies show that the proposed approach achieves high power while maintaining accurate error control. Applications to the Framingham Heart Study and the Multi-Ethnic Study of Atherosclerosis reveal that the mediation role of gene expression in sexs effect on high-density lipoprotein varies across subgroups defined by body mass index and age. Our framework provides a practical tool for uncovering heterogeneous biological mechanisms in high-dimensional genomic studies. Author SummaryBiological processes linking risk factors to disease often differ across individuals, but many existing methods assume these processes are the same for everyone. This can hide important differences between groups. We developed a powerful method to identify when these pathways vary across subgroups using large-scale molecular data. Our approach detects differences in how intermediate biological factors contribute to outcomes in populations defined by characteristics such as age and body mass index. Applying our method to population studies, we found that some biological pathways operate differently across groups, suggesting that key mechanisms may be missed when differences are ignored. Our work provides a tool to better understand how disease-related processes vary across individuals, which may support more targeted and personalized approaches to health research.

bioinformatics

Intelligent differential ion mobility spectrometry (iDMS): A deep neural network that predicts optimal space-resolved ion mobility parameters for isomeric monoglycosphingolipids

Simultaneous quantification of monoglycosphingolipid stereoisomers is required to monitor changes in defective enzymatic pathways linked to diseases such as Gaucher Disease, Parkinson's Disease, and Krabbe Disease. Resolution of beta-glucosyl and beta-galactosyl epimers cannot be achieved by standard liquid chromatography, electrospray ionization, tandem mass spectrometry (LC-ESI-MS/MS). Separation becomes possible when field asymmetric ion mobility spectrometry (FAIMS), also known as differential mobility mass spectrometry (DMS), is added as an orthogonal separation technique to LC. FAIMS/DMS separates epimeric ion clusters in a high versus low electric field (separation voltage, SV) then redirects the target epimeric ions to the mass spectrometer through the application of a direct current (compensation voltage, CoV). Resolving SVs and CoVs must be manually determined for each lipid. Manual derivation is a labour-intensive process that requires pure synthetic standards, limiting the number of stereoisomers a user can include in an assay. To address this problem, we introduce here intelligent DMS (iDMS). iDMS is an in silico supervised neural network model that learns the ion mobility relationships between SV and CoV and the monoglycosphingolipid structural features of sugar headgroup, N-acyl chain length, and N-acyl degree of unsaturation. iDMS predicts the SV and CoV combinations capable of resolving any stereoisomer pair from a training dataset of composed of measured signal intensities across a range of SVs and CoVs of 12 lipids. This machine learning alternative to manual DMS optimization promises to accelerate the deployment of multiple-reaction-monitoring mode (MRM) RPLC-ESI-DMS-MS/MS assays for the routine and rapid quantification of biologically relevant monoglycosphingolipid stereoisomers.

bioinformatics

BARCS: beta-binomial regression for multivariate CRISPR screen design

Pooled CRISPR screens increasingly use longitudinal, donor-adjusted, and factorial designs, but beta-binomial screen methods have largely remained limited to pairwise comparisons. BARCS extends the library-total-conditional beta-binomial model to guide-level regression with an arbitrary design matrix, enabling direct estimation of time, covariate, and interaction effects. In four replicate-complete Cas13 screens, adding the intermediate time point modestly improved essential-gene recovery. Applying the same non-targeting-control scaling rule to BARCS, MAGeCK-MLE, edgeR-QL, DESeq2, and limma--voom produced similar calibration across all five methods, while the four alternatives ranked essential genes more strongly than BARCS. In an ordered-bin IL2RA screen, donor-adjusted BARCS recovered more validated regulators with fewer total calls than the matched four-bin MAGeCK-MLE fit, and cross-fitted controls exposed excess guide-level significance. Simulations showed gains from dispersion moderation and control-based denominators, but seed-specific results exposed denominator sensitivity and a null grid localized substantial gene-level error to correlated-guide aggregation rather than dispersion alone. Aggregation-matched control scaling reduced but did not eliminate this error. An external audit prompted by concerns about beta-binomial false discoveries showed that the reported CB2 null-discovery count disappeared when full-library totals were restored. This corrected one denominator-dependent result but did not refute the broader calibration concern; nominal-level calibration remained unresolved. BARCS therefore contributes a multivariable extension of the library-total-conditional beta-binomial model together with an explicit account of where its inference is valid: guide-level coefficients are supported by independent biological libraries, whereas gene-level summaries and partitioned-bin designs require correlation-aware aggregation or joint modelling that the present implementation provides diagnostically rather than generatively. We report this boundary because complex pooled designs make it consequential, not because it is unique to the beta-binomial model.

bioinformatics

Interactive downstream proteomics analysis with MiraProt using Mueller cell proteomes from equine recurrent uveitis

Mass spectrometry-based proteomics requires downstream analysis of processed protein abundance data, including data inspection, filtering, statistical testing, functional enrichment, protein set comparison, network analysis, and visualization. MiraProt was developed as a modular, metadata-aware R Shiny platform that integrates these steps in a single interactive workflow for processed protein-level proteomics data. Its metadata-aware design enables identifiers, sample information, experimental conditions, transformations, and derived data columns to be defined during data preparation and reused consistently across downstream analyses. To demonstrate its use, we reanalyzed a previously published label-free proteomic dataset of primary retinal Mueller cells from healthy horses and horses with equine recurrent uveitis (ERU). ERU is a naturally occurring autoimmune eye disease of horses characterized by recurrent intraocular inflammation triggered by autoreactive T-cells. Mueller cells are specialized retinal macroglia with various functions such as maintaining retinal ion homeostasis and supporting retinal neuron metabolism. Of 193 proteins with an adjusted p-value [≤] 0.05, 187 also showed at least a twofold abundance difference between ERU-derived and control Mueller cells. Functional enrichment highlighted nuclear RNA processing, chromatin-associated structures, DNA and RNA binding, interferon responses, and cell-cycle-associated programs. Gene set enrichment analysis identified positive enrichment of Interferon Alpha Response, Interferon Gamma Response, and MYC-, E2F-, and G2M-associated gene sets. Network analysis of shared proteins further linked this signature to DNA replication, mitotic checkpoint control, and RNA processing. ERU-derived Mueller cells also showed increased abundance of MHC class II-associated proteins. Together, these findings identified an interferon-responsive, cell-cycle-associated, and MHC class II-associated Mueller cell protein signature in ERU and generated experimentally testable hypotheses for further mechanistic studies. MiraProt provides an accessible, metadata-aware framework for reproducible downstream exploration of processed proteomic datasets and prioritization of candidate proteins and pathways for experimental follow-up.

bioinformatics

RECON infers regions of interest from H&E images and reconstructs whole-slide molecular profiles at single-cell resolution

Spatial omics technologies resolve molecular expression and spatial architecture at single-cell resolution, but profiling whole slides remains costly. In practice, only a few regions of interest (ROIs) are profiled, leaving the rest of the tissue unmeasured. S2-omics was the first framework to unify ROI selection with out-of-ROI prediction, but it operates on superpixels rather than individual cells and predicts discrete cell types rather than continuous molecular profiles. Superpixel-based representations do not explicitly preserve cell boundaries, while categorical cell-type labels cannot quantify molecular expression within cells. Here we present RECON, a two-stage framework that performs ROI inference and whole-slide molecular reconstruction at single-cell resolution, predicting both continuous molecular profiles and discrete cell-type labels. In the first stage, RECON extracts morphological and microenvironmental features from individual cells to identify a representative ROI for spatially resolved single-cell molecular profiling. In the second stage, RECON trains deep learning models on molecular measurements acquired within the selected ROI and reconstructs transcriptomic or proteomic profiles for all remaining cells on the slide. Benchmarked against pathologist annotations, RECONs ROI selection outperforms the superpixel-based S2-omics approaches (IoU: 0.75 versus 0.64). For transcriptomics, refining the modeling unit from superpixels to single cells improves per-gene Pearson correlation by 22%. For proteomics, RECON surpasses the current state-of-the-art method, ROSIE, across all 16 markers, with a median per-cell Pearson correlation of 0.91 versus 0.84. Moreover, RECON delineates tumour boundaries and regions with distinct immune-cell densities, and highlights candidate tertiary lymphoid structures. Together, these results demonstrate that RECON enables informative ROI selection and whole-slide molecular reconstruction at single-cell resolution for both spatial transcriptomics and spatial proteomics.

bioinformatics

From Prompt to Provenance: BloClaw, a Capability-Gated AI4S Workstation for Auditable Computational Biology

Scientific agents can produce plausible answers while remaining unable to establish whether the computation behind an answer is executable, recoverable, or reproducible. We present BloClaw, an AI4S workstation built around a simple principle: a scientific agent should know what it can do, show how it did it, and state what remains unvalidated. Each capability declares an execution state, input constraints, dependencies, expected outputs, and scientific limitations. Natural-language requests are translated into structured tasks, validated against this registry, executed through scientific tools, and recorded in a provenance-aware Living Lab Notebook. The system is designed to detect invalid inputs, failed tool calls, missing dependencies, and remote timeouts, and to route them to repair, retry, or escalation. The implemented and tested scope comprises RDKit-based molecular property and rule screening, protein structure analysis, docking-pose inspection, 3D visualization, and structured reporting. We demonstrate the workflow on a PubChem-retrieved osimertinib structure and a supplied 6LU7 docking artifact: the former yields deterministic descriptors (molecular weight 499.619 Da, cLogP 4.5098, TPSA 87.55 A^2), while the latter contains 2,387 protein ATOM records, 309 residues, and nine pose records. These examples are workflow demonstrations, not efficacy or affinity studies. Beyond retrospective prediction, the manuscript specifies a prior-minimized constructive mode in which a desired function is compiled into explicit physical, chemical, and systems constraints, candidate mechanisms are simulated, and observations are reintroduced for calibration and falsification; this is a proposed extension rather than a result of the present case studies. We describe an evaluation protocol that compares BloClaw with a standard single-agent workflow and fixed-script execution using task completion, scientific correctness, recovery success, provenance completeness, reproducibility, human review time, latency, and cost. This manuscript reports the system design, verified capability boundary, deterministic software artifacts, and a reproducible evaluation protocol; it does not claim benchmark improvements before those experiments are run. BloClaw is an execution and accountability layer for AI-assisted research, complementing expert review and experimental validation rather than replacing them.

bioinformatics

High-Resolution Subtyping of Pediatric Low-Grade Glioma Using an Integrated Meta-Clustering Framework

Pediatric low-grade glioma (pLGG) is the most common type of brain tumor in children, accounting for approximately 30% of all central nervous system tumors in children. pLGG has multiple molecular subtypes that differ in disease progression, recurrence patterns, and treatment responses. Conventional wet lab approaches including molecular profiling and histopathological studies for pLGG characterization are time consuming, costly, and laborious. Recently, methods based on artificial intelligence (AI) or machine learning (ML) have been widely used for pLGG molecular categorization, but most of them can only identify two or three pLGG subtypes. To more comprehensively characterize the molecular subtypes of pLGG and their potential biological and therapeutic significance, we develop an integrated meta-clustering approach, namely Meta-pLGG, that can explore high resolution molecular subtypes and their transcriptional heterogeneity for pLGG. Specifically, we first performed multiple rounds of random projection (RP) to generate dimension-reduced feature vectors from pLGG transcriptomics data, each of which was subsequently clustered by different clustering algorithms including hierarchical clustering, K-means, Self-Organizing Maps (SOM), Non-negative Matrix Factorization (NMF), Gaussian Mixture Model (GMM), and Spectral Clustering, as base clustering methods. Then, to yield robust clustering performance, we integrated the clustering results of these RP based individual clustering algorithms by adopting a weighted meta-clustering (wMetaC) approach. Results based on 532 pLGG patients suggested that our proposed approach demonstrated superior stability and discriminative powers for higher resolution pLGG subtyping compared to conventional approaches. Based on consensus matrix analysis, we identified two major pLGG mega-subtypes, with one further subdivided into three subgroups and the other into two. Then, we performed cluster specific differential gene expression analysis, molecular pathway analysis, and gene-drug-disease association analysis. The results showed that the identified five subgroups exhibited significant subtype-specific transcriptomic heterogeneity. In summary, our meta-clustering approach demonstrated much higher performance and robustness in identifying higher resolution molecular subtypes of pLGG, revealing the molecular heterogeneity within pLGG and potentially providing new insights for more precise molecular subtyping and precision therapy.

bioinformatics

Spatial Transcriptomics Reveals Compartment-Specific Immune Activation Signatures in Ileal and Lymph Node Tissue in Treated HIV Infection

People with HIV (PWH) on long-term antiretroviral therapy (ART) continue to experience elevated rates of morbidities and mortality driven by persistent immune activation despite viral suppression. Known contributors include low-level HIV provirus activity, microbial translocation in part from epithelial barrier dysfunction, microbiome dysfunction, and co-infections. However, how these interact and where they predominate across tissue compartments remains incompletely defined. Here, we applied spatial transcriptomics to characterize compartment-specific transcriptional programs in ileum (epithelium, Peyer's patches, lamina propria) and inguinal lymph nodes (B Cell follicles and T cell zone) from ten PWH on long-term ART, stratified by CD4/CD8 ratio into low-ratio and high-ratio groups, with low-ratio as a proxy for immune activation and increased risk for non-AIDS related serious event. Comparison of global expression found significant differences between groups in four of five compartments. Differential expression analysis identified 483 differentially expressed genes across four of five compartments, with the greatest burden in the T-cell zone and none in the lamina propria. Gene set enrichment analysis identified 116 enriched pathways predominantly in the low-ratio group, spanning immune activation, infection-associated, and metabolic programs, with Peyer's patches showing the broadest transcriptional divergence of any compartment. Cross-compartment signals included higher expression of ORMDL3 and ARL17B in the low-ratio group implicating mitochondrial stress and inflammasome activation, lower expression of CCL3L3 and FCMR in the low-ratio group suggesting impaired immune execution, and divergent ribosomal protein programs between B-cell follicles and the T-cell zone. Cell deconvolution identified compartment-specific differences in estimated immune cell proportions, and T-cell zone gene expression showed significant associations with HIV reservoir measures and plasma markers of microbial translocation and immune activation. Together these findings support spatially heterogeneous immune activation as a feature of persistent immune dysregulation in treated HIV infection and provide compartment-resolved, hypothesis-generating evidence for the tissue-specific mechanisms driving inflammation in this population.

bioinformatics

Taxonomic classification cost tracks neither sequencing depth nor community richness at single-sample scale: a measured resource protocol for 16S rRNA amplicon pipelines

Marker-gene amplicon workflows are routinely run on shared compute, yet the cores, memory and wall time they are given are chosen by convention and not by measurement. We present a protocol for measuring them, applied to the two dominant stages of a QIIME 2 16S rRNA pipeline, DADA2 denoising and Naive Bayes taxonomic classification, across nine upper-respiratory samples from a paediatric otitis media cohort. The two stages do not consume the same input: denoising reads every sequence, classification only those surviving it. Subsampling one library across a 27-fold range of sequencing depth, denoising wall time rose 14.3-fold while classification changed by 1% and its peak memory not at all (3.11 GiB). Amplicon sequence variant (ASV) richness rose 2.8-fold over that range, so this is not richness saturating: the stage is dominated by a fixed per-invocation cost. Across a body-site gradient of 5 to 70 ASVs, denoising followed read count (exponent 0.75) while classification followed neither: a 5-ASV effusion and a 70-ASV adenoid community cost 40.81 s and 40.79 s. One ASV took 36.20 s and 218 took 37.27 s, 97% fixed cost. Thread-level parallelism offered little benefit. Denoising peaked at 1.18x near 8 threads and then declined; classification was slower at every setting above one job, consuming 10.5 times the CPU at 40. Representative sequences and their taxonomic assignments were identical at 1, 4 and 40 threads, so a reduced allocation changes what the analysis costs, not what it reports. Extending the query set to 10,000 sequences located two distinct boundaries: eight jobs first beat one at roughly 5,000 queries, and fitted fixed and per-query costs become equal at 15,248. Both lie roughly two orders of magnitude above the richest single sample measured. Practically: size denoising by read count, calibrate classification once against the reference in use, request one job for classification below a few thousand sequences, and take throughput from sample-level parallelism. Protocol, data and analysis code are released with the pipeline.

bioinformatics