Search bioRxivSearch

SEARCH · Search bioRxiv

Search Search bioRxiv

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

33 records · Page 2Linked to original sources

dnoise: Fast Native Data Reduction for Bruker timsTOF

Bruker timsTOF acquisitions produce dense native .d files whose storage, transfer, and archival become substantial at high throughput. We present dnoise, an open-source Rust tool that removes points directly from timsTOF frames and writes a native-compatible .d directory. dnoise retains ions that form coherent streaks across the ion-mobility dimension and applies acquisition-aware gates to signal that cannot be selected for fragmentation. On a three-species benchmark spanning ddaPASEF and diaPASEF at 5- and 15-minute gradients, default MS1-only denoising reduced the frame binary by 35 to 53%. Label-free quantification accuracy was preserved in both modes. ddaPASEF peptide-spectrum-match, peptide, and protein-group counts were unchanged, as expected with the searched MS/MS spectra untouched, and diaPASEF precursor and protein-group counts changed only slightly. Every tested processing run completed in 69 seconds or less on the benchmark workstation. Optional MS/MS denoising produced greater reduction but sacrificed several percent of identifications. Thus, a substantial fraction of native timsTOF frame data can be removed with little analytical change.

bioinformatics

XpBrew and PanXpresso - automatic RNA-seq processing workflow and comprehensive collection of gene expression data

Rapid developments in sequencing technologies have reduced the costs of transcriptomic experiments and resulted in a plethora of publicly available RNA-seq datasets. This is a valuable resource that can be harnessed to obtain novel biological insights through data upcycling. In this wake, we introduce XpBrew, an end-to-end Python workflow that was applied to generate PanXpresso, a comprehensive collection of gene expression datasets covering the taxonomic breadth of plants, animals, fungi, bacteria and archaea. XpBrew (https://github.com/PuckerLab/XpBrew) and PanXpresso (https://doi.org/10.60507/FK2/OBIGQH) are freely available.

bioinformatics

CyChat: a conversational Cytoscape app for no-code, reproducible network analysis

Network-based analyses of molecular interactions are useful for interpreting high-throughput omics data and identifying therapeutic targets. Cytoscape is the standard platform for these tasks, but users face a trade-off between accessible graphical workflows that are difficult to document and reproducible automation in Python or R that requires programming expertise. General-purpose coding assistants can generate Cytoscape Automation scripts, but remain external to Cytoscape. We present CyChat, a Cytoscape Desktop app that integrates a chat interface and a large language model (LLM) agent into the application. CyChat translates natural language into executable Cytoscape Automation workflows, runs generated Python code, and exports chat sessions with executed code as standalone Jupyter notebooks. To reduce setup barriers, CyChat includes an embedded Python runtime and supports both cloud-based and locally hosted LLMs. CyChat was evaluated across ten Cytoscape workflows using seven LLM providers, each represented by one LLM. The strongest configuration achieves a pass rate above 99%. In a qualitative evaluation based on a published network visualization, CyChat completes the task in 1.5-5 minutes, compared with 15-20 minutes for manual GUI workflows by computational biologists. CyChat is available through the Cytoscape App Store at https://apps.cytoscape.org/apps/cychat.

bioinformatics

Spatial Transcriptomics Reveals Compartment-Specific Immune Activation Signatures in Ileal and Lymph Node Tissue in Treated HIV Infection

People with HIV (PWH) on long-term antiretroviral therapy (ART) continue to experience elevated rates of morbidities and mortality driven by persistent immune activation despite viral suppression. Known contributors include low-level HIV provirus activity, microbial translocation in part from epithelial barrier dysfunction, microbiome dysfunction, and co-infections. However, how these interact and where they predominate across tissue compartments remains incompletely defined. Here, we applied spatial transcriptomics to characterize compartment-specific transcriptional programs in ileum (epithelium, Peyer's patches, lamina propria) and inguinal lymph nodes (B Cell follicles and T cell zone) from ten PWH on long-term ART, stratified by CD4/CD8 ratio into low-ratio and high-ratio groups, with low-ratio as a proxy for immune activation and increased risk for non-AIDS related serious event. Comparison of global expression found significant differences between groups in four of five compartments. Differential expression analysis identified 483 differentially expressed genes across four of five compartments, with the greatest burden in the T-cell zone and none in the lamina propria. Gene set enrichment analysis identified 116 enriched pathways predominantly in the low-ratio group, spanning immune activation, infection-associated, and metabolic programs, with Peyer's patches showing the broadest transcriptional divergence of any compartment. Cross-compartment signals included higher expression of ORMDL3 and ARL17B in the low-ratio group implicating mitochondrial stress and inflammasome activation, lower expression of CCL3L3 and FCMR in the low-ratio group suggesting impaired immune execution, and divergent ribosomal protein programs between B-cell follicles and the T-cell zone. Cell deconvolution identified compartment-specific differences in estimated immune cell proportions, and T-cell zone gene expression showed significant associations with HIV reservoir measures and plasma markers of microbial translocation and immune activation. Together these findings support spatially heterogeneous immune activation as a feature of persistent immune dysregulation in treated HIV infection and provide compartment-resolved, hypothesis-generating evidence for the tissue-specific mechanisms driving inflammation in this population.

bioinformatics

GNMCADS: Sampling For Protein Conformation Diversity With Gaussian Network Model Guided Condition Annealed Diffusion Sampler

Proteins are dynamic molecules existing in diverse conformational states underlying their biological functions. Although recent approaches have enabled diverse conformational sampling by emulating molecular dynamics simulations, perturbing evolutionary information, or steering internal mechanisms of structure prediction models, predicting conformations resulting from major domain motions or motions that occur over long timescales still remains a challenge. To this end, we introduce GNMCADS, a conformational sampling strategy that enhances the diversity of protein diffusion models by selectively annealing the conditioning signal guided by the intrinsic dynamical organization of the sampled protein. Further, we implement GNMCADS in the diffusion module of AlphaFold3, enabling the generation of diverse protein conformations. When benchmarked across 92 proteins that include 54 class A GPCRs, 15 transporters, and 23 proteins with major domain movements, GNMCADS exhibits improved sampling diversity compared to other current conformational sampling methods.

bioinformatics

Calibration-free compression brings Evo 2 to its full million-token context on a single GPU

Evo 2 is the largest openly available genomic foundation model, but its forty billion parameter configuration cannot be loaded onto a single 80 GB accelerator, placing genome-scale analysis beyond most laboratories. We present TurboQuant-Bio, an open toolkit that compresses Evo 2s weights and attention cache to four bits without calibration data, and serves both through fused kernels. Compression is near-lossless across perplexity spanning the tree of life, genomic classification, splice-site prediction, gene completion and clinically relevant variant-effect prediction. It brings Evo 2 40B onto one 80 GB GPU and Evo 2 7B to its full million-token context within a 40 GB memory budget, an eightfold gain in reachable context. We further show that the released chunked-prefill path is silently incorrect, returning plausible but uncorrelated likelihoods, and derive the block-wise continuation that repairs it: a complete 580-kilobase bacterial genome is now scored in one context in 22 minutes rather than 13.7 hours.

bioinformatics

Integrated Transcriptomic and CRISPR Dependency Analysis Prioritizes a CDK1-AURKB Mitotic Vulnerability Axis in Diffuse Intrinsic Pontine Glioma

Diffuse intrinsic pontine glioma (DIPG), now classified within diffuse midline glioma, H3K27-altered, remains a lethal pediatric brainstem tumor with limited therapeutic options. Here, we integrated public DIPG transcriptomic datasets, protein-protein interaction modeling, functional enrichment, immune deconvolution, survival analysis, and DepMap CRISPR dependency data to nominate candidate mitotic vulnerabilities. Differential expression analysis comparing 27 DIPG tumors with 6 brainstem low-grade glioma comparator samples identified a proliferative transcriptional program enriched for chromosome segregation, nuclear division, and cell-cycle pathways. Network analysis prioritized a compact mitotic hub module containing CDK1, AURKB, TOP2A, CDC20, CDCA8, and related G2/M regulators. CIBERSORT analysis of an independent DIPG cohort inferred low cytotoxic T-cell signal, consistent with an immune-cold phenotype, although immune-cell fractions require orthogonal validation. Survival analysis showed that neither inferred immune scores nor a composite mitotic hub score significantly stratified overall survival. DepMap CRISPR gene-effect data nominated CDK1, AURKB, TOP2A, and BIRC5 as candidate dependencies across brain tumor models. These findings provide a computational framework for prioritizing mitotic vulnerabilities in DIPG and support experimental validation in disease-relevant models.

bioinformatics

From Public Archive to Reusable Resource: Characterizing Gut Microbiome Metadata in the NCBI SRA

Public sequencing repositories contain large amounts of gut microbiome data that could support cross-study comparison, reproducibility analysis, and microbiome foundation model development. However, the extent to which these data are structured, harmonized, and reusable at archive scale remains unclear. Here, we characterized publicly available gut microbiome sequencing metadata from the NCBI Sequence Read Archive using Google BigQuery, focusing on human gut metagenome, mouse gut metagenome, and broadly annotated gut metagenome records. We evaluated temporal growth, sequencing depth, BioSample and BioProject structure, platform and instrument use, metadata completeness, host attribution, publication linkage, and research themes from linked literature. Public gut microbiome data increased substantially over time and were dominated by human-associated datasets and Illumina sequencing platforms. Core technical metadata fields were highly complete, but biological context needed for reuse, including host identity, phenotype, study design, and disease status, was often inconsistently encoded or required recovery from BioSample attributes and linked publications. In the generic "gut metagenome" cohort, host identity could be assigned for only 13.00% of BioSamples, highlighting the limitations of broad organism annotations for automated cohort construction. Publication linkage was also incomplete at the archive level, although usable text was recovered for most linked publications. Topic modeling of SRA-linked literature showed persistent emphasis on core gut microbiota composition and increasing representation of human cohort and infant microbiome studies. Overall, these findings show that public gut microbiome data are extensive and technically rich but not uniformly analysis ready. Improved metadata harmonization, publication linkage, and biological context recovery will be necessary to support reliable large-scale reuse and AI-ready microbiome data resources.

bioinformatics

Scaling recipes for single-cell RNA sequencing foundation models: when do scaling laws hold?

Deep learning models exhibit empirical scaling laws whereby performance changes predictably with model size, dataset size, and training compute. Although these relationships are well established in domains such as language and image modelling, their applicability to biological data remains unclear. Here, we investigate scaling behaviour in foundation models trained on large collec tions of single-cell transcriptomes. We show that pre-training loss decreases systematically with model capacity and training compute, exhibiting a power law dependence on model size. The strength and regularity of these trends differ between model formulations. We identify and quantify empirical relationships linking the optimal learning rate and depth-to-width ratio to model size and depth or compute. These results demonstrate that scaling principles extend to transcriptomic modelling. More broadly, they provide a quantitative framework for estimating the expected returns from additional resources and selecting suit able hyperparameters and architectures, thereby supporting the development of increasingly capable foundation models for omics data.

bioinformatics

TomatoPGFM: A graph-conditioned foundation model for tomato pangenomes

Most genomic foundation models are pretrained on independent linear assemblies and therefore do not explicitly represent population-level segment sharing or local graph connectivity. We developed TomatoPGFM, a graph-conditioned model pretrained on 54.65 Gb of sequence from 66 tomato (Solanum spp.) accessions. Sequence tokens were conditioned on pangenome node attributes and local adjacency, and the model was optimised using masked language modelling and graph-feature reconstruction. To evaluate model responses to graph-conditioned input, we compared aligned, shuffled and disabled graph inputs in 25,000 windows from the training panel. Sequence-aligned graph input produced lower masked language modelling loss than graph-off at all five curriculum stages in both training-panel strata, while the shuffled perturbation generally yielded intermediate losses. We then assessed sequence-only transfer in Solanum sitiens LA1974 and S. lycopersicum MicroTom, neither of which was used for graph construction or pretraining. Frozen-probe AUROC values for gene-versus-intergenic and coding-sequence-versus-intergenic classification ranged from 0.8489 to 0.9593. TomatoPGFM produced higher AUROC point estimates than DNABERT-2 in all four comparisons. Enabling the zero-feature GraphAdapter pathway with adjacency messaging disabled changed throughput by less than 1% at 512-2,048 positions under the tested configuration. Together, these results show that TomatoPGFM responds consistently to sequence-aligned pangenome context in training-panel sequences and provides informative sequence representations for genic-region classification in accessions excluded from graph construction and pretraining.

bioinformatics

BARCS: beta-binomial regression for multivariate CRISPR screen design

Pooled CRISPR screens increasingly use longitudinal, donor-adjusted, and factorial designs, but beta-binomial screen methods have largely remained limited to pairwise comparisons. BARCS extends the library-total-conditional beta-binomial model to guide-level regression with an arbitrary design matrix, enabling direct estimation of time, covariate, and interaction effects. In four replicate-complete Cas13 screens, adding the intermediate time point modestly improved essential-gene recovery. Applying the same non-targeting-control scaling rule to BARCS, MAGeCK-MLE, edgeR-QL, DESeq2, and limma--voom produced similar calibration across all five methods, while the four alternatives ranked essential genes more strongly than BARCS. In an ordered-bin IL2RA screen, donor-adjusted BARCS recovered more validated regulators with fewer total calls than the matched four-bin MAGeCK-MLE fit, and cross-fitted controls exposed excess guide-level significance. Simulations showed gains from dispersion moderation and control-based denominators, but seed-specific results exposed denominator sensitivity and a null grid localized substantial gene-level error to correlated-guide aggregation rather than dispersion alone. Aggregation-matched control scaling reduced but did not eliminate this error. An external audit prompted by concerns about beta-binomial false discoveries showed that the reported CB2 null-discovery count disappeared when full-library totals were restored. This corrected one denominator-dependent result but did not refute the broader calibration concern; nominal-level calibration remained unresolved. BARCS therefore contributes a multivariable extension of the library-total-conditional beta-binomial model together with an explicit account of where its inference is valid: guide-level coefficients are supported by independent biological libraries, whereas gene-level summaries and partitioned-bin designs require correlation-aware aggregation or joint modelling that the present implementation provides diagnostically rather than generatively. We report this boundary because complex pooled designs make it consequential, not because it is unique to the beta-binomial model.

bioinformatics

Designing antimicrobials with programmable mechanism and safety

Antimicrobial peptides (AMPs) are a promising solution to antimicrobial resistance, yet generative models for their design cannot control the physicochemical properties and motifs that shape activity and selectivity. Here, we present OmegAMP, a conditional diffusion framework controlling net charge, mean hydrophobicity, and sequence length, supporting de novo, analog, and motif-guided design. Across 204 wet-lab characterized peptides, de novo generation yielded antimicrobials with broad activity against multidrug-resistant Gram-negative isolates. Analog generation converted six inactive prototypes into antimicrobials, with the prototype determining each analog's membrane-disruption mode and mammalian-cell safety. Motif-guided analog generation preserved lipopolysaccharide engagement of active prototypes, and a redesigned non-antimicrobial leucine zipper acquired antimicrobial activity while retaining DNA-perturbing character in vitro. In murine skin and thigh infection models, leads reduced bacterial burden, with a motif-guided DNA-perturbing lead matching the fluoroquinolone control systemically. OmegAMP opens a programmable route to new peptide antibiotics whose mechanism and safety follow from the chosen prototype.

bioinformatics

Taxonomic classification cost tracks neither sequencing depth nor community richness at single-sample scale: a measured resource protocol for 16S rRNA amplicon pipelines

Marker-gene amplicon workflows are routinely run on shared compute, yet the cores, memory and wall time they are given are chosen by convention and not by measurement. We present a protocol for measuring them, applied to the two dominant stages of a QIIME 2 16S rRNA pipeline, DADA2 denoising and Naive Bayes taxonomic classification, across nine upper-respiratory samples from a paediatric otitis media cohort. The two stages do not consume the same input: denoising reads every sequence, classification only those surviving it. Subsampling one library across a 27-fold range of sequencing depth, denoising wall time rose 14.3-fold while classification changed by 1% and its peak memory not at all (3.11 GiB). Amplicon sequence variant (ASV) richness rose 2.8-fold over that range, so this is not richness saturating: the stage is dominated by a fixed per-invocation cost. Across a body-site gradient of 5 to 70 ASVs, denoising followed read count (exponent 0.75) while classification followed neither: a 5-ASV effusion and a 70-ASV adenoid community cost 40.81 s and 40.79 s. One ASV took 36.20 s and 218 took 37.27 s, 97% fixed cost. Thread-level parallelism offered little benefit. Denoising peaked at 1.18x near 8 threads and then declined; classification was slower at every setting above one job, consuming 10.5 times the CPU at 40. Representative sequences and their taxonomic assignments were identical at 1, 4 and 40 threads, so a reduced allocation changes what the analysis costs, not what it reports. Extending the query set to 10,000 sequences located two distinct boundaries: eight jobs first beat one at roughly 5,000 queries, and fitted fixed and per-query costs become equal at 15,248. Both lie roughly two orders of magnitude above the richest single sample measured. Practically: size denoising by read count, calibrate classification once against the reference in use, request one job for classification below a few thousand sequences, and take throughput from sample-level parallelism. Protocol, data and analysis code are released with the pipeline.

bioinformatics

Basophilic Erythroblast Emerges as the Key Turning Point in Polycythemia Vera

Abstract Polycythemia vera (PV) is a rare, chronic myeloproliferative neoplasm driven by the JAK2V617F mutation and characterized by uncontrolled erythroid proliferation. Although the mutation arises in hematopoietic stem cells, the differentiation stage at which its transcriptional consequences first become biologically meaningful has remained undefined. Using a multi-layer transcriptomics integration approach that combined differential gene expression, NicheNet ligand-receptor analysis, pseudotime trajectory inference, and CNV profiling on scRNA seq data, alongside bulk transcriptome validation, we identified basophilic erythroblasts as the critical transition point at which JAK2V617F shifts from a genomically present but transcriptionally silent state to an actively trajectory-altering and treatment-responsive disease driver. Differential expression revealed a qualitatively distinct disease signature at this stage, including ERFE-mediated iron dysregulation, MAP2K2-driven RAS/MAPK co-activation, and epigenetic reprogramming. NicheNet showed the establishment of a TGF{beta} superfamily and chemokine-driven niche-remodeling axis, and pseudotime analysis demonstrated that basophilic erythroblasts are the first erythroid population to exhibit condition-dependent trajectory divergence, whereas earlier progenitors showed none despite carrying the mutation. Interferon- treatment showed its broadest counterresponse at this stage but declined sharply thereafter, identifying basophilic erythroblasts as both the principal therapeutic target and the point of maximum vulnerability in PV.

bioinformatics

Model-based evaluation of Targeted-Antibacterial-Plasmids (TAPs) transfer kinetics and resensitization of pOXA-48 carbapenem-resistant Escherichia coli

Background Targeted-Antibacterial Plasmids (TAPs) are engineered mobile genetic elements that use bacterial conjugation to deliver selective CRISPR/Cas9 antibacterial activity against a specific target strain. Yet, the efficiency of TAPs is typically evaluated at a single time point, whereas the success of TAP-mediated resensitization critically depends on the dynamics of plasmid transfer and the complex interactions between bacterial subpopulations. This is the first study to evaluate the efficiency of a conjugation-based antibacterial approach at the subpopulation level, using an analytical framework analogous to that used for conventional antibiotics. Here, we investigate which process limits resensitization by TAPF-dCas9-OXA48: plasmid delivery, dCas9 activity, or the emergence of refractory and escape populations. Methods We fitted a mechanistic model of five interacting subpopulations (donors, recipients, transconjugants, escapers, and recusants) to 44 longitudinal conjugation experiments and used the fitted model to explore a range of biologically relevant scenarios. Results Using longitudinal conjugation data spanning 24 h, we show that up to 24% of recipients become recusants within 24h, refractory to further conjugation via entry exclusion, while secondary transconjugant emergence stays below 0.01%. Overall resensitization efficiency reaches up to 80%. Conclusion Plasmid transfer, rather than dCas9 repression, therefore appears to be the main bottleneck limiting the efficiency of TAPF-dCas9-OXA48 efficiency. These results identify plasmid delivery as a key engineering target for improving the performance of future TAPs.

bioinformatics