Search bioRxivSearch

FIND YOUR NEXT DISCOVERY

Search the archive

Original records, connected by a shared subject.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

33 recordsLinked to original sources

Interactive downstream proteomics analysis with MiraProt using Mueller cell proteomes from equine recurrent uveitis

Mass spectrometry-based proteomics requires downstream analysis of processed protein abundance data, including data inspection, filtering, statistical testing, functional enrichment, protein set comparison, network analysis, and visualization. MiraProt was developed as a modular, metadata-aware R Shiny platform that integrates these steps in a single interactive workflow for processed protein-level proteomics data. Its metadata-aware design enables identifiers, sample information, experimental conditions, transformations, and derived data columns to be defined during data preparation and reused consistently across downstream analyses. To demonstrate its use, we reanalyzed a previously published label-free proteomic dataset of primary retinal Mueller cells from healthy horses and horses with equine recurrent uveitis (ERU). ERU is a naturally occurring autoimmune eye disease of horses characterized by recurrent intraocular inflammation triggered by autoreactive T-cells. Mueller cells are specialized retinal macroglia with various functions such as maintaining retinal ion homeostasis and supporting retinal neuron metabolism. Of 193 proteins with an adjusted p-value [≤] 0.05, 187 also showed at least a twofold abundance difference between ERU-derived and control Mueller cells. Functional enrichment highlighted nuclear RNA processing, chromatin-associated structures, DNA and RNA binding, interferon responses, and cell-cycle-associated programs. Gene set enrichment analysis identified positive enrichment of Interferon Alpha Response, Interferon Gamma Response, and MYC-, E2F-, and G2M-associated gene sets. Network analysis of shared proteins further linked this signature to DNA replication, mitotic checkpoint control, and RNA processing. ERU-derived Mueller cells also showed increased abundance of MHC class II-associated proteins. Together, these findings identified an interferon-responsive, cell-cycle-associated, and MHC class II-associated Mueller cell protein signature in ERU and generated experimentally testable hypotheses for further mechanistic studies. MiraProt provides an accessible, metadata-aware framework for reproducible downstream exploration of processed proteomic datasets and prioritization of candidate proteins and pathways for experimental follow-up.

bioinformatics

High-Resolution Subtyping of Pediatric Low-Grade Glioma Using an Integrated Meta-Clustering Framework

Pediatric low-grade glioma (pLGG) is the most common type of brain tumor in children, accounting for approximately 30% of all central nervous system tumors in children. pLGG has multiple molecular subtypes that differ in disease progression, recurrence patterns, and treatment responses. Conventional wet lab approaches including molecular profiling and histopathological studies for pLGG characterization are time consuming, costly, and laborious. Recently, methods based on artificial intelligence (AI) or machine learning (ML) have been widely used for pLGG molecular categorization, but most of them can only identify two or three pLGG subtypes. To more comprehensively characterize the molecular subtypes of pLGG and their potential biological and therapeutic significance, we develop an integrated meta-clustering approach, namely Meta-pLGG, that can explore high resolution molecular subtypes and their transcriptional heterogeneity for pLGG. Specifically, we first performed multiple rounds of random projection (RP) to generate dimension-reduced feature vectors from pLGG transcriptomics data, each of which was subsequently clustered by different clustering algorithms including hierarchical clustering, K-means, Self-Organizing Maps (SOM), Non-negative Matrix Factorization (NMF), Gaussian Mixture Model (GMM), and Spectral Clustering, as base clustering methods. Then, to yield robust clustering performance, we integrated the clustering results of these RP based individual clustering algorithms by adopting a weighted meta-clustering (wMetaC) approach. Results based on 532 pLGG patients suggested that our proposed approach demonstrated superior stability and discriminative powers for higher resolution pLGG subtyping compared to conventional approaches. Based on consensus matrix analysis, we identified two major pLGG mega-subtypes, with one further subdivided into three subgroups and the other into two. Then, we performed cluster specific differential gene expression analysis, molecular pathway analysis, and gene-drug-disease association analysis. The results showed that the identified five subgroups exhibited significant subtype-specific transcriptomic heterogeneity. In summary, our meta-clustering approach demonstrated much higher performance and robustness in identifying higher resolution molecular subtypes of pLGG, revealing the molecular heterogeneity within pLGG and potentially providing new insights for more precise molecular subtyping and precision therapy.

bioinformatics

Transcriptomic profile of a rat jaw opener (anterior digastric) and a jaw closer (superficial masseter).

Mammalian skeletal muscle research predominantly focuses on locomotor muscles, and feeding related muscles remain less extensively characterized despite their role in mastication, mandibular stabilization, and swallowing. In this study, we investigated the transcriptomic specialization of three functionally and developmentally unique rat muscles: the anterior digastric (AD), a jaw opening muscle; the superficial masseter (SM), a jaw closing muscle; and the Sternohyoid (SH), a non-mandibular muscle involved in swallowing. Differential gene expression and weighted gene co-expression network analysis were used to characterize the transcription level features associated with their distinct roles. Our results indicated that all three muscles predominantly expressed fast-twitch contractile isoforms. However, the AD showed lower overall expression of several contractile gene families, including myosin heavy chain, myosin light chain, and tropomyosin isoforms, while exhibiting elevated expression of slow/oxidative myosin isoforms like Myh7 and Myh2. Network analysis revealed that modules correlated with AD are strongly enriched for fatty acid catabolism, mitochondrial energy production, and vascular/extracellular matrix remodeling. Additionally, AD and SM shared a distinct gene set compared to SH, highlighting their common developmental origin from the first branchial arch. Our findings show that the rat feeding related muscles possess unique transcriptomic profiles shaped by their contractile functions, developmental origins, and metabolic functions.

bioinformatics

AmPair: automating housekeeping-gene primer design for species-level metataxonomics

Amplicon sequencing of the 16S rRNA gene is the most widely used approach for profiling bacterial communities, but its taxonomic resolution is typically limited to the genus level. Many species carry multiple divergent 16S rRNA alleles that overlap across species boundaries, an ambiguity that even full-length, long-read sequencing cannot fully resolve. Shotgun metagenomics achieves species-level resolution but remains costly, particularly when only a single genus is of interest. Amplicon sequencing of rapidly evolving, protein-coding housekeeping genes offers a cost-effective alternative, yet no tool exists to identify suitable primer sets for a given target taxon. Here we present AmPair, a Snakemake pipeline that, given a target genus and one or more candidate housekeeping genes, designs and ranks primer pairs binding conserved regions while flanking a variable region capable of species-level discrimination, and validates them in silico across all available genomes. Using the genus Bacillus and the housekeeping gene tuf as a case study, the primer set recommended by AmPair amplified 99% of 2,392 genomes; only 0.04% carried multiple alleles and none showed inter-species allele overlap, compared with 91.41% and 69.49%, respectively, for the standard 16S rRNA V1-V9 region. Applied to a Bacillus community profiled by Nanopore sequencing, the same primers resolved closely related species. AmPair thus offers a generalizable and accessible route to species-level community profiling.

bioinformatics

An M-learner approach for heterogeneous mediation analysis with high-dimensional omics mediators

Causal mediation analysis is widely used to identify biological pathways linking exposures to outcomes, but most methods assume homogeneous mediation effects across individuals. In high-dimensional omics settings, this assumption can mask important heterogeneity driven by demographic, genetic, or environmental factors. We propose the M-high-learner, a flexible framework for detecting heterogeneous mediation effects with high-dimensional mediators. The method identifies mediators with subgroup-specific indirect effects while distinguishing them from null or homogeneous signals and controlling the type I error rate. It is computationally efficient, scalable, and yields interpretable sub-types. Simulation studies show that the proposed approach achieves high power while maintaining accurate error control. Applications to the Framingham Heart Study and the Multi-Ethnic Study of Atherosclerosis reveal that the mediation role of gene expression in sexs effect on high-density lipoprotein varies across subgroups defined by body mass index and age. Our framework provides a practical tool for uncovering heterogeneous biological mechanisms in high-dimensional genomic studies. Author SummaryBiological processes linking risk factors to disease often differ across individuals, but many existing methods assume these processes are the same for everyone. This can hide important differences between groups. We developed a powerful method to identify when these pathways vary across subgroups using large-scale molecular data. Our approach detects differences in how intermediate biological factors contribute to outcomes in populations defined by characteristics such as age and body mass index. Applying our method to population studies, we found that some biological pathways operate differently across groups, suggesting that key mechanisms may be missed when differences are ignored. Our work provides a tool to better understand how disease-related processes vary across individuals, which may support more targeted and personalized approaches to health research.

bioinformatics

RECON infers regions of interest from H&E images and reconstructs whole-slide molecular profiles at single-cell resolution

Spatial omics technologies resolve molecular expression and spatial architecture at single-cell resolution, but profiling whole slides remains costly. In practice, only a few regions of interest (ROIs) are profiled, leaving the rest of the tissue unmeasured. S2-omics was the first framework to unify ROI selection with out-of-ROI prediction, but it operates on superpixels rather than individual cells and predicts discrete cell types rather than continuous molecular profiles. Superpixel-based representations do not explicitly preserve cell boundaries, while categorical cell-type labels cannot quantify molecular expression within cells. Here we present RECON, a two-stage framework that performs ROI inference and whole-slide molecular reconstruction at single-cell resolution, predicting both continuous molecular profiles and discrete cell-type labels. In the first stage, RECON extracts morphological and microenvironmental features from individual cells to identify a representative ROI for spatially resolved single-cell molecular profiling. In the second stage, RECON trains deep learning models on molecular measurements acquired within the selected ROI and reconstructs transcriptomic or proteomic profiles for all remaining cells on the slide. Benchmarked against pathologist annotations, RECONs ROI selection outperforms the superpixel-based S2-omics approaches (IoU: 0.75 versus 0.64). For transcriptomics, refining the modeling unit from superpixels to single cells improves per-gene Pearson correlation by 22%. For proteomics, RECON surpasses the current state-of-the-art method, ROSIE, across all 16 markers, with a median per-cell Pearson correlation of 0.91 versus 0.84. Moreover, RECON delineates tumour boundaries and regions with distinct immune-cell densities, and highlights candidate tertiary lymphoid structures. Together, these results demonstrate that RECON enables informative ROI selection and whole-slide molecular reconstruction at single-cell resolution for both spatial transcriptomics and spatial proteomics.

bioinformatics

Automatic bioinformatic software named entity recognition from literature

Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.

bioinformatics

PathFold: Predicting the Entire Protein Folding Pathway from Protein Sequence Alone

Recent advances in protein structure prediction, exemplified by AlphaFold, have largely addressed the determination of static structures, one aspect of the protein folding problem. However, predicting folding pathways, by which proteins reach their native states, remains a significant challenge. Here, we present PathFold, a deep learning framework that predicts protein folding pathways directly from sequence information. PathFold leverages an AlphaFold-based module to extract structural information from the sequence and generates a progressive folding trajectory from an extended conformation using a diffusion model. By modeling the full trajectory, it enables prediction of folding intermediates and transition pathways, analogous to those observed in steered molecular dynamics (SMD) simulations. The predicted pathways reveal well-defined intermediates and sequential folding events, and show agreement with experimental folding data, including measured {Phi}-values.

bioinformatics

Intelligent differential ion mobility spectrometry (iDMS): A deep neural network that predicts optimal space-resolved ion mobility parameters for isomeric monoglycosphingolipids

Simultaneous quantification of monoglycosphingolipid stereoisomers is required to monitor changes in defective enzymatic pathways linked to diseases such as Gaucher Disease, Parkinson's Disease, and Krabbe Disease. Resolution of beta-glucosyl and beta-galactosyl epimers cannot be achieved by standard liquid chromatography, electrospray ionization, tandem mass spectrometry (LC-ESI-MS/MS). Separation becomes possible when field asymmetric ion mobility spectrometry (FAIMS), also known as differential mobility mass spectrometry (DMS), is added as an orthogonal separation technique to LC. FAIMS/DMS separates epimeric ion clusters in a high versus low electric field (separation voltage, SV) then redirects the target epimeric ions to the mass spectrometer through the application of a direct current (compensation voltage, CoV). Resolving SVs and CoVs must be manually determined for each lipid. Manual derivation is a labour-intensive process that requires pure synthetic standards, limiting the number of stereoisomers a user can include in an assay. To address this problem, we introduce here intelligent DMS (iDMS). iDMS is an in silico supervised neural network model that learns the ion mobility relationships between SV and CoV and the monoglycosphingolipid structural features of sugar headgroup, N-acyl chain length, and N-acyl degree of unsaturation. iDMS predicts the SV and CoV combinations capable of resolving any stereoisomer pair from a training dataset of composed of measured signal intensities across a range of SVs and CoVs of 12 lipids. This machine learning alternative to manual DMS optimization promises to accelerate the deployment of multiple-reaction-monitoring mode (MRM) RPLC-ESI-DMS-MS/MS assays for the routine and rapid quantification of biologically relevant monoglycosphingolipid stereoisomers.

bioinformatics

The DYNAM-O Toolbox: Characterizing Individualized Neural Signatures in Sleep EEG

Conventional sleep electroencephalography (EEG) measures often rely on predefined bands, thresholds, and averages that incompletely capture transient oscillatory dynamics across an entire night. Here, we introduce the Dynamic Oscillation (DYNAM-O) Toolbox, an open-source, cross-platform (MATLAB, Python, and Rust) software package for data-driven characterization of individualized neural dynamics in sleep EEG. DYNAM-O identifies transient oscillations as time-frequency peaks on multitaper spectrograms using a novel multi-resolution procedure, computes intrinsic and sleep-state-dependent extrinsic features for each event, and represents the overnight distributions of tens of thousands of TF-peaks as feature histograms spanning oscillation frequency, slow oscillation power, and slow oscillation phase. This distributional representation preserves continuous brain-state variation that could be obscured by averaging within conventional sleep stages. The toolbox further provides Gaussian and spline basis-based dimensionality reduction, visualization, and whole-histogram statistical testing tools to support both exploratory and hypothesis-driven analyses. To demonstrate its use for group-level inference, we analyzed overnight C3-channel EEG from 133 adults (71 females, 72 males; ages 20-35 years) in the Cleveland Family Study. Whole-histogram and parameterized-mode analyses reproduced the established higher center frequency of fast-spindle activity in females and additionally revealed greater low-alpha transient oscillatory activity in females, a pattern outside the conventional sleep spindle range. By completing the analysis cycle from TF-peak extraction to statistical inference, DYNAM-O provides an accessible and interpretable framework for studying individualized sleep physiology and identifying subtle, reproducible electrophysiological patterns.

bioinformatics

Sphingolipid metabolism-related genes as key regulatory hubs in white smoke inhalation induced lung injury

Objective White smoke inhalation injury (WSI) causes severe acute lung damage with no specific therapy currently available. Sphingolipid metabolism is implicated in pulmonary inflammation, but its transcriptional regulatory landscape in WSI remains unexplored. This study aimed to identify key sphingolipid metabolism related genes and evaluate their regulatory roles and therapeutic potential in WSI. Methods We established a rat model of WSI and performed integrated bulk RNA sequencing, weighted gene coexpression network analysis (WGCNA), and single-cell RNA sequencing (scRNAseq) to screen for differentially expressed sphingolipid metabolism-related genes (DESRGs). Protein-protein interaction (PPI) network with four centrality algorithms was used to prioritize hub genes. In silico gene knockout and molecular docking were conducted to assess regulatory functions and identify potential drug candidates. Results We identified 22 DESRGs that were predominantly enriched in DNA replication and cell cycle pathways rather than canonical sphingolipid metabolic processes. PPI consensus prioritized three hub genes--Top2a, Ttk, and Ccna2--with Top2a exhibiting the highest expression in epithelial cells and significant downregulation after smoke exposure. ScRNAseq revealed immune cell infiltration and epithelial differentiation trajectories. Virtual knockout showed that Top2a depletion affected the largest transcriptomic fraction (~0.4%) and was enriched in lysosome biogenesis, innate immunity, phagocytosis, and lipid catabolism. Molecular docking identified thalidomide as a high affinity ligand for Top2a (Vina score: -8.5 kcal/mol). Conclusion Our multiomics integrative framework identifies Top2a as a central regulatory hub linking sphingolipid associated inflammation to epithelial responses in WSI, and nominates thalidomide as a potential drug repurposing candidate. These findings provide prioritized targets for future translational investigation.

bioinformatics

From Prompt to Provenance: BloClaw, a Capability-Gated AI4S Workstation for Auditable Computational Biology

Scientific agents can produce plausible answers while remaining unable to establish whether the computation behind an answer is executable, recoverable, or reproducible. We present BloClaw, an AI4S workstation built around a simple principle: a scientific agent should know what it can do, show how it did it, and state what remains unvalidated. Each capability declares an execution state, input constraints, dependencies, expected outputs, and scientific limitations. Natural-language requests are translated into structured tasks, validated against this registry, executed through scientific tools, and recorded in a provenance-aware Living Lab Notebook. The system is designed to detect invalid inputs, failed tool calls, missing dependencies, and remote timeouts, and to route them to repair, retry, or escalation. The implemented and tested scope comprises RDKit-based molecular property and rule screening, protein structure analysis, docking-pose inspection, 3D visualization, and structured reporting. We demonstrate the workflow on a PubChem-retrieved osimertinib structure and a supplied 6LU7 docking artifact: the former yields deterministic descriptors (molecular weight 499.619 Da, cLogP 4.5098, TPSA 87.55 A^2), while the latter contains 2,387 protein ATOM records, 309 residues, and nine pose records. These examples are workflow demonstrations, not efficacy or affinity studies. Beyond retrospective prediction, the manuscript specifies a prior-minimized constructive mode in which a desired function is compiled into explicit physical, chemical, and systems constraints, candidate mechanisms are simulated, and observations are reintroduced for calibration and falsification; this is a proposed extension rather than a result of the present case studies. We describe an evaluation protocol that compares BloClaw with a standard single-agent workflow and fixed-script execution using task completion, scientific correctness, recovery success, provenance completeness, reproducibility, human review time, latency, and cost. This manuscript reports the system design, verified capability boundary, deterministic software artifacts, and a reproducible evaluation protocol; it does not claim benchmark improvements before those experiments are run. BloClaw is an execution and accountability layer for AI-assisted research, complementing expert review and experimental validation rather than replacing them.

bioinformatics

AVOCODO: An open-source multimodal annotation platform for developmental EEG

Behavioral annotation of synchronized video recordings is an essential step in developmental electroencephalography (EEG) research, supporting both the identification of behavior-related artifacts and the investigation of brain-behavior relationships. Existing annotation workflows, however, are often fragmented: proprietary EEG software provides limited flexibility for behavioral coding, whereas dedicated behavioral annotation platforms typically lack native integration with EEG data. We developed AVOCODO (Audio/VideO CODing Optimization), an open-source MATLAB-based software platform that integrates synchronized behavioral annotation directly into the EEG workflow. AVOCODO reads native EGI MFF recordings, synchronizes embedded video with EEG, visualizes the audio spectrogram to facilitate precise annotation of vocalizations, and writes user-defined behavioral events directly back into the original MFF recording as native EEG event markers while simultaneously exporting annotations as CSV files. The software supports fully customizable behavioral coding schemes, optional EEG visualization for quality control, and reloading of previously annotated recordings for review and inter-rater verification. Since its initial development in 2024, AVOCODO has been applied internally across five developmental EEG studies involving approximately 500 pediatric participants and more than 3,000 EEG recordings. By bridging behavioral annotation and EEG preprocessing within a unified open-source workflow, AVOCODO has the potential to improve the efficiency, reproducibility, and scalability of behavioral annotation in developmental EEG research.

bioinformatics

Rclade: automated taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R

Background: Reproducible taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R often require coordination among several packages and repeated code for label parsing, clade validation, plotting, and export. Workflow-managed analyses additionally benefit from non-interactive configuration, predictable diagnostics, and machine-readable exit status. Results: We present Rclade, an R package that consolidates the multi-package coordination required for taxonomic collapsing into a streamlined, single-function interface. Rclade provides (1) custom ggproto objects (GeomPolygonStraight/GeomSegmentStraight) that bypass coord_munch() interpolation to achieve straight-edge rendering of collapsed triangles in circular layouts; (2) automatic detection and parsing of four taxonomic-label formats (GTDB, Silva, NCBI, embedded) plus user-supplied custom regex, with explicit input-validation contracts and parsing-accuracy evaluation on real and derived test sets; and (3) workflow embeddability through YAML configuration, library-mode APIs, and standard Unix exit codes. Benchmarks on synthetic and real datasets (200-10,000 synthetic tips and real reference trees up to 10,122 tips; 5 replicates at every scale under a unified fully rendered measurement protocol) show that the full-pipeline overhead is modest for interactive use (median {approx}0.87 s in-session rendering and {approx}8.4 s process-level wall-clock at 10,000 tips). Conclusions: Rclade is a convenience layer over the ggtree/deeptime ecosystem that reduces boilerplate while adding targeted technical improvements for circular-layout rendering and format heterogeneity management.

bioinformatics

Constructing microbiome co-occurrence networks with confidence: A conditional, nonparametric, inference-based approach

Constructing microbial association networks is a common strategy for exploring relationships among taxa in microbiome studies. Although marginal correlation methods are easy to implement and allow formal inference, they can produce spurious edges driven by indirect associations through other taxa. Conditional graphical-modeling methods aim to recover direct associations, but many rely on Gaussian or linear assumptions and often provide limited uncertainty quantification. We propose a conditional, nonparametric approach based on the scaled expected conditional covariance (SEcov). SEcov measures population-level conditional association by residualizing each taxon with respect to the remaining taxa and scaling the resulting expected conditional covariance. The resulting estimator can incorporate flexible machine-learning methods for conditional-mean estimation and admits asymptotic normal inference, enabling p-values and confidence intervals for taxon-pair associations. We demonstrate through simulation studies that our proposed approach improves network recovery relative to other methods, and we illustrate the new method via construction of a co-occurrence network for the vaginal microbiome during pregnancy. IMPORTANCEHigh-throughput sequencing has made it possible to characterize microbial communities at large scale, and network analysis is widely used to summarize relationships among taxa. However, networks based on marginal correlations may include indirect associations, whereas many conditional graphical models rely on assumptions that may be difficult to justify for sparse, zero-inflated, compositional microbiome data. SEcov offers a practical alternative by estimating conditional associations nonparametrically and attaching inferential uncertainty to individual edges. This allows investigators to construct microbiome networks using statistically interpretable evidence for taxon-pair associations, rather than relying solely on arbitrary correlation cutoffs or regularization tuning parameters.

bioinformatics

Mural-VISTA: A tool for mural cell-vessel interaction assessment and multiscale single-cell topo-morphological analysis

Three-dimensional (3D) mural cell morphology is heterogeneous and coupled to vessel geometry, however, measurements from two-dimensional (2D) maximum intensity projections (MIP) obscure overlapping processes and cell-vessel contacts. Accordingly, we developed Mural-VISTA, a semi-automated Python workflow for mural cell-vessel interaction and single-cell topo-morphology analysis of reconstructed surface meshes. This workflow integrates mesh pretreatment, interactive centerline extraction, hierarchical segmentation of cell soma, main axis and secondary processes (branches), and extraction of 36 multiscale (cell process segment level, process level, and whole cell level) topo-morphological and vessel-referenced metrics. Mural-VISTA identified morphological changes in pericytes and vascular smooth muscle cells (vSMCs) with altered RhoA activity. Constitutive active RhoA (RhoA CA) over-expression reduced branch complexity and increased process alignment in both cell types, while increased whole-cell and branch solidity only in vSMCs. Dominant negative RhoA (RhoA DN) over-expression increased branch abundance and reduced branch solidity in pericytes but not vSMCs, suggesting cell-type specific effect of reduced RhoA activity. In conclusion, Mural-VISTA enables quantitative 3D profiling of mural cell architecture and its spatial relationship with the vessel.

bioinformatics

EXTRARNAS: A Framework for Extracting RNA Structures with Multiple Tools

Accurate annotation of RNA base-pairing interactions is essential for structural analysis, benchmarking, and data-driven RNA structure prediction. Several tools can extract RNA interactions from three-dimensional coordinates, but their outputs are heterogeneous and may disagree, particularly for non-canonical base pairs. We present EXTRARNAS, a Java-based framework for automated, reproducible, and user-friendly large-scale extraction of RNA structural annotations with multiple tools. EXTRARNAS processes batches of RNA structures specified by PDB identifier and chain, or provided as local PDB files, executes annotation tools through a Docker-based environment, and parses tool-specific outputs using ANTLR4-based grammars. For each structure-tool pair, the framework generates standard BPSEQ files for canonical cis Watson-Crick interactions and introduces BPSEQE, a standardized text format for representing the extended secondary structure, preserving canonical, non-canonical, and multiple interactions per nucleotide. The current prototype supports RNAView, MC-Annotate, and RNAPolis Annotator. We demonstrate EXTRARNAS on eight RNA structures containing triple-helix motifs, comparing extracted canonical pairs against curated BPSEQ references and evaluating the recovery of manually validated Hoogsteen interactions. The results show consistent differences among tools, especially for non-canonical interactions, highlighting the need for standardized representations such as BPSEQE to support reproducible comparison and future consensus-based annotation.

bioinformatics

The first OpenBind release: An open experimental structure-affinity dataset and benchmark for structure-based AI

High-quality experimental datasets that link protein-ligand structures with binding affinity data are essential for developing and evaluating structure-based machine learning methods. To help address this need, we established OpenBind as an open-science initiative to generate large-scale experimental datasets for structure-based AI and molecular discovery. Here, we describe the first public OpenBind release, which, to the best of our knowledge, is the largest public single-target experimental structure-affinity dataset. The dataset focuses on enteroviral 2A protease, comprising 925 crystallographic binding events from 699 compounds and associated affinity measurements for 601 compounds. It combines structures from an initial fragment screen and follow-on molecules, together with affinity data, linking experimentally determined protein-ligand binding modes to biophysical measurements within a coherent antiviral discovery campaign. We used this dataset to evaluate protein-ligand structure prediction, binding-affinity prediction, and virtual screening using representative structure-based methods, including docking and cofolding. This exposed several challenges that are central to practical structure-based modelling: docking performance depends strongly on binding-pocket conformation, poses are difficult to rank, and structure-based affinity prediction remains challenging. Fine-tuning OpenFold3-p2 on the fragment-screen structures substantially improved pose prediction and virtual screening for related follow-on compounds, demonstrating how early-stage experimental structures can support target-specific model adaptation.

bioinformatics