Search bioRxiv⌕ Search

Biology subjects

Onawole, A.

Publications and source records attributed to Onawole, A..

3 recordsLinked to original sources

Trust-Aware Sequence-to-Function Modelling in Regulatory Genomics

Objective: Sequence-to-function models increasingly predict regulatory activity, such as chromatin accessibility, directly from DNA sequence, and are used to interpret non-coding genetic variation. Standard accuracy metrics, computed over a held-out set of genomic regions, do not establish whether an individual prediction remains reliable once the input sequence departs from that set, nor whether a model's attribution-based explanation is biologically grounded rather than coincidental. We develop and evaluate RegTrust-XAI, a trust-aware framework separating these questions using three inference-time signals: ensemble consensus, motif-grounded attribution coherence, and applicability-domain distance. Methods: A five-model convolutional ensemble was trained on 517,790 K562 ATAC-seq windows and evaluated on a held-out chromosome test set (chr8/chr9, n = 42,844). Consensus, coherence, and applicability-domain distance were each tested against prediction error, alongside complementary sequence-novelty analyses and validation against an independent lentiMPRA reporter assay and saturation-mutagenesis MPRA data at the PKLR promoter. Results: The ensemble reached Spearman {rho} = 0.782, with skill of 0.328 over a constant-value null predictor. High-consensus predictions (Scenarios A+B) were consistently enriched for lower error than low-consensus predictions (Scenarios C+D), and attribution coherence further separated error within the high-consensus population (mean absolute error 0.396 versus 0.435, p = 9.6e-10). Applicability-domain distance showed a monotonic error gradient across six distance bands. A 4-mer composition-divergence metric was negatively associated with error and anti-correlated with applicability-domain distance, so composition-based and model-relevant novelty are not equivalent. Attribution transfer to lentiMPRA was assay- and subgroup-dependent, and predicted allele-substitution effects correlated with measured saturation-mutagenesis effects at the PKLR promoter at both 24 h and 48 h ({rho} = 0.227 and 0.235). Motif-specific perturbation further showed that regulatory attributions were strongly context-dependent, with more than 90% of multi-instance motif modules exhibiting superadditive joint effects. Conclusions: Prediction reliability, explanation validity, and sequence novelty are related but distinct properties of a sequence-to-function model. Evaluating each explicitly gives a more complete basis for deciding when to act on a prediction than accuracy alone.

genomics↗

Mapping the Competence Boundary of a Protein Property Model: A Case Study on Plastic-Degrading Enzymes using ProtTrust-XAI

Machine learning models for protein properties are usually reported by a single accuracy figure, which says how a model behaves on average but not whether to act on any one prediction, especially for a protein unlike anything in the training set. That gap is both a black box problem and an out-of-distribution problem, and it is worst exactly where discovery work happens, on sequences the model has not seen. We present ProtTrust-XAI, a framework that scores each prediction by ensemble consensus and by the structural coherence of its own attribution, and separately tracks a third signal, distance from the training distribution, to catch cases the first two cannot see. We demonstrate it on per-protein thermostability, training a relational graph convolutional network on melting temperatures for over 20,000 proteins using AlphaFold-derived contact graphs and frozen protein language model embeddings. On family-level held-out proteins the model reaches a Spearman correlation of 0.65 and a mean absolute error of 4.1{degrees}C, and predictions the framework labels most trustworthy fall to 3.0{degrees}C, below the assays own reproducibility floor, so a practitioner can act on the label with the same confidence as on the measurement itself. Applying the framework across the full dataset also exposes two representational blind spots, one around cofactor chemistry and one around membrane proteins, each with a distinct mechanistic explanation that points to a specific fix. Transferred to plastic-degrading enzymes at low sequence identity to the training data, absolute predictions collapse while the ranking survives, and a controlled ablation shows this is a general property of distribution shift rather than something particular to that external set. The same transfer identifies where the distance-based signal itself needs recalibrating before deployment, which is a diagnosis the framework produces about itself and not a hidden failure. The result is a practical rule. Inside a models competence domain, trust its labels. Outside it, trust its ranking. A model that reports its own limits, rather than only its average accuracy, is one an experimentalist can actually build on.

bioinformatics↗

TrustPGS: When can a polygenic score be trusted? A per-individual reliability framework across ancestries

Polygenic scores summarise genetic predisposition to a trait, but a population-level accuracy figure cannot tell a clinician whether a given prediction is reliable for the person in front of them. This gap is most consequential for individuals whose ancestry is under-represented in the discovery cohort, precisely the patients for whom a wrong trust call carries the highest clinical cost. We present TrustPGS, a framework that tells clinicians and downstream models which individual predictions can be trusted and which cannot, so that polygenic scores can inform clinical decisions rather than being acted on uniformly regardless of how well-supported each prediction is. The framework rests on two axes calibrated on a discovery cohort, the consensus of a Bayesian posterior-sample ensemble and the directional agreement of the top-magnitude linkage-disequilibrium blocks. We computed SBayesRC posterior-sample scores for ten polygenic traits in the 1000 Genomes Project phase-3 cohort and tested whether the resulting trust labels transfer, without recalibration, to the ancestrally diverse Simons Genome Diversity Project, comparing strict application of the European cutoffs, percentile-rank rescaling, and within-cohort recalibration. Percentile-rank rescaling preserved an enrichment factor above one in non-European populations for five of ten traits (Alzheimer disease, breast cancer, body mass index, LDL cholesterol, and systolic blood pressure), traits whose European and target-cohort distributions were shifted but comparable in shape. Three traits (coronary artery disease, height, and schizophrenia) carried distributions that differed in shape rather than location, a pattern traceable to discovery-cohort bias that recalibration could not repair either, and two further traits (type 2 diabetes and educational attainment) showed intermediate behaviour, present but never enriched in one case, and an apparent success that rank-mapping correctly unmasked as artefactual in the other. Because each of these patterns is detectable before any individual-level claim is made, TrustPGS gives clinicians and downstream models a falsifiable, per-trait basis for deciding when a reliability label can be trusted on a new population, rather than a single portability promise that holds or fails silently.

genomics↗