Search bioRxiv⌕ Search

Biology subjects

Brimo, N.

Publications and source records attributed to Brimo, N..

2 recordsLinked to original sources

Hierarchical temporal transformer for cancer grade prediction and cross cancer transfer learning from pathology reports

Language models for cancer clinical reports carry two blind spots. They read each report in isolation, ignoring how a patients disease changes across visits, and they are evaluated only on cancer types present in their training data. We present the Hierarchical Temporal Transformer (HTT), a two-level architecture that addresses both. Level 1 encodes each report with BiomedBERT adapted by low-rank adaptation (LoRA). Level 2 is a temporal transformer that reads a patients full report sequence using a continuous-time positional encoding built from the measured number of days between visits, with learnable cancer-type embeddings supplying per-family conditioning. Two experiments test the two capabilities separately, since no fully open corpus contains longitudinal reports for many cancer types. On a controlled synthetic corpus of sequential radiology reports, in which progression phrases are inserted from templated trajectories, HTT reaches a validation AUROC of 0.942 against 0.881 for a single-report baseline and transfers to held-out pancreatic cancer at 0.995 against 0.949 while the two models are indistinguishable on a 60-patient test set. On 4,786 real pathology reports from the TCGA-Reports corpus spanning 14 cancer types HTT predicts tumor grade for three types withheld entirely from training, reaching AUROC 1.000 on thyroid carcinoma, 0.960 on sarcoma and 0.808 on lung squamous cell carcinoma. The mean held-out AUROC of 0.923 equals the in-distribution test AUROC of 0.923, so transfer to unseen cancer families incurred no measurable penalty. Ablation on the real corpus shows that the transfer is carried by the pre-trained encoder rather than by the temporal components, which, with one report per patient, contribute 0.39 AUROC points. Grade-related pathological language therefore appears to be learnable in a cancer-type agnostic way, which points toward unified cancer NLP systems that require no per-type retraining.

bioinformatics↗

Medicament identity rather than total loading governs the morphology of electrospun poly(vinylpyrrolidone) nanofibers for regenerative endodontics: a machine learning analysis of a failure-inclusive dataset

Electrospun fibers loaded with antibiotics or calcium hydroxide are being developed as intracanal carriers for regenerative endodontics, where the dose must stay low enough to spare the stem cells that repopulate the canal. Formulation development sweeps the medicament concentration while holding the polymer and machine settings fixed. We asked whether that sweep targets the right variable. We assembled ENDOSPIN-29, a dataset of 29 poly(vinylpyrrolidone) formulations produced under a single process backbone and loaded with metronidazole, ciprofloxacin, minocycline or calcium hydroxide, alone and in combination, retaining the five that produced no submicron fibers. Across nine regression models, those given per-medicament composition predicted fiber diameter far better than the same models given only total loading. The best reached a leave-one-out coefficient of determination of 0.84 and a median relative error of 18%, whereas every loading-only model performed at or below a mean baseline. Uniformity and distribution span behaved likewise; asymmetry was unpredictable. Dose response ran in opposite directions for different actives: ciprofloxacin thinned fibers monotonically from 406 to 257 nm, while metronidazole thickened them and destroyed fiber formation above 10% w/w. Holding out an entire medicament class removed the advantage, bounding the method to interpolation within a known drug panel.

bioengineering↗