Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.02.02.703416

Predicting Post-Stroke Aphasia Speech Performance from Multimodal Data with Explainable Machine Learning

Abstract

Aphasia, an acquired language deficit, is the most common post-stroke focal cognitive impairment, and roughly 60% cases become chronic (duration >6 months). Aphasia therapies could be optimized if clinicians could make personalized predictions of how individual persons with aphasia (PWA) would be likely to perform on particular language tasks. However, current approaches relying on imaging, lesion volume, patient demographics, and clinical scores achieve less than 50% accuracy in predicting performance in PWA. Research algorithms using complex imaging and fMRI can make binary predictions about the presence or absence of aphasia but do not give more clinically relevant information. We aim to predict word-by-word speech accuracy in PWA to better enable personalized speech therapies. To be clinically informative, machine learning models developed for this purpose should use clinically available inputs, explain key features behind a prediction, and generalize to new PWA and previously unseen words. This study combines multimodal input features from clinical testing scores and structural MRI neuroimaging with a novel data source: word-by-word linguistic difficulty. We computed metrics of cognitive burden, such as semantic selection and recall demands, and articulatory burden, such as word length in phonemes and syllables, using naturalistic corpora containing over a billion words of English text. Retrospective training, ten-fold cross validation and 500-run bootstrapping of different machine learning models with various combinations of input features was conducted using 4620 trials. A simplified version of the best model using widely available inputs was deployed clinically through a web app, and prospective generalization was tested on 570 trials with unseen words and different naming tasks in new PWA. We found the best performances with random forest classifiers using linguistic difficulty combined with either clinical information (AUROC {+/-} SEM = 0.87 {+/-} 0.07), or all together with structural imaging connectivity (0.90 {+/-} 0.04). Classifiers using multimodal inputs significantly outperformed others employing single inputs (range 0.66-0.85, p<0.05). Extracting feature importances from the best model showed that Western Aphasia Battery scores, semantic demands, number of phonemes, and syllables were predictive of PWA speech accuracy. Structural integrity in peri-lesional brain regions predicted better language performance whereas higher connectivity of select contralateral homotopes contributed to prediction of worse speech. Without the inclusion of MRI data, lesion volume was a key predictor of PWA speech as well. A simplified, clinically ready, explainable model (publicly available as AphasiaLENS web application) predicted PWA accuracy for any user-entered word, not restricted to a standardized battery. Its prospective generalization performance was not significantly different from the best model using full inputs (AUROC ranges 0.81-0.89, p>0.05). Thus, our research can help inform individualized treatment planning for PWA, while also suggesting research targets through better understanding of brain-behavior relationships.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Parchure, S., Gupta, A., Kelkar, A., Vnenchak, L., Faseyitan, O., Medaglia, J. D., Harvey, D. Y., Coslett, H. B., Hamilton, R. H.. 2026-02-05. Predicting Post-Stroke Aphasia Speech Performance from Multimodal Data with Explainable Machine Learning. https://doi.org/10.64898/2026.02.02.703416

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Comparative study of chlorophyll measurement in Physcomitrium patens moss using a conventional microscope adapted for combined 2D+1D imaging and spectral analysis

Imaging spectroscopy often requires expensive and complex equipment. Here we show a simple procedure for attaching a standard miniature fiber spectrometer to a conventional microscope, allowing easy integration of 2D imaging with 1D high-resolution spectral measurements. This combination provides much of the benefit of a full imaging spectrometer without the large equipment investment, and we provide instructions for modifying microscopes to this setup and the present measurements of living cells that demonstrate their performance. Using this setup, we compare the quantitative measurement of chlorophyll concentration in Physcomitrium patens moss using color imaging and spectral sampling.

bioengineering↗

De novo designed single-domain antibodies protect against lethal cobra venom neurotoxicity in vivo

Generative protein design can now rapidly produce de novo binders with high affinity and functional activity against a wide range of targets, including lethal snake venom toxins. However, so far most reported successes rely on new-to-nature scaffolds with limited therapeutic precedent. Single-domain antibodies (VHHs) offer a clinically validated alternative scaffold that can bind and neutralize long-chain -neurotoxins, which are some of the most lethal components in snake venoms. Here we compare three recently established de novo design models with VHH-design capabilities (Germinal, RFantibody, and BoltzGen) for their ability to generate VHHs against the neurotoxin -cobratoxin from the monocled cobra (Naja kaouthia). Using standardized model inputs and evaluation criteria based on AlphaFold3 interface confidence (ipTM) and RMSD self-consistency, we find that Germinal was the only method to generate designs passing stringent in silico criteria for experimental testing. We therefore performed a larger Germinal design campaign employing three different VHH frameworks and experimentally validated 46 designs in vitro. Of these, 42 expressed as soluble proteins and we identified four binding hits derived from two of the three tested frameworks. Of the four binders, two lead candidates were further characterized and demonstrated high affinity (KDs of 4.1 nM and 10.8 nM), monomeric behavior and low polyreactivity, indicating favorable biophysical and developability properties, as well as functional toxin neutralization in vitro. To assess their therapeutic potential we investigated their ability to protect against -cobratoxin toxicity in vivo. Both candidates fully protected mice after -cobratoxin challenge, with 100% survival compared to a lethal control. One candidate also retained notable neutralization capacity against whole venom of Naja kaouthia with a survival of 56%, while the other protected 22% when tested in a rescue setting. Together, we demonstrate that de novo VHH design can generate high affinity single-domain antibodies with in vivo protection against lethal cobra venom neurotoxicity, and provide practical insights into method- and framework-dependent performance.

bioengineering↗

Simple Feedback for Complex Movement: Capturing Whole-Limb Reorganization during Single-IMU Gait Retraining

Clinical gait retraining typically relies on multi-sensor arrays and high-dimensional feedback displays, imposing setup and interpretation burdens that limit routine clinical deployment. We developed a single-IMU visual biofeedback system that delivers real-time feedback of Lower Limb Trajectory Error (LLTE), a composite kinematic error metric integrating knee position and shank angle across the stance phase. Twenty able-bodied adults walked on a treadmill under two visual biofeedback targets (flexed-knee, extended-knee) while receiving either corrected (n=10) or uncorrected (n=8) feedback, where the correction accounted for limb orientation at initial contact. LLTE and stance-phase knee kinematics adapted consistently under the flexed-knee target for both feedback groups, with feedback formulation moderating the temporal trajectory of change. Adaptation toward the extended-knee target was limited, likely because participants were already operating near terminal knee extension and because the scalar error metric provided limited directional information for correction. Ankle range of motion (ROM) changed significantly across the stance phase under both target conditions, while hip ROM did not. Multiscale multivariate sample entropy (MSMVSE) increased monotonically with time scale across all conditions, with no statistically distinguishable difference between corrected and uncorrected feedback. These results suggest that single-IMU LLTE biofeedback can modify gait mechanics and that adaptation was expressed across multiple lower-limb segments rather than through changes at a single joint.

bioengineering↗