Search bioRxivSearch

Biology subjects

Keiser, M. J.

Publications and source records attributed to Keiser, M. J..

2 recordsLinked to original sources

Adding stochastic negative examples into machine learning improves molecular bioactivity prediction

Multitask deep neural networks learn to predict ligand-target binding by example, yet public pharmacological datasets are sparse, imbalanced, and approximate. We constructed two hold-out benchmarks to approximate temporal and drug-screening test scenarios whose characteristics differ from a random split of conventional training datasets. We developed a pharmacological dataset augmentation procedure, Stochastic Negative Addition (SNA), that randomly assigns untested molecule-target pairs as transient negative examples during training. Under the SNA procedure, ligand drug-screening benchmark performance increases from R2 = 0.1926 {+/-} 0.0186 to 0.4269{+/-}0.0272 (121.7%). This gain was accompanied by a modest decrease in the temporal benchmark (13.42%). SNA increases in drug-screening performance were consistent for classification and regression tasks and outperformed scrambled controls. Our results highlight where data and feature uncertainty may be problematic, but also show how leveraging uncertainty into training improves predictions of drug-target relationships.

bioinformatics

Interpretable classification of Alzheimer’s disease pathologies with a convolutional neural network pipeline

Neuropathologists assess vast brain areas to identify diverse and subtly-differentiated morphologies. Standard semi-quantitative scoring approaches, however, are coarse-grained and can lack precise neuroanatomic localization. We report a proof-of-concept deep learning pipeline identifying specific neuropathologies--amyloid plaques and cerebral amyloid angiopathy--in immunohistochemical-stained archival slides. Using automated segmentation of stained objects and a cloud-based interface, we annotated >70,000 plaque candidates from 43 whole slide images (WSIs) to train and evaluate convolutional neural networks. Networks achieved strong plaque classification (0.993 and 0.744 areas under the receiver operating characteristic and precision recall curve, respectively) on a 10 WSI hold-out set. Prediction confidence maps visualized morphology distributions from the full-WSI level down to 20x magnification. Resulting plaque-burden scores correlated well with established semi-quantitative scores. Finally, saliency mapping demonstrated that networks learned patterns agreeing with accepted pathologic features. This scalable means to augment a neuropathologists ability may suggest a route to neuropathologic deep phenotyping.

pathology