Search bioRxiv⌕ Search

Biology subjects

Hwang, M.-J.

Publications and source records attributed to Hwang, M.-J..

2 recordsLinked to original sources

Predicting FDA approvability of small-molecule drugs

A high rate of compound attrition makes drug discovery via conventional methods time-consuming and expensive. Here, we showed that machine learning models can be trained to classify compounds into distinctive groups according to their status in the drug development process, which can significantly reduce the compound attrition rate. Using molecular structure fingerprints and physicochemical properties as input, our models accurately predicted which drug compounds would proceed to trial, with an area under the receiver operating curve (AUC) of 0.94 {+/-} 0.01 (mean {+/-} standard deviation). Our models also identified which drugs in clinical trials would be approved by the US Food and Drug Administration (FDA) to go on the market, with an AUC of 0.73 {+/-} 0.02. The predictive power of our models could reduce the attrition rate of preclinical compounds to enter clinical trials from 65%, as with conventional methods, to 12% (with 92% sensitivity) and the clinical trial failure rate from 80-90% to 29% (with 83% sensitivity). The results largely held in additional tests on new clinical trial compounds and new FDA-approved drugs, as well as on drugs uniquely approved for use in Europe and Japan. SIGNIFICANCE STATEMENTThe odds of developing a drug approved by the US Food and Drug Administration (FDA) are slim, meaning that the vast majority of drug candidates would fail tests for safety and efficacy in the drug discovery process, rendering it highly inefficient and costly. Here, we have developed machine learning models to predict drug compounds worthy of clinical trials with high accuracy, and clinical-trial compounds to receive FDA approval with a much higher success rate than that achieved by the traditional approach. Our computational prediction requires input of only the drug compounds chemical structure and physicochemical properties. It can help mitigate the long-standing problem of drug discovery.

bioinformatics↗

Detection and Classification of Cardiac Arrhythmias by a Challenge-Best Deep Learning Neural Network Model

BackgroundElectrocardiogram (ECG) is widely used to detect cardiac arrhythmia (CA) and heart diseases. The development of deep learning modeling tools and publicly available large ECG data in recent years has made accurate machine diagnosis of CA an attractive task to showcase the power of artificial intelligence (AI) in clinical applications.\n\nMethods and FindingsWe have developed a convolution neural network (CNN)-based model to detect and classify nine types of heart rhythms using a large 12-lead ECG dataset (6877 recordings) provided by the China Physiological Signal Challenge (CPSC) 2018. Our model achieved a median overall F1-score of 0.84 for the 9-type classification on CPSC2018s hidden test set (2954 ECG recordings), which ranked first in this latest AI competition of ECG-based CA diagnosis challenge. Further analysis showed that concurrent CAs observed in the same patient were adequately predicted for the 476 patients diagnosed with multiple CA types in the dataset. Analysis also showed that the performances of using only single lead data were only slightly worse than using the full 12 lead data, with leads aVR and V1 being the most prominent. These results are extensively discussed in the context of their agreement with and relevance to clinical observations.\n\nConclusionsAn AI model for automatic CA diagnosis achieving state-of-the-art accuracy was developed as the result of a community-based AI challenge advocating open-source research. In- depth analysis further reveals the models ability for concurrent CA diagnosis and potential use of certain single leads such as aVR in clinical applications.\n\nAbbreviationsCA, cardiac arrhythmia; AF, Atrial fibrillation; I-AVB, first-degree atrioventricular block; LBBB, left bundle branch block; RBBB, right bundle branch block; PAC, premature atrial contraction; PVC, premature ventricular contraction; STD, ST-segment depression; STE, ST-segment elevation.

bioinformatics↗