Search bioRxiv⌕ Search

Biology subjects

Howard, F. M.

Publications and source records attributed to Howard, F. M..

5 recordsLinked to original sources

Validating a low-cost, open-source, locally manufactured workstation and computational pipeline for automated histopathology evaluation using deep learning

Deployment and access to state-of-the-art diagnostic technologies remains a fundamental challenge in providing equitable global cancer care to low-resource settings. The expansion of digital pathology in recent years and its interface with computational biomarkers provides an opportunity to democratize access to personalized medicine. Here we describe a low-cost platform for digital side capture and computational analysis composed of open-source components. The platform provides low-cost ($200) digital image capture from glass slides and is capable of real-time computational image analysis using an open-source deep learning (DL) algorithm and Raspberry Pi ($35) computer. We validate the performance of deep learning models performance using images captured from the open-source workstation and show similar model performance when compared against significantly more expensive standard institutional hardware.

pathology↗

Latent transcriptional programs reveal histology-encoded tumor features spanning tissue origins

Precision medicine in cancer treatment depends on deciphering tumor phenotypes to reveal the underlying biological processes. Molecular profiles, including transcriptomics, provide an information-rich tumor view, but their high-dimensional features and assay costs can be prohibitive for clinical translation at scale. Recent studies have suggested jointly leveraging histology and genomics as a strategy for developing practical clinical biomarkers. Here, we use machine learning techniques to identify de novo latent transcriptional processes in squamous cell carcinomas (SCCs) and to accurately predict their activity levels directly from tumor histology images. In contrast to analyses focusing on pre-specified, individual genes or sample groups, our latent space analysis reveals sets of genes associated with both histologically detectable features and clinically relevant processes, including immune response, collagen remodeling, and fibrosis. The results demonstrate an approach for discovering clinically interpretable histological features that indicate complex, potentially treatment-informing biological processes.

cancer biology↗

Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers

BackgroundLarge language models such as ChatGPT can produce increasingly realistic text, with unknown information on the accuracy and integrity of using these models in scientific writing. MethodsWe gathered ten research abstracts from five high impact factor medical journals (n=50) and asked ChatGPT to generate research abstracts based on their titles and journals. We evaluated the abstracts using an artificial intelligence (AI) output detector, plagiarism detector, and had blinded human reviewers try to distinguish whether abstracts were original or generated. ResultsAll ChatGPT-generated abstracts were written clearly but only 8% correctly followed the specific journals formatting requirements. Most generated abstracts were detected using the AI output detector, with scores (higher meaning more likely to be generated) of median [interquartile range] of 99.98% [12.73, 99.98] compared with very low probability of AI-generated output in the original abstracts of 0.02% [0.02, 0.09]. The AUROC of the AI output detector was 0.94. Generated abstracts scored very high on originality using the plagiarism detector (100% [100, 100] originality). Generated abstracts had a similar patient cohort size as original abstracts, though the exact numbers were fabricated. When given a mixture of original and general abstracts, blinded human reviewers correctly identified 68% of generated abstracts as being generated by ChatGPT, but incorrectly identified 14% of original abstracts as being generated. Reviewers indicated that it was surprisingly difficult to differentiate between the two, but that the generated abstracts were vaguer and had a formulaic feel to the writing. ConclusionChatGPT writes believable scientific abstracts, though with completely generated data. These are original without any plagiarism detected but are often identifiable using an AI output detector and skeptical human reviewers. Abstract evaluation for journals and medical conferences must adapt policy and practice to maintain rigorous scientific standards; we suggest inclusion of AI output detectors in the editorial process and clear disclosure if these technologies are used. The boundaries of ethical and acceptable use of large language models to help scientific writing remain to be determined.

scientific communication and education↗

Multimodal Prediction of Breast Cancer Recurrence Assays and Risk of Recurrence

Gene expression-based recurrence assays are strongly recommended to guide the use of chemotherapy in hormone receptor-positive, HER2-negative breast cancer, but such testing is expensive, can contribute to delays in care, and may not be available in low-resource settings. Here, we describe the training and independent validation of a deep learning model that predicts recurrence assay result and risk of recurrence using both digital histology and clinical risk factors. We demonstrate that this approach outperforms an established clinical nomogram (area under the receiver operating characteristic curve of 0.833 versus 0.765 in an external validation cohort, p = 0.003), and can identify a subset of patients with excellent prognoses who may not need further genomic testing.

cancer biology↗

The Impact of Digital Histopathology Batch Effect on Deep Learning Model Accuracy and Bias

The Cancer Genome Atlas (TCGA) is one of the largest biorepositories of digital histology. Deep learning (DL) models have been trained on TCGA to predict numerous features directly from histology, including survival, gene expression patterns, and driver mutations. However, we demonstrate that these features vary substantially across tissue submitting sites in TCGA for over 3,000 patients with six cancer subtypes. Additionally, we show that histologic image differences between submitting sites can easily be identified with DL. This site detection remains possible despite commonly used color normalization and augmentation methods, and we quantify the digital image characteristics constituting this histologic batch effect. As an example, we show that patient ethnicity within the TCGA breast cancer cohort can be inferred from histology due to site-level batch effect, which must be accounted for to ensure equitable application of DL. Batch effect also leads to overoptimistic estimates of model performance, and we propose a quadratic programming method to guide validation that abrogates this bias.

bioinformatics↗