Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.04.26.591355

EYE-Llama, an in-domain large language model for ophthalmology

Abstract

1BackgroundTraining Large Language Models (LLMs) with in-domain data can significantly enhance their performance, leading to more accurate and reliable question-answering (Q&A) systems essential for supporting clinical decision-making and educating patients. MethodsThis study introduces ophthalmic LLMs trained on in-domain, well-curated datasets. We present an open-source substantial ophthalmic language dataset for model training. Our models (EYE-Llama), were pre-trained on an ophthalmology-specific dataset, including paper abstracts, textbooks, and Wikipedia articles. Subsequently, the models underwent fine-tuning using a diverse range of QA pairs. Our models were compared to baseline Llama 2, ChatDoctor, Meditron, Llama 3, and ChatGPT (GPT3{middle dot}5) models, using four distinct test sets, and evaluated quantitatively (Accuracy, F1 score, BERTScore, BARTScore and BLEU score) and qualitatively by two ophthalmologists. FindingsUpon evaluating the models using the synthetic dialogue test set with three different metrics (BERTScore, BARTScore, and BLEU score), our models demonstrated superior performance. Specifically, when evaluated using BERTScore, our models surpassed Llama 2, Llama 3, Meditron, and ChatDoctor in terms of F1 score, and performed on par with ChatGPT, which was trained with 175 billion parameters (EYE-Llama: 0.57, Llama 2: 0.56, Llama 3: 0.55, Meditron: 0.50, ChatDoctor: 0.56, and ChatGPT: 0.57). Additionally, the EYE-Llama model outperformed the above models when evaluated using BARTScore and BLEU scores. When tested on the MedMCQA test set, the fine-tuned models exhibited higher accuracy compared to Llama 2, Meditron, and ChatDoctor models (EYE-Llama: 0.39, Llama 2: 0.33, ChatDoctor: 0.29, Meditron: 0.22). However, ChatGPT, and Llama 3 models outperformed EYE-Llama, achieving accuracies of 0.55, 0.78, and 0.90, respectively. On the PubmedQA test set, our model showed improved accuracy over all other models (EYE-Llama: 0.96, Llama 2: 0.90, Llama 3: 0.92, Meditron: 0.76, ChatGPT: 0.93, ChatDoctor: 0.92). InterpretationThe study shows that pre-training and fine-tuning LLMs like EYE-Llama enhances their performance in specific medical domains. Our EYE-Llama models surpass baseline Llama 2 in all evaluations, highlighting the effectiveness of specialized LLMs in medical QA systems. FundingFunded by NEI R15EY035804 (MNA), R21EY035271 (MNA), and UNC Charlotte Faculty Research Grant (MNA)

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

haghighi, t., Gholami, S., Sokol, J. T., Kishnani, E., Ahsaniyan, A., Rahmanian, H., Hedayati, F., Leng, T., Alam, M. N.. 2024-04-29. EYE-Llama, an in-domain large language model for ophthalmology. https://doi.org/10.1101/2024.04.26.591355

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Dynamic Compression Platform for Live Imaging of Scaffold-Transmitted Cellular Mechanoresponses

Mechanical characterization of biomaterial scaffolds is essential to evaluate their capacity to meet the functional demands of target tissues in tissue engineering and regenerative medicine applications. Scaffolds designed to interface with living tissues must support the transmission of mechanical cues to resident cells and stimulate mechanosignaling pathways that are essential to their function. In joints, bone and cartilage cells act as primary mechanosensors, converting mechanical stimuli into biochemical signals that regulate tissue homeostasis and remodelling. Therefore, evaluating cellular mechanoresponses to scaffold-transmitted compression in vitro can inform the development of functional tissue-engineered constructs. For example, poly({epsilon}-caprolactone) (PCL) scaffolds are highly relevant for bone and cartilage tissue engineering due to their biocompatibility, stable mechanical properties and slow degradation. Here, we applied a custom-built device to study compression-induced mechanosignaling in MC3T3-E1 pre-osteoblast cells. The device is composed of a polydimethylsiloxane (PDMS) pillar, a force-sensing load cell, and a piezoelectric linear track. A protocol is described in which MC3T3-E1 cells are repeatedly compressed, while in parallel live tracking of force measurements and live imaging of intracellular calcium dynamics in MC3T3-E1 cells are recorded. PCL scaffolds fabricated by melt electrowriting (MEW) were subsequently integrated into the platform. Scaffold-transmitted compression triggered dynamic increases in cytosolic calcium; in MC3T3-E1 cells located directly under the PCL microfibers, but also in cells located in the interfiber spaces. This device and workflow facilitate in vitro investigations of real-time cellular mechanoresponses to dynamic compression applied with biomaterial scaffolds, and provides a testing platform for evaluating the mechanotransductive properties of scaffolds intended for tissue engineering applications.

bioengineering↗

Ultrasound Tracking Reveals Progressive Regional Strain Differences in Human Achilles Tendons During Fatigue Loading

Ultrasound is commonly used to assess structural changes in symptomatic Achilles tendons, but quantitative biomechanical metrics for progressive tendon deterioration remain limited. The goal of this study was to develop and validate an automated ultrasound tracking algorithm for regional tendon deformation and evaluate strain progression in survived and ruptured tendons during fatigue loading. We hypothesized that maximum strain, average strain, and strain heterogeneity would exhibit different trajectories between groups. Ten cadaveric Achilles tendons underwent cyclic loading with stress tests every 500 cycles until rupture or 150,000 cycles. Ultrasound images acquired during stress tests were analyzed using an automated tracking algorithm to generate spatially resolved regional strain fields. Ultrasound-derived bulk strain was highly correlated with actuator-derived strain in survived (R^2 = 0.968 +/- 0.017) and ruptured tendons (R^2 = 0.972 +/- 0.014). Maximum and average longitudinal strains progressively diverged between groups across fatigue life (Group x FatigueLife: p = 0.003 and p < 0.0001, respectively). During the first 10,000 cycles, average strain decreased in survived tendons ({beta} = -0.0268%, p = 0.0215) but not ruptured tendons ({beta} = 0.0147%, p = 0.1197), with a significant Group x Cycle interaction (p = 0.0061). This study demonstrates that the algorithm quantified Achilles tendon deformation with high fidelity and enabled spatially resolved strain assessment throughout fatigue loading. Maximum and average strain followed different trajectories between groups, whereas strain heterogeneity did not. Early differences in tendon biomechanics suggest that regional strain behavior may change before pronounced differences in absolute magnitude develop.

bioengineering↗

Brain organoid computing for robotic decision-making

Biomimicry has inspired the evolution of robotics toward greater autonomy, adaptability, and symbiosis with humans and dynamic environments. However, current robotic systems still face major challenges in recapitulating the high-efficiency decision-making capabilities of the human brain under complex and dynamic conditions. Here, we present Brainobot, a biohybrid robotic system that establishes a brain organoid controller as a high-level robotic decision-making layer for closed-loop embodiment. By leveraging brain organoid reservoir computing, Brainobot interacts with dynamic environments by receiving and processing sensory inputs and generating motor actions. As a proof-of-concept demonstration, Brainobot is implemented in a humanoid robotic system to perform real-world tasks, including object grasping and laser chasing. Interestingly, Brainobot exhibits unique features, including cross-task adaptivity, high computing efficiency, and low energy consumption. Thus, our approach may provide insights for advancing robotic embodiment and understanding biological decision-making.

bioengineering↗