Search bioRxiv⌕ Search

Biology subjects

Joy, J.

Publications and source records attributed to Joy, J..

4 recordsLinked to original sources

Federated Knowledge Retrieval Elevates Large Language Model Performance on Biomedical Benchmarks

BackgroundLarge language models (LLMs) have significantly advanced natural language processing in biomedical research, however, their reliance on implicit, statistical representations often results in factual inaccuracies or hallucinations, posing significant concerns in high-stakes biomedical contexts. ResultsTo overcome these limitations, we developed BTE-RAG, a retrieval-augmented generation framework that integrates the reasoning capabilities of advanced language models with explicit mechanistic evidence sourced from BioThings Explorer, an API federation of more than sixty authoritative biomedical knowledge sources. We systematically evaluated BTE-RAG in comparison to traditional LLM-only methods across three benchmark datasets that we created from DrugMechDB. These datasets specifically targeted gene-centric mechanisms (798 questions), metabolite effects (201 questions), and drug-biological process relationships (842 questions). On the gene-centric task, BTE-RAG increased accuracy from 51% to 75.8% for GPT-4o mini and from 69.8% to 78.6% for GPT-4o. In metabolite-focused questions, the proportion of responses with cosine similarity scores of at least 0.90 rose by 82% for GPT-4o mini and 77% for GPT-4o. While overall accuracy was consistent in the drug-biological process benchmark, the retrieval method enhanced response concordance, producing a greater than 10% increase in high-agreement answers (from 129 to 144) using GPT-4o. ConclusionFederated knowledge retrieval provides transparent improvements in accuracy for large language models, establishing BTE-RAG as a valuable and practical tool for mechanistic exploration and translational biomedical research.

bioinformatics↗

Catalyzing computational biology research at an academic institute through an interest network

Biology has been transformed by the rapid development of computing and the concurrent rise of data-rich approaches such as -omics or high-resolution imaging. However, there is a persistent computational skills gap in the biomedical research workforce. Inherent limitations of classroom teaching and institutional core support highlight the need for accessible ways for researchers to explore developments in computational biology. An analysis of the Scripps Research Genomics Core revealed an increasingly diverse set of experiments: the share of experiments other than bulk RNA- or DNA-seq increased from 34% to 60% within 10 years, requiring more tailored computational analyses. These challenges were tackled by forming a volunteer-led affinity group of over 300 academic biomedical researchers interested in computational biology referred to as the Computational Biology and Bioinformatics (CBB) affinity group. This adaptive group has provided continuing education and networking opportunities through seminars, workshops and coding sessions while evolving along with the needs of its members. A survey of CBBs impact confirmed the groups events increased the members exposure to computational biology educational and research events (79% respondents) and networking opportunities (61% respondents). Thus, volunteer-led affinity groups may be a viable complement to traditional institutional resources for enhancing the application of computing in biomedical research.

bioinformatics↗

Adaptation of a transmitted/founder simian-human immunodeficiency virus for enhanced replication in rhesus macaques

Transmitted/founder (TF) simian-human immunodeficiency viruses (SHIVs) express HIV-1 envelopes modified at position 375 to efficiently infect rhesus macaques while preserving authentic HIV-1 Env biology. TF SHIV.C.CH505 is an extensively characterized virus shown to recapitulate key features of HIV-1 immunobiology, including CCR5-tropism, a tier 2 neutralization profile, reproducible early viral kinetics, and authentic immune responses. SHIV.C.CH505 is used frequently in nonhuman primate studies of HIV, but viral loads after months of infection are variable and typically lower than those in people living with HIV. We hypothesized that additional mutations besides {Delta}375 might further enhance virus fitness without compromising essential components of CH505 Env biology. From sequence analysis of SHIV.C.CH505-infected macaques across multiple experiments, we identified a signature of envelope mutations associated with higher viremia. We then used short-term in vivo mutational selection and competition to identify a minimally adapted SHIV.C.CH505 with just five amino acid changes that substantially improve virus replication fitness in macaques. Next, we validated the performance of the adapted SHIV in vitro and in vivo and identified the mechanistic contributions of selected mutations. In vitro, the adapted SHIV shows improved virus entry, enhanced replication on primary rhesus cells, and preserved neutralization profiles. In vivo, the minimally adapted virus rapidly outcompetes the parental SHIV with an estimated growth advantage of 0.14 days-1 and persists through suppressive antiretroviral therapy to rebound at treatment interruption. Here, we report the successful generation of a well-characterized, minimally adapted virus, termed SHIV.C.CH505.v2, with enhanced replication fitness and preserved native Env properties that can serve as a new reagent for NHP studies of HIV-1 transmission, pathogenesis, and cure. Author SummaryThe power of the nonhuman primate model of HIV to predict outcomes in people living with HIV (PLWH) depends on authentic virus-host interactions. In pursuit of viruses that generate infection that mirrors the effects of HIV-1 in PLWH, we developed a minimally adapted version of a commonly used virus, SHIV.C.CH505, which has better fitness than the parental virus while retaining important biological properties. First, we studied virus sequences from SHIV.C.CH505-infected rhesus macaques to identify a signature of mutations common to animals with higher viral loads. We then tested viruses containing the various mutations in the lab and in animals to determine the most fit version and to identify the contribution of each mutation. Ultimately, we identified a minimally adapted version of SHIV.C.CH505 with just 5 amino acid substitutions that enhances virus replication and preserves CH505 envelope properties, including sensitivity to clinically relevant broadly neutralizing antibodies. This new virus, called SHIV.C.CH505.v2 replicates well in macaques over time and persists through antiretroviral therapy. SHIV.C.CH505.v2 could be an important component of nonhuman primate studies of HIV prevention, therapy, and cure.

microbiology↗

Extensive heterogeneity of the HIV-1 infected CD4+ T-cell reservoir revealed by single-cell viral ASAPseq

Understanding the complexity of the long-lived HIV reservoir during antiretroviral therapy (ART) remains a major impediment for HIV cure research. To address this, we developed single-cell viral ASAPseq to precisely define the unperturbed peripheral blood HIV-infected memory CD4+ T cell reservoir from antiretroviral treated people living with HIV (ART-PLWH) via the presence of integrated accessible proviral DNA in concert with epigenetic and cell surface protein profiling. We identified profound reservoir heterogeneity within and between ART-PLWH, characterized by novel and known surface markers within total and individual memory CD4+ T cell subsets. We further uncovered novel epigenetic profiles and transcription factor motifs enriched in HIV-infected cells that suggest infected cells with accessible provirus, irrespective of reservoir distribution, are poised for reactivation during ART treatment. Together, our findings reveal the extensive inter- and intrapersonal cellular heterogeneity of the HIV reservoir, and establish an initial multiomic atlas to develop targeted reservoir elimination strategies.

microbiology↗