Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.11.05.622126

Advancing Plant Metabolic Research By Using Large Language Models To Expand Databases And Extract Labelled Data

Abstract

Premise: Recently, plant science has seen transformative advances in scalable data collection for sequence and chemical data. These large datasets, combined with machine learning, revealed that conducting plant metabolic research on large scales yields remarkable insights. A key next step in increasing scale has been revealed with the advent of accessible large language models, which, even in their early stages, can distill structured data from literature. This brings us closer to creating specialized databases that consolidate virtually all published knowledge on a topic. Methods: Here, we first test different prompt engineering technique / language model combinations in the identification of validated enzyme-product pairs. Next, we evaluate automated prompt engineering and retrieval augmented generation applied to identifying compound-species associations. Finally, we build and determine the accuracy of a multimodal language model-based pipeline that transcribes images of tables into machine-readable formats. Results: When tuned for each specific task, these methods perform with high accuracies (80-90 percent for enzyme-product pair identification and table image transcription), or with modest accuracies (50 percent) but lower false-negative rates than previous methods (down to 40 percent from 55 percent) for compound-species pair identification. Discussion: We enumerate several suggestions for working with language models as researchers, among which is the importance of the users domain-specific expertise and knowledge. Significance StatementScientific databases have played a major role in advancing metabolic research. However, even todays advanced databases are incomplete and/or are not built to best suit certain research tasks. Here, we explored and evaluated the use of large language models and various prompt engineering techniques to expand and subset existing databases in task-specific ways. Our results illustrate the potential for high-accuracy additions and restructurings of existing databases using language models, assuming the specific methods by which the models are used are tuned and validated for the specific task. These findings are important because they outline a method by which we could greatly expand existing databases and rapidly tailor them to specific research efforts, leading to greater research productivity and effective utilization of past research findings. All authors collected data, analyzed data, prepared the manuscript, and approved its final version. The authors declare that they have no competing interests.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Knapp, R., Johnson, B., Busta, L.. 2024-11-06. Advancing Plant Metabolic Research By Using Large Language Models To Expand Databases And Extract Labelled Data. https://doi.org/10.1101/2024.11.05.622126

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

AtNHR2A and AtNHR2B participate in unconventional protein secretion in response to environmental stress

The Arabidopsis thaliana nonhost resistance proteins 2A (AtNHR2A) and 2B (AtNHR2B) play crucial roles in plant immunity as the single mutants Atnhr2a and Atnhr2b and the double mutant Atnhr2bAtnhr2a are susceptible to the non-adapted pathogen Pseudomonas syringae pv. tabaci that is unable to infect wild-type Col-0 plants. The localization of fluorescent versions of AtNHR2A and AtNHR2B to compartments of the endomembrane system together with their interaction with secreted proteins suggested a function in endomembrane-mediated secretory processes participating in plant immunity. Comparative apoplastic proteomics analysis between wild type Col-0 and the double mutant Atnhr2bAtnhr2a after treatment P. syringae pv. tabaci, revealed that AtNHR2A and AtNHR2B are indeed required for the secretion of proteins containing N-terminal signal peptides that occurs through the conventional protein secretion pathway. In this work, we leveraged these apoplastic proteomics datasets to identify proteins lacking N-terminal signal peptide and expected to be secreted through unconventional secretion pathway(s). We discovered that AtNHR2A and AtNHR2B are also required for the secretion of proteins through an unconventional secretion pathway that, intriguingly, included proteins previously associated with abiotic stress. These findings led us to define the subcellular dynamics of AtNHR2A and AtNHR2B, and through co-localization analyses and the use of vesicle trafficking inhibitors, we uncovered their trafficking pathways transitioning through Golgi-dependent and Golgi-independent pathways to ultimately reach the central vacuole. Our findings suggest that AtNHR2A and AtNHR2B participate in a multivesicular bodies-vacuole-mediated unconventional secretion pathway that results in the release of proteins involved in plant responses to environmental stresses.

plant biology↗

Low-cost rhizotron imaging and zero-shot deep-learning resolve temporal, spatial, and genetic variation in grapevine rootstock root systems

Root system architecture shapes how grapevine rootstocks take up water and nutrients, yet roots remain the least phenotyped grapevine organ because they are hidden and hard to image. We present a low-cost phenotyping pipeline that pairs custom acrylic rhizotrons (about US$30 each) with a consumer flatbed scanner and BiRefNet, a general-purpose deep-learning model used without training on root images, followed by automated mask cleaning, skeleton-based trait extraction, and soil moisture mapping. We tested it on nine commercial rootstocks scanned 16 times over 42 days after transplanting (DAT), with half under a ten-day water deficit. From 1,108 images we extracted 21 whole-root, depth-resolved, and topological traits. Genotypes differed in nearly every trait and in how they changed over time. Heritability of size and branching traits peaked at 0.92-0.93 between 21 and 31 DAT and fell for width, depth, and convex hull once roots reached the rhizotron walls, defining the best measurement window. The image-derived soil moisture map accurately tracked the deficit and its recovery. Deficit plants shifted new root growth to deeper soil without growing less overall, and the substrate dried fastest around older and denser roots. Root brightness decreased with root age and local moisture, and transport segments (axes serving several tips) were brighter than terminal laterals in every genotype. Root system size was associated with stomatal conductance in well-watered plants, and stomatal recovery after re-watering correlated with new root growth. The pipeline turns simple hardware into a quantitative, time-resolved root phenotyping platform suitable for breeding.

plant biology↗

Engineering chromatin to encode transcriptional immune memory in Arabidopsis

Transcriptional memory enables organisms to respond more rapidly to recurrent stress, yet the underlying features of chromatin that contribute to this transcriptional recalibration remain poorly defined. Here we identify the genes displaying transcriptional memory in response to the bacterial immune elicitor, flg22, in Arabidopsis thaliana. In comparison to non-memory response genes, these memory genes show a preference for tissue-specific over uniform spatial expression patterning. The chromatin architecture of these genes in the resting state displays depletion of H3K4me3, elevation H3K27me3 and a subset are marked by H3K27me3-H3K4me3 bivalency. The H3K4me3 demethylase, JMJ14, is required for transcriptional memory, with JMJ14 occupancy enriched over memory gene loci. Upon priming, chromatin is reconfigured, with H3K4me3 levels increasing in a sustained manner at memory gene loci. To assess the function of this H3K4me3 accrual, we employ epigenome-engineering, observing that its targeted deposition at memory gene loci, including the WRKY29 locus, is sufficient to drive transcriptional memory and can endow plants with enhanced resistance to the bacterial pathogen, Pseudomonas syringae. Together, the findings demonstrate a causal role for H3K4me3 in transcriptional memory, under the regulation of JMJ14, and open the door for rational rewriting of chromatin to enhance organismal resilience.

plant biology↗