Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.01.12.575330

Large language models help facilitate the automated synthesis of information on potential pest controllers

Abstract

The body of ecological literature, which informs much of our knowledge of the global loss of biodiversity, has been experiencing rapid growth in recent decades. The increasing difficulty to synthesise this literature manually has simultaneously resulted in a growing demand for automated text mining methods. Within the domain of deep learning, large language models (LLMs) have been the subject of considerable attention in recent years by virtue of great leaps in progress and a wide range of potential applications, however, quantitative investigation into their potential in ecology has so far been lacking. In this work, we analyse the ability of GPT-4 to extract information about invertebrate pests and pest controllers from abstracts of a body of literature on biological pest control, using a bespoke, zero-shot prompt. Our results show that the performance of GPT-4 is highly competitive with other state-of-the-art tools used for taxonomic named entity recognition and geographic location extraction tasks. On a held-out test set, we show that species and geographic locations are extracted with F1-scores of 99.8% and 95.3%, respectively, and highlight that the model is able to distinguish very effectively between the primary roles of interest (predators, parasitoids and pests). Moreover, we demonstrate the ability of the model to effectively extract and predict taxonomic information across various taxonomic ranks, and to automatically correct spelling mistakes. However, we do report a small number of cases of fabricated information (hallucinations). As a result of the current lack of specialised, pre-trained ecological language models, general-purpose LLMs may provide a promising way forward in ecology. Combined with tailored prompt engineering, such models can be employed for a wide range of text mining tasks in ecology, with the potential to greatly reduce time spent on manual screening and labelling of the literature.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Scheepens, D., Millard, J., Farrell, M., Newbold, T.. 2024-01-15. Large language models help facilitate the automated synthesis of information on potential pest controllers. https://doi.org/10.1101/2024.01.12.575330

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

PlanktonLake-CEREEP- A Freshwater Plankton Image Dataset with Semi-Automated Label Cleaning

Plankton plays a fundamental role in aquatic ecosystems, influencing biogeochemical cycles and serving as a key food source for many organisms. Recent high-throughput imaging technologies enable the rapid acquisition of large volumes of microscopic images, creating new opportunities for monitoring planktonic ecosystems. However, the manual processing and annotation of the vast amounts of data generated by these devices remain time-consuming tasks. In this context, machine learning-based classification models offer a promising solution. In this data paper, we introduce a new labeled freshwater plankton dataset comprising approximately 88,000 images distributed across 43 taxa. We also present the labeling assistance method we used to facilitate dataset annotation. Finally, we present a baseline based on a convolutional neural network (CNN), which achieves a classification accuracy of 93% on our dataset.

ecology↗

A training protocol for human classification of Asian elephant images from trail cameras

Trail cameras have become ubiquitous tools for ecological data collection over recent decades. Despite progress in the development of automated algorithms and artificial intelligence for image classification, our ability to process large volumes of data remain limited by the need for trained human observers to make refined judgements. We provide guidance on placement of trail cameras for observing Asian elephants (Elephas maximus) and outline a protocol for training and testing naive human observers in performing image classifications (age/sex class and group composition) that cannot yet be automated. This process can be used to develop a high-throughput workflow capable of extracting useful data from large volumes of images. Our training material consisted of 14,007 images collected from 6 trail cameras around Udawalawe National Park in Sri Lanka from 2017-2019. In the first stage, expert observers (n=3) trained a group of inexperienced participants (n=4), who engaged in an iterative process to develop a protocol document. The document was then tested on a second set of subjects (n=6) each of whom classified 350 test images in four separate sequential batches using quantitative measures of precision and accuracy. The test set was sampled from 54,435 images from an additional 25 cameras. When compared to expert observers, they achieved a fair level of precision (Fleiss' kappa = 0.247) and 82.6% accuracy. Our approach can usefully be extended to other species and contexts.

ecology↗

Forest belowground productivity and carbon allocation predominantly driven by soil properties rather than climate

Forests are threatened by a multitude of stressors, including anthropogenic disturbances and climate change. Assessing how forests will respond to these stressors requires a comprehensive understanding of net primary productivity (Npp), environmental constraints on growth, and adaptive capacity. A parameter of significant uncertainty is belowground Npp (bNpp), which can account for up to 80% of total Npp but is poorly estimated and rarely measured directly. We used a cross-biome dataset of direct, field-based measurements of aboveground and belowground primary productivity and 21 climatic and soil variables to identify potential constraints on bNpp and belowground carbon allocation in boreal and cold temperate forests. Soil variables, rather than climate variables, were the main drivers of bNpp and belowground allocation across biomes. The importance of soil variables suggests that soil nutrient dynamics, especially soil nutrient pool and flux variables, must be explicitly modeled to more accurately predict feedbacks between climate, productivity, and within-tree carbon allocation. Within biomes, environmental drivers of belowground allocation varied between low versus high allocation forests, indicating that environmental drivers are site-specific and the development of within-biome, site-scale classifications for forest ecosystems could be useful. Changes in soil variables, such as increasing soil nitrogen pools, caused abrupt and large decreases in bNpp for boreal, but not cold temperate forests. Threshold-like shifts indicate that boreal forests might have lower adaptive capacity and higher sensitivity to disturbances than cold temperate forests. With 70% of boreal forests characterized by low bNpp, disturbances such as anthropogenic nitrogen deposition could cause large-scale decreases in bNpp that could push these forests beyond their adaptive capacity.

ecology↗