Search bioRxiv⌕ Search

Biology subjects

Pipek, P.

Publications and source records attributed to Pipek, P..

2 recordsLinked to original sources

Can data mining from various internet platforms systematically accelerate detection of alien species invasions across the EU?

Invasive alien species (IAS) expansions are increasingly impacting the biodiversity and economy of Europe. To more effectively allocate the limited resources available for their management, it is pertinent to accelerate detection of IAS spread and distribution. One largely untapped secondary data source showing much potential lies in the automated tracking of internet activity such as IAS search intensity or mentions across different internet platforms. In this study, we tested if internet activity increases systematically when IAS expand into new EU countries utilizing the combined data of 88 invasive species from various internet platforms. In total, 14 internet platforms were screened and evaluated based on their database accessibility, mined data quality and utility for systematic IAS expansion tracking. We found that the procedure to obtain researcher access to minimal data required for IAS tracking (i.e., information about location, time and place) varies widely across platforms, and is particularly difficult without incurring significant costs for many of the larger ones (X, Google and Tiktok). From the explored species, more charismatic species (i.e., mammals) overall gained more online traction than more cryptic ones (i.e., plants), though online activity of the first proved a worse representation of real-world occurrence patterns. Moreover, while the final five selected internet platforms showed increased activity surrounding the year of invasion in many of the explored invasion scenarios (particularly Wikipedia and Facebook), inconsistencies between species groups, trends per platform and the large variability in data quality currently still hampers systematic integration of such data into existing databases. We conclude that combining IAS activity data from various internet platforms shows potential to accelerate IAS expansion detection across the EU (especially for fish, crustaceans, reptiles, birds and plants). However, incorporation in automated early warning systems is currently hampered by variable data quality, limited researcher access to online data and the few open, accurate and generalizable species classification algorithms with API access.

ecology↗

Large language models overcome the challenges of unstructured text data in ecology

The vast volume of currently available unstructured text data, such as research papers, news, and technical report data, shows great potential for ecological research. However, manual processing of such data is labour-intensive, posing a significant challenge. In this study, we aimed to assess the application of three state-of-the-art prompt-based large language models (LLMs), GPT 3.5, GPT 4, and LLaMA-2-70B, to automate the identification, interpretation, extraction, and structuring of relevant ecological information from unstructured textual sources. We focused on species distribution data from two sources: news outlets and research papers. We assessed the LLMs for four key tasks: classification of documents with species distribution data, identification of regions where species are recorded, generation of geographical coordinates for these regions, and supply of results in a structured format. GPT 4 consistently outperformed the other models, demonstrating a high capacity to interpret textual data and extract relevant information, with the percentage of correct outputs often exceeding 90% (average accuracy across tasks: 87-100%). Its performance also depended on the data source type and task, with better results achieved with news reports, in the identification of regions with species reports and presentation of structured output. Its predecessor, GPT 3.5, exhibited reasonably low accuracy across all tasks and data sources (average accuracy across tasks: 81-97%), whereas LLaMA-2-70B showed the worst performance (37- 73%). These results demonstrate the potential benefit of integrating prompt-based LLMs into ecological data assimilation workflows as essential tools to efficiently process large volumes of textual data.

ecology↗