Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.11.12.623213

Capacity building needed to reap the benefits of access to biodiversity collections

Abstract

SummaryO_LIThis research examines biodiversity specimens from two areas of the Caribbean to understand patterns of collection and the roles of the people involved. Using open data from the Global Biodiversity Information Facility (GBIF) and Wikidata, we aimed to uncover geographic and historical trends in specimen use. This study aims to provide concrete evidence to guide collaboration between collection-holding institutions and the communities that need their resources most. C_LIO_LIWe analysed biodiversity specimens from Montserrat and the Cayman Islands in three steps. First, we extracted specimen data from GBIF, disambiguated collector names, and linked them to unique biographical entries. Next, we connected collectors to their publications and specimens. Finally, we analysed the modern use of these specimens through citation data, mapping author affiliations and research themes. C_LIO_LISpecimens are predominantly housed in the Global North and were initially used by their collectors, whose focus was largely on taxonomy and biogeography. With digitisation, use of these collections remains concentrated in the Global North and covers a broader range of subjects, although Brazil and China stand out as significant users of digital collection data compared to other similar countries. C_LIO_LIThe availability of open digital data from collections in the Global North has led to a substantial increase in the reuse of these data across biodiversity science. Nonetheless, most research using these data is still conducted in the Global North. For the non-monetary benefits of digitisation to extend to the countries of origin, capacity building in the Global South is crucial, Open Data alone are insufficient. C_LI Societal Impact StatementDigital biodiversity data from herbaria and museums hold significant potential for nature conservation in the Global South, yet many regions, like Montserrat and the Cayman Islands in the Caribbean, are, for multiple reasons, unable to fully leverage this information. This lack of skills and resources limits local conservation efforts, showing the need for more investment in training, facilities, and expertise. Although past funding has helped improve coordination and build skills, our findings show that more work is needed to make sure conservation in these biodiverse areas can continue in the long term.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Groom, Q., Meeus, S., Barrios, S., Childs, C., Clubbe, C., Corbett, E., Francis, S., Gray, A., Harding, L., Jackman, A., Machin, R., McGovern, A., Pienkowski, M., Ryan, D., Sealys, C., Wensink, C., Peyton, J.. 2024-11-15. Capacity building needed to reap the benefits of access to biodiversity collections. https://doi.org/10.1101/2024.11.12.623213

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

Employing Microsoft Sway as a Repository for Data Generated in Pharmaceutics Experimental Courses: A Case study on Topical Formulation Characterization

Microsoft Sway is an easy-to-use online creative software developed by Microsoft. It has many advantages as follows: easy to operate and does not require long-term practice and learning; network links can be created and accessed by devices such as Personal Computer (PC) and mobile terminals; Educators and Learners are allowed to browse this repository at anytime to facilitate the application prospects of mining relevant data. Therefore, the author proposes that Microsoft Sway may be a very promising data repository for pharmaceutical experimental teaching courses, and is expected to be extended to undergraduate experimental teaching in other natural science fields.

scientific communication and education↗

Bridging the Python Training Gap for Bioscientists in Brazil: Improvements and Challenges

The rapid evolution of high-throughput technologies in biosciences has generated diverse and voluminous datasets, requiring bioscientists to develop data manipulation and analysis skills. Python, known for its versatility and powerful libraries, has become a crucial tool for managing these datasets. However, there is a significant lack of programming training for bioscientists in many countries. To address this knowledge gap among scientists in Brazil, the Brazilian Python Workshop for Biological Data was introduced several years ago, focusing on basic programming concepts and data handling techniques using popular Python libraries. Despite the progress and positive feedback from earlier editions, challenges persisted, necessitating continuous adaptation and improvement to meet the evolving needs of bioscientists.This work describes the advancements made in the 2021 and 2022 editions of the workshop and discusses new suggestions for its ongoing enhancement. Key innovations were introduced in the workshop structure and coordination, including the creation of new committees and the establishment of a code of conduct. Feedback forms were updated to enable real-time adjustments during the event, improving its overall effectiveness. The workshop also expanded its reach by increasing geographical diversity among participants. New didactic strategies, such as pair-teaching, code clubs, and the integration of information and communication technologies (ICTs), were implemented to enhance learning outcomes. Programming best practices and scientific reproducibility were emphasized through talks and hands-on activities, guided by PEP8 conventions. Furthermore, efforts to enhance scientific dissemination were intensified, with an increased presence on social media and participation in international scientific events and communication networks. Finally, we present updated recommendations for students, researchers, and educators interested in organizing and promoting similar events, building on those previously described.

scientific communication and education↗