bioRxiv · 10.1101/2021.09.13.460161
Embracing imperfection: machine-assisted invertebrate classification in real-world datasets
Abstract
O_LIDespite growing concerns over the health of global invertebrate diversity, terrestrial invertebrate monitoring efforts remain poorly geographically distributed. Machine-assisted classification has been proposed as a potential solution to quickly gather large amounts of data; however, previous studies have often used unrealistic or idealized datasets to train their models. C_LIO_LIIn this study, we describe a practical methodology for including machine learning in ecological data acquisition pipelines. Here we train and test machine learning algorithms to classify over 56,000 bulk terrestrial invertebrate specimens from morphometric data and contextual metadata. All vouchered specimens were collected in pitfall traps by the National Ecological Observatory Network (NEON) at 27 locations across the United States in 2016. Specimens were photographed, and morphometric data was extracted as feature vectors using ImageJ. Issues stemming from inconsistent taxonomic label specificity were resolved by making classifications at the lowest identified taxonomic level (LITL). Taxa with too few specimens to be included in the training dataset were classified by the model using zero-shot classification. C_LIO_LIWhen classifying specimens that were known and seen by our models, we reached an accuracy of 72.7% using extreme gradient boosting (XGBoost) at the LITL. Models that were trained without contextual metadata underperformed models with contextual metadata by an average of 7.2%. We also classified invertebrate taxa that were unknown to the model using zero-shot classification, with an accuracy of 39.4%, resulting in an overall accuracy of 71.5% across the entire NEON dataset. C_LIO_LIThe general methodology outlined here represents a realistic application of machine learning as a tool for ecological studies. Hierarchical and LITL classifications allow for flexible taxonomic specificity at the input and output layers. These methods also help address the long tail problem of underrepresented taxa missed by machine learning models. Finally, we encourage researchers to consider more than just morphometric data when training their models, as we have shown that the inclusion of contextual metadata can provide significant improvements to accuracy. C_LI
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Blair, J. D., Weiser, M. D., deBeurs, K., Kaspari, M., Siler, C. D., Marshall, K. E.. 2021-09-15. Embracing imperfection: machine-assisted invertebrate classification in real-world datasets. https://doi.org/10.1101/2021.09.13.460161
Cite the original work for its findings. Save a collection to share your selection of sources.