bioRxiv · 10.1101/2020.06.29.177261
Frame-by-frame annotation of video recordings using deep neural networks
Abstract
Video data are widely collected in ecological studies but manual annotation is a challenging and time-consuming task, and has become a bottleneck for scientific research. Classification models based on convolutional neural networks (CNNs) have proved successful in annotating images, but few applications have extended these to video classification. We demonstrate an approach that combines a standard CNN summarizing each video frame with a recurrent neural network (RNN) that models the temporal component of video. The approach is illustrated using two datasets: one collected by static video cameras detecting seal activity inside coastal salmon nets, and another collected by animal-borne cameras deployed on African penguins, used to classify behaviour. The combined RNN-CNN led to a relative improvement in test set classification accuracy over an image-only model of 25% for penguins (80% to 85%), and substantially improved classification precision or recall for four of six behaviour classes (12–17%). Image-only and video models classified seal activity with equally high accuracy (90%). Temporal patterns related to movement provide valuable information about animal behaviour, and classifiers benefit from including these explicitly. We recommend the inclusion of temporal information whenever manual inspection suggests that movement is predictive of class membership.Competing Interest StatementThe authors have declared no competing interest.View Full Text
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Conway, A. M., Durbach, I. N., McInnes, A., Harris, R. N.. 2020-06-29. Frame-by-frame annotation of video recordings using deep neural networks. https://doi.org/10.1101/2020.06.29.177261
Cite the original work for its findings. Save a collection to share your selection of sources.