bioRxiv · 10.64898/2026.07.02.735695
BehaviorScope-X: reusing pose-trained visual representations for full-video ethology
Abstract
Pose-estimation pipelines usually export keypoint coordinates and discard the intermediate visual representations learned to localize animals in a specific assay. We asked whether those discarded representations can be reused for full-video ethology. BehaviorScope-X tests this idea by treating a trained pose checkpoint as both a keypoint estimator and a reusable visual encoder: the pose model is run once to cache detections, keypoints, pose-derived social geometry, and frozen intermediate descriptors, after which compact temporal classifiers are trained on cached multimodal windows. Across MARS resident-intruder videos, cached pose-trained descriptors and pose-derived geometry provided complementary evidence for behavior decoding, recovering sustained behavioral episodes and local sequence structure while revealing a main limitation in dense short-bout regions. The same cache-and-classify design generalized across pose routes, including a MobileNetV3 backbone and a DeepLabCut SuperAnimal HRNet-W32 checkpoint, showing that standard pose workflows can expose behavior-relevant visual descriptors without giving up their keypoint-estimation role. We further tested the approach in Fly-v-Fly aggression, extending the analysis to a second species and shorter behavioral time scale, where sub-second events and annotation-boundary uncertainty limited strict bout recovery. End-to-end profiling showed that the workflow can operate near-real-time or real-time on consumer hardware. Together, these experiments support amortized pose vision as a practical strategy for reusing assay-trained pose models as stable sources of visual and geometric evidence for scalable behavioral analysis. Author summaryPose-estimation models are usually used to convert animal video into body landmarks, while the same models internal visual representations are discarded. We ask whether the model that estimates pose can also provide visual features for behavior analysis. Our workflow runs a trained pose model once, caches its landmarks, pose-derived interaction geometry, and internal visual features, and trains compact behavior classifiers on that cache. Across multiple pose backbones and pose-estimation workflows, these shared pose-and-visual signals supported full-video ethogram recovery without training a separate video network. This turns pose training into a reusable source of assay-specific visual and geometric evidence for scalable behavioral analysis. HighlightsO_LIA pose-estimation checkpoint is reused as both keypoint estimator and visual encoder. C_LIO_LIA single pose-inference pass caches keypoints, social geometry, and intermediate visual descriptors. C_LIO_LICached pose-derived visual and geometric evidence improves full-video behavior decoding beyond pose-derived features alone. C_LIO_LIBehaviorScope-X turns existing pose workflows into reusable full-video behavior-analysis pipelines. C_LI
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Augustine, F., Murray, V.. 2026-07-07. BehaviorScope-X: reusing pose-trained visual representations for full-video ethology. https://doi.org/10.64898/2026.07.02.735695
Cite the original work for its findings. Save a collection to share your selection of sources.