Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.07.02.735695

BehaviorScope-X: reusing pose-trained visual representations for full-video ethology

Abstract

Pose-estimation pipelines usually export keypoint coordinates and discard the intermediate visual representations learned to localize animals in a specific assay. We asked whether those discarded representations can be reused for full-video ethology. BehaviorScope-X tests this idea by treating a trained pose checkpoint as both a keypoint estimator and a reusable visual encoder: the pose model is run once to cache detections, keypoints, pose-derived social geometry, and frozen intermediate descriptors, after which compact temporal classifiers are trained on cached multimodal windows. Across MARS resident-intruder videos, cached pose-trained descriptors and pose-derived geometry provided complementary evidence for behavior decoding, recovering sustained behavioral episodes and local sequence structure while revealing a main limitation in dense short-bout regions. The same cache-and-classify design generalized across pose routes, including a MobileNetV3 backbone and a DeepLabCut SuperAnimal HRNet-W32 checkpoint, showing that standard pose workflows can expose behavior-relevant visual descriptors without giving up their keypoint-estimation role. We further tested the approach in Fly-v-Fly aggression, extending the analysis to a second species and shorter behavioral time scale, where sub-second events and annotation-boundary uncertainty limited strict bout recovery. End-to-end profiling showed that the workflow can operate near-real-time or real-time on consumer hardware. Together, these experiments support amortized pose vision as a practical strategy for reusing assay-trained pose models as stable sources of visual and geometric evidence for scalable behavioral analysis. Author summaryPose-estimation models are usually used to convert animal video into body landmarks, while the same models internal visual representations are discarded. We ask whether the model that estimates pose can also provide visual features for behavior analysis. Our workflow runs a trained pose model once, caches its landmarks, pose-derived interaction geometry, and internal visual features, and trains compact behavior classifiers on that cache. Across multiple pose backbones and pose-estimation workflows, these shared pose-and-visual signals supported full-video ethogram recovery without training a separate video network. This turns pose training into a reusable source of assay-specific visual and geometric evidence for scalable behavioral analysis. HighlightsO_LIA pose-estimation checkpoint is reused as both keypoint estimator and visual encoder. C_LIO_LIA single pose-inference pass caches keypoints, social geometry, and intermediate visual descriptors. C_LIO_LICached pose-derived visual and geometric evidence improves full-video behavior decoding beyond pose-derived features alone. C_LIO_LIBehaviorScope-X turns existing pose workflows into reusable full-video behavior-analysis pipelines. C_LI

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Augustine, F., Murray, V.. 2026-07-07. BehaviorScope-X: reusing pose-trained visual representations for full-video ethology. https://doi.org/10.64898/2026.07.02.735695

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Functional validation of allele-specific LMNB1 silencing in patient-derived astrocytes as a therapeutic option for Autosomal Dominant Leukodystrophy

Adult-onset Autosomal Dominant Leukodystrophy (ADLD) is a rare fatal leukodystrophy caused by increased LMNB1 gene dosage, most commonly resulting from duplication of the LMNB1 locus. Because ADLD is a gene dosage disorder, selective reduction of pathological LMNB1 expression represents a rational therapeutic strategy. Although allele-specific RNA interference has previously been shown to lower LMNB1 levels in patient-derived fibroblasts and directly reprogrammed neurons, its therapeutic effects have not been evaluated in disease-relevant human glial cells or using functional efficacy endpoints. Here, we established human induced pluripotent stem cell-derived astrocytes from ADLD patients as a human glial model in which to validate allele-specific LMNB1 silencing across molecular, cellular, and functional readouts. ADLD astrocytes recapitulated increased LMNB1 expression and characteristic nuclear abnormalities and displayed transcriptional alterations affecting extracellular matrix organization, calcium homeostasis, metabolism and RNA processing. Functionally, these cells also exhibited functional phenotypes suitable for therapeutic evaluation: astrocyte-conditioned medium impaired the viability of both murine and human oligodendroglial cultures, while conditioned-medium and direct astrocyte-seeding paradigms revealed impaired post-lesion myelin recovery in lysolecithin-treated cerebellar organotypic slices. Allele-specific LMNB1 silencing restored physiological LMNB1 levels, corrected nuclear abnormalities, attenuated astrocyte-mediated oligodendroglial toxicity, improved post-lesion myelin recovery, and was associated with selective transcriptional programs associated with extracellular support and cholesterol metabolism. Together, these findings provide molecular, cellular, and functional validation of allele-specific LMNB1 dosage correction in patient-derived human astrocytes and offer key support for LMNB1-lowering strategies in disease-relevant human glial cells.

neuroscience↗

Perceptual integration of multisensory haptic, visual, and auditory feedback for roughness discrimination in augmented reality

Understanding how our different senses interact to shape our perception is essential to design realistic and immersive virtual and augmented reality (VR/AR) experiences. The present study investigated how roughness perception can be modulated through haptic, visual, and auditory cues in AR using a vibrotactile wristband. Participants compared virtual textures varying in vibration frequency/amplitude, visual grain size, and friction sound. Results revealed strong linear relationships between stimulus parameters and perceived roughness, with haptic frequency and visual cues driving the highest discrimination performance. Adding non-informative sensory feedback reduced perceptual sensitivity, acting as noise. Individual differences emerged: participants who rated haptic as the easiest modality showed greater sensitivity to haptic variations, while visual-reliant participants performed better with visual cues. We conclude that roughness in AR can be systematically manipulated, but is vulnerable to perceptual interference from irrelevant inputs, where our work provides actionable insights for implementing optimized and adaptive AR/VR interfaces.

neuroscience↗

Structural and functional MRI signatures of Gambling Disorder: a case-control study

Gambling disorder (GD) is a behavioural addiction that may help identify addiction-related neural features without the direct neurobiological effects of a primary substance of dependence. We examined regional grey matter volume (GMV) and resting-state functional connectivity (rsFC) in the same well-characterised sample. Eighteen men with GD and 21 matched healthy controls underwent high-resolution structural and resting-state functional MRI. GMV was quantified across 214 cortical and subcortical regions, and seed-based rsFC analyses focused on striatal subdivisions and mesocorticolimbic regions. Group differences were evaluated using permutation testing and cluster-corrected mixed-effects modelling. GD was associated with lower GMV in the ventromedial prefrontal cortex, orbitofrontal regions and other cortical and subcortical areas, alongside higher GMV in a subset of limbic and default-mode regions. Participants with GD also showed lower connectivity between the limbic striatum and the hippocampus, thalamus and putamen. In exploratory analyses, somatomotor connectivity was positively associated with gambling severity (Problem Gambling Severity Index: Spearman's rho = 0.71, p = 0.003, false-discovery-rate-adjusted q = 0.016). Structural and functional findings overlapped spatially in regions associated with valuation, memory, reward and habit formation, but regional GMV did not mediate group differences in rsFC. These findings are broadly consistent with corticostriatal models of GD and identify candidate circuit-level differences for independent replication. Larger, more diverse and longitudinal samples are required to establish their reproducibility, temporal direction and clinical relevance.

neuroscience↗