A Vision-Language Model as a Teacher for Bird Vocalization Detection
Time-frequency detection of bird vocalizations is an important step toward turning weakly annotated field recordings into usable data for studying avian communication. Supervised detectors trained on human expert annotations scale poorly across species and recording conditions, so we introduce a teacher-student setup in which the teacher, a vision-language model, labels bounding boxes on spectrograms of citizen-scientist recordings, and those labels train a student, a self-supervised bioacoustic encoder. We find that the student generally exceeds both the teacher and supervised models trained on human annotations and performs strongly on held-out datasets for both time-frequency and onset-offset localization. YOLO detectors trained on our teacher labels match those trained on human annotations, suggesting that VLM labels can substitute for costly expert labeling.