BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//University of Iceland//AI Centre//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
BEGIN:VEVENT
UID:03568865-8c7e-4a8b-a354-3b2c59324cec@ai.hi.is
DTSTAMP:20261011T141338Z
DTSTART:20260224T090000Z
DTEND:20260224T100000Z
SUMMARY:How Far Can Vision-Language Models Take You?  Zero-Shot Detection a
 nd Data-Efficient Species Classification in Videos
DESCRIPTION:Can pre-trained vision-language models replace task-specific tr
 aining for species classification? How can we leverage the temporal nature
  of videos to improve performance? We explore these questions using CLIP/S
 igLIP architectures as both zero-shot classifiers and feature extractors. 
 For detection\, we show that prompt engineering alone achieves 99.1% accur
 acy without any labeled data for training. For the species classification\
 , we combine frozen SigLIP embeddings with temporal feature aggregation an
 d linear classifiers\, achieving 96.8% macro F1\, outperforming a fine-tun
 ed ResNet-50 while using less than 1/3 of the training samples. We demonst
 rate the approach on underwater fish species classification from Icelandic
  river monitoring footage\, but the method could generalize to a broader s
 et of video classification tasks with limited labeled data.\n\nhttps://ai.
 hi.is/is/events/colloquium-in-statistics-and-ai-12/
LOCATION:Íslensk erfðagreining\, Tjarnarsalur
URL:https://ai.hi.is/is/events/colloquium-in-statistics-and-ai-12/
STATUS:CONFIRMED
END:VEVENT
END:VCALENDAR
