[Paper Review] GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
The paper evaluates four multimodal LLMs (Video-LLaMA, MiniGPT4-Video, Video-Chat2, GPT-4o) on piglet activity understanding using close-up and full-shot videos, focusing on counting, actor referring, semantic correspondence, time perception, and robustness, with GPT-4o achieving strongest overall performance especially in time perception and robustness.
Animal ethology is an crucial aspect of animal research, and animal behavior labeling is the foundation for studying animal behavior. This process typically involves labeling video clips with behavioral semantic tags, a task that is complex, subjective, and multimodal. With the rapid development of multimodal large language models(LLMs), new application have emerged for animal behavior understanding tasks in livestock scenarios. This study evaluates the visual perception capabilities of multimodal LLMs in animal activity recognition. To achieve this, we created piglet test data comprising close-up video clips of individual piglets and annotated full-shot video clips. These data were used to assess the performance of four multimodal LLMs-Video-LLaMA, MiniGPT4-Video, Video-Chat2, and GPT-4 omni (GPT-4o)-in piglet activity understanding. Through comprehensive evaluation across five dimensions, including counting, actor referring, semantic correspondence, time perception, and robustness, we found that while current multimodal LLMs require improvement in semantic correspondence and time perception, they have initially demonstrated visual perception capabilities for animal activity recognition. Notably, GPT-4o showed outstanding performance, with Video-Chat2 and GPT-4o exhibiting significantly better semantic correspondence and time perception in close-up video clips compared to full-shot clips. The initial evaluation experiments in this study validate the potential of multimodal large language models in livestock scene video understanding and provide new directions and references for future research on animal behavior video understanding. Furthermore, by deeply exploring the influence of visual prompts on multimodal large language models, we expect to enhance the accuracy and efficiency of animal behavior recognition in livestock scenarios through human visual processing methods.
Motivation & Objective
- Investigate how current multimodal LLMs perform in piglet activity understanding in livestock video data.
- Compare models across multiple visual-perception dimensions, including counting, actor identification, semantic understanding, temporal perception, and robustness.
- Assess the impact of visual prompting (close-up vs. full-shot with marks) on model performance.
- Provide guidance on prompts and evaluation frameworks for livestock scene understanding using multimodal LLMs.
Proposed method
- Construct a piglet video dataset with close-up and full-shot clips and annotated behaviors (standing, lying, feeding, drinking, moving, socializing).
- Track individual piglets in 1-minute segments using the DEVA method to ensure consistent actor labeling.
- Develop two visual-prompt templates (text-only for close-up, text + visual cues for full-shot with marks) to probe models.
- Evaluate four multimodal LLMs (Video-LLaMA 7B, MiniGPT4-Video, Video-Chat2, GPT-4o) under five metrics using a 0–5 scoring scale normalized per video clip.
- Define five metrics: counting, actor referring, semantic correspondence, time perception, robustness, with rules to compute scores from model outputs.
Experimental results
Research questions
- RQ1How do state-of-the-art multimodal LLMs perform on piglet activity understanding in both close-up and full-shot video settings?
- RQ2Which aspects (counting, actor referring, semantic correspondence, time perception, robustness) limit current models in livestock scene understanding?
- RQ3Does visual prompting (close-up vs. full-shot with marks) affect the models' performance across the identified metrics?
- RQ4Among the tested models, which demonstrates the strongest temporal understanding and robustness in piglet behavior tasks?
Key findings
- GPT-4o generally outperforms peers across several metrics, especially in time perception and robustness on full-shot clips.
- Video-Chat2 and GPT-4o show stronger semantic correspondence and time perception on close-up clips compared to other models.
- All four models struggle with semantic correspondence, indicating room for improvement in interpreting piglet behaviors from video data.
- Close-up videos tend to yield better results for video understanding tasks in multimodal LLMs than full-shot videos with scene cues.
- Counting performance is weak for most models, with GPT-4o showing the largest gains among the tested models on full-shot data.
- Overall, multimodal LLMs have initial visual-perception capabilities for livestock activity understanding but require further optimization for specialized livestock-scene tasks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.