Skip to main content
QUICK REVIEW

[论文解读] GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding

Yiqi Wu, Xiaodan Hu|arXiv (Cornell University)|Jun 14, 2024
Animal Behavior and Welfare Studies被引用 6
一句话总结

该论文在小猪幼崽活动理解上评估了四种多模态大模型(Video-LLaMA、MiniGPT4-Video、Video-Chat2、GPT-4o),在近景与全景视频的基础上,聚焦于计数、演员指称、语义对应、时间感知和鲁棒性,GPT-4o 在总体性能上最强,特别是在时间感知和鲁棒性方面。

ABSTRACT

Animal ethology is an crucial aspect of animal research, and animal behavior labeling is the foundation for studying animal behavior. This process typically involves labeling video clips with behavioral semantic tags, a task that is complex, subjective, and multimodal. With the rapid development of multimodal large language models(LLMs), new application have emerged for animal behavior understanding tasks in livestock scenarios. This study evaluates the visual perception capabilities of multimodal LLMs in animal activity recognition. To achieve this, we created piglet test data comprising close-up video clips of individual piglets and annotated full-shot video clips. These data were used to assess the performance of four multimodal LLMs-Video-LLaMA, MiniGPT4-Video, Video-Chat2, and GPT-4 omni (GPT-4o)-in piglet activity understanding. Through comprehensive evaluation across five dimensions, including counting, actor referring, semantic correspondence, time perception, and robustness, we found that while current multimodal LLMs require improvement in semantic correspondence and time perception, they have initially demonstrated visual perception capabilities for animal activity recognition. Notably, GPT-4o showed outstanding performance, with Video-Chat2 and GPT-4o exhibiting significantly better semantic correspondence and time perception in close-up video clips compared to full-shot clips. The initial evaluation experiments in this study validate the potential of multimodal large language models in livestock scene video understanding and provide new directions and references for future research on animal behavior video understanding. Furthermore, by deeply exploring the influence of visual prompts on multimodal large language models, we expect to enhance the accuracy and efficiency of animal behavior recognition in livestock scenarios through human visual processing methods.

研究动机与目标

  • 研究当前的多模态大语言模型在牲畜视频数据中的小猪幼崽活动理解表现
  • 在多种视觉感知维度上比较模型,包括计数、演员识别、语义理解、时间感知和鲁棒性
  • 评估视觉提示(近景 vs. 带标记的全景)对模型性能的影响
  • 就使用多模态LLMs进行牲畜场景理解的提示设计与评估框架提供指导

提出的方法

  • 构建包含近景与全景片段的小猪幼崽视频数据集,并对行为进行注释(站立、躺卧、进食、喝水、移动、社交)
  • 使用 DEVA 方法在1分钟片段内跟踪单个猪仔,确保参与者标注的一致性
  • 开发两种视觉提示模板(近景文本优先,带标记的全景文本+视觉线索)以探测模型
  • 在五个指标下对四种多模态大模型(Video-LLaMA 7B、MiniGPT4-Video、Video-Chat2、GPT-4o)进行评估,使用0–5的得分尺度并对每个视频片段归一化
  • 定义五个指标:计数、演员指称、语义对应、时间感知、鲁棒性,并制定从模型输出计算分数的规则

实验结果

研究问题

  • RQ1最先进的多模态大语言模型在近景和全景视频设置下对小猪幼崽活动理解的表现如何?
  • RQ2哪些方面(计数、演员指称、语义对应、时间感知、鲁棒性)限制了当前在牲畜场景理解中的模型?
  • RQ3视觉提示(近景 vs. 带标记的全景)是否影响模型在上述指标上的表现?
  • RQ4在测试的模型中,哪一个在小猪行为任务中展现出最强的时间理解和鲁棒性?

主要发现

  • GPT-4o 在多个指标上通常优于同侪,尤其是在全景片段的时间感知和鲁棒性方面。
  • Video-Chat2 和 GPT-4o 在近景片段上的语义对应和时间感知优于其他模型。
  • 所有四个模型在语义对应方面仍存在挑战,表明在从视频数据中解释小猪行为方面仍有提升空间。
  • 近景视频在多模态LLMs的視頻理解任务中往往比带场景线索的全景视频取得更好结果。
  • 大多数模型的计数表现较弱,在全景数据上,GPT-4o 显示出最大的提升。
  • 总体而言,多模态LLMs 已具备初步的视觉感知能力以理解牲畜活动,但在针对专业牲畜场景任务上仍需进一步优化。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。