[논문 리뷰] GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
이 논문은 close-up 및 full-shot 비디오를 활용하여 네 가지 다중모달 LLM(Video-LLaMA, MiniGPT4-Video, Video-Chat2, GPT-4o)을 쥐새끼의 활동 이해에 대해 평가하며, 계수 세기, 배우 지칭, 의미 일치, 시간 지각, 견고성에 중점을 두고, GPT-4o가 특히 시간 지각 및 견고성에서 가장 강력한 전반 성능을 달성한다.
Animal ethology is an crucial aspect of animal research, and animal behavior labeling is the foundation for studying animal behavior. This process typically involves labeling video clips with behavioral semantic tags, a task that is complex, subjective, and multimodal. With the rapid development of multimodal large language models(LLMs), new application have emerged for animal behavior understanding tasks in livestock scenarios. This study evaluates the visual perception capabilities of multimodal LLMs in animal activity recognition. To achieve this, we created piglet test data comprising close-up video clips of individual piglets and annotated full-shot video clips. These data were used to assess the performance of four multimodal LLMs-Video-LLaMA, MiniGPT4-Video, Video-Chat2, and GPT-4 omni (GPT-4o)-in piglet activity understanding. Through comprehensive evaluation across five dimensions, including counting, actor referring, semantic correspondence, time perception, and robustness, we found that while current multimodal LLMs require improvement in semantic correspondence and time perception, they have initially demonstrated visual perception capabilities for animal activity recognition. Notably, GPT-4o showed outstanding performance, with Video-Chat2 and GPT-4o exhibiting significantly better semantic correspondence and time perception in close-up video clips compared to full-shot clips. The initial evaluation experiments in this study validate the potential of multimodal large language models in livestock scene video understanding and provide new directions and references for future research on animal behavior video understanding. Furthermore, by deeply exploring the influence of visual prompts on multimodal large language models, we expect to enhance the accuracy and efficiency of animal behavior recognition in livestock scenarios through human visual processing methods.
연구 동기 및 목표
- 현재 다중모달 LLM이 가축 비디오 데이터에서 쥐새끼 활동 이해를 어떻게 수행하는지 조사한다.
- 계수 세기, 배우 식별, 의미 이해, 시간 지각, 견고성 등 다양한 시각 인지 차원에서 모델을 비교한다.
- 시각 프롬프트의 영향(마크가 포함된 full-shot 대 close-up)을 모델 성능에 미치는 영향을 평가한다.
- 다중모달 LLM을 사용한 가축 현장 이해를 위한 프롬프트 및 평가 프레임워크에 대한 지침을 제공한다.
제안 방법
- close-up 및 full-shot 클립과 주석된 행동(서 있기, 눕기, 먹이주기, 물 마시기, 이동, 사회적 상호작용)을 포함하는 돼지새끼 비디오 데이터세트를 구성한다.
- DEVA 방법을 사용하여 1분 단위로 개별 돼지새끼를 추적하여 일관된 배우 표기를 보장한다.
- 모델을 탐색하기 위한 두 가지 시각 프롬프트 템플릿을 개발한다( close-up의 경우 텍스트만, full-shot의 경우 텍스트 + 시각 단서로 마크 포함 ).
- 다섯 가지 지표를 0–5 척도로 비디오 클립당 정규화하여 평가한다( Video-LLaMA 7B, MiniGPT4-Video, Video-Chat2, GPT-4o의 네 가지 다중모달 LLM ).
- 지표 다섯 가지를 정의한다: counting, actor referring, semantic correspondence, time perception, robustness, 모델 출력에서 점수를 계산하는 규칙을 설정한다.
실험 결과
연구 질문
- RQ1최신 다중모달 LLM이 close-up 및 full-shot 비디오 설정에서 쥐새끼의 활동 이해를 어떻게 수행하는가?
- RQ2현재 모델을 제한하는 측면은 무엇이며, 가축 현장 이해의 어떤 측면(계수, 배우 지칭, 의미 일치, 시간 지각, 견고성)인가?
- RQ3시각 프롬프트(close-up 대 full-shot with marks)가 식별된 지표 전반에 걸쳐 모델의 성능에 영향을 주는가?
- RQ4테스트된 모델들 중 쥐새끼 행동 과제에서 가장 강한 시간 이해도와 견고성을 보이는 모델은 무엇인가?
주요 결과
- GPT-4o는 일반적으로 여러 지표에서 타 모델을 능가하는 편이며, 특히 full-shot 클립에서의 시간 지각 및 견고성에서 두드러진다.
- Video-Chat2와 GPT-4o는 close-up 클립에서 다른 모델들에 비해 의미 일치 및 시간 지각이 더 강한 것으로 나타난다.
- 네 모델 모두 의미 일치에 어려움을 겪으며, 비디오 데이터로부터의 쥐새끼 행동 해석에 개선 여지가 있음을 시사한다.
- close-up 비디오가 다중모달 LLM의 비디오 이해 과제에 대해 full-shot 비디오보다 더 나은 결과를 내는 경향이 있다.
- 대부분의 모델에서 counting 성능은 약하며, full-shot 데이터에서 GPT-4o가 측정된 모델들 가운데 가장 큰 향상을 보인다.
- 전반적으로 다중모달 LLM은 가축 활동 이해를 위한 초기 시각 인지 능력을 갖추고 있지만, 특수한 가축 현장 과제에 맞춘 추가 최적화가 필요하다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.