Skip to main content
QUICK REVIEW

[논문 리뷰] Technical Report of the Video Event Reconstruction and Analysis (VERA) System -- Shooter Localization, Models, Interface, and Beyond

Junwei Liang, Jay D. Aronson|arXiv (Cornell University)|2019. 05. 26.
Video Analysis and Summarization참고 문헌 25인용 수 5
한 줄 요약

VERA는 메타데이터가 없는 비정형 소셜 미디어 영상에서 총격 사건을 정확하게 국지화할 수 있도록 기계학습과 물리 기반 기법을 결합한 시스템이다. 영상 간 동기화, 음성 신호를 통한 총성 탐지, 초음속 탄도 및 음파 전파 물리 모델을 적용하여, 3개의 영상과 첫 번째 총성만으로 2017년 라스베이거스 총격 사건에서 총격자 위치를 정확히 특정했다.

ABSTRACT

Every minute, hundreds of hours of video are uploaded to social media sites and the Internet from around the world. This material creates a visual record of the experiences of a significant percentage of humanity and can help illuminate how we live in the present moment. When properly analyzed, this video can also help analysts to reconstruct events of interest, including war crimes, human rights violations, and terrorist acts. Machine learning and computer vision can play a crucial role in this process. In this technical report, we describe the Video Event Reconstruction and Analysis (VERA) system. This new tool brings together a variety of capabilities we have developed over the past few years (including video synchronization and geolocation to order unstructured videos lacking metadata over time and space, and sound recognition algorithms) to enable the reconstruction and analysis of events captured on video. Among other uses, VERA enables the localization of a shooter from just a few videos that include the sound of gunshots. To demonstrate the efficacy of this suite of tools, we present the results of estimating the shooter's location of the Las Vegas Shooting in 2017 and show that VERA accurately predicts the shooter's location using only the first few gunshots. We then point out future directions that can help improve the system and further reduce unnecessary human labor in the process. All of the components of VERA run through a web interface that enables human-in-the-loop verification to ensure accurate estimations. All relevant source code, including the web interface and machine learning models, is freely available on Github. We hope that researchers and software developers will be inspired to improve and expand this system moving forward to better meet the needs of human rights and public safety.

연구 동기 및 목표

  • 메타데이터가 제거된 비정형 소셜 미디어 영상에서 실제 사건을 재구성하는 도전 과제 해결.
  • 영상에 시간 또는 GPS 메타데이터가 없더라도 총격 사건에서 총격자 국지화를 정확하게 가능하게 하기.
  • 자동화된 영상 동기화, 총성 탐지, 물리 기반 국지화를 통합하여 수동 분석에 대한 의존도 감소.
  • 사람이 개입하는 검증 방식을 통합하여 인간 권리 조사 및 공공안전 활동을 지원하는 확장 가능한 오픈소스 도구 제공.
  • 생산용으로 사용 가능한 웹 기반 시스템을 제공하여 기계학습, 음성 처리, 지리공간 모델링을 융합해 사건 재구성 수행.

제안 방법

  • 시각적 및 청각적 신호를 활용해 비정형 영상을 글로벌 타임라인에 자동으로 동기화하는 반자동 영상 동기화 기법 적용.
  • 기계학습 기반 총성 탐지를 통해 음성 세그먼트 내에서 발사 소리와 충격파 소리를 식별.
  • 초음속 탄도 물리 모델을 활용해 다수의 카메라 위치에서 도착한 음성 신호 간 도착 시간 차이 추정.
  • 총성 구성 요소(발사 소리 및 충격파)의 도착 시간 차이를 활용해 쌍곡선 삼각측량 기법으로 가능한 총격자 위치 추정.
  • 지도상에 열지도 및 쌍곡선 신뢰 영역을 시각화하여 잠재적 총격자 위치 표시하며, 사용자가 편집 가능한 카메라 위치 마커 제공.
  • 실시간 처리 및 진행 상황 업데이트를 위해 PHP 프론트엔드와 Python 백엔드 간 암호화된 통신을 구현한 웹 인터페이스 통합.

실험 결과

연구 질문

  • RQ1메타데이터 없이도 비정형 소셜 미디어 영상 간 효과적인 동기화 및 시간적 정렬이 가능한가?
  • RQ2공개 영상에서 저품질, 노이즈가 많은 음성 신호에서 총격 사건을 얼마나 정확하게 탐지하고 시간적으로 국지화할 수 있는가?
  • RQ3초음속 탄도 및 음파 전파 물리 모델을 활용해 음성 신호만으로 총격자 국지화 정확도를 얼마나 향상시킬 수 있는가?
  • RQ4사람이 개입하는 시스템이 최소한의 사용자 입력으로도 실세계 사건 재구성의 오류를 줄이고 신뢰도를 높일 수 있는가?
  • RQ5최소한의 인프라 부담으로도 인간 권리 및 공공안전 응용을 지원할 수 있는 확장 가능한 오픈소스 시스템 아키텍처는 어떻게 설계할 수 있는가?

주요 결과

  • VERA는 공개된 3개의 영상과 첫 번째 총성만으로 2017년 라스베이거스 총격 사건에서 총격자를 성공적으로 국지화했다.
  • 시스템이 추정한 총격자 위치는 음성 도착 시간 차이를 기반으로 하였으며, 실제로는 만달레이 베이 호텔의 북쪽 날개에 해당하여 정확도가 높았다.
  • 다수의 영상에서 유도된 열지도 및 쌍곡선 삼각측량 시각화 결과는 실제 총격자 위치로 수렴하여 높은 공간 정확도를 입증했다.
  • 최소한의 메타데이터에 의존하여 오직 음성 신호 처리와 물리 모델만으로도 정확한 결과를 도출했다.
  • 웹 인터페이스를 통해 효율적인 사람-개입 기반 검증이 가능하여 사용자가 카메라 위치를 수정하고 국지화 신뢰도를 향상시킬 수 있었다.
  • 모든 소스 코드, 모델, 웹 인터페이스가 GitHub에 공개되어 있어 인권 및 공공안전 분야의 커뮤니티 확장 및 개선을 위한 활용이 가능하다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.