Skip to main content
QUICK REVIEW

[논문 리뷰] Scene Graphs: A Survey of Generations and Applications.

Xiaojun Chang, Pengzhen Ren|arXiv (Cornell University)|2021. 03. 17.
Multimodal Machine Learning Applications참고 문헌 201인용 수 11
한 줄 요약

이 종합적 서베이는 컴퓨터 시각 분야에서 장면 그래프 생성(SGG) 및 그 응용에 대한 체계적인 종합적 검토를 제공하며, 사전 지식이 있는지 여부에 따라 SGG를 위한 방법, 핵심 데이터셋, 그리고 시각적 질의 응답 및 이미지 편집과 같은 새로운 응용 분야를 포함한다. 이는 구조적 장면 이해 분야의 향후 연구를 위한 기초 참고 자료를 마련한다.

ABSTRACT

Scene graph is a structured representation of a scene that can clearly express the objects, attributes, and relationships between objects in the scene. As computer vision technology continues to develop, people are no longer satisfied with simply detecting and recognizing objects in images; instead, people look forward to a higher level of understanding and reasoning about visual scenes. For example, given an image, we want to not only detect and recognize objects in the image, but also know the relationship between objects (visual relationship detection), and generate a text description (image captioning) based on the image content. Alternatively, we might want the machine to tell us what the little girl in the image is doing (Visual Question Answering (VQA)), or even remove the dog from the image and find similar images (image editing and retrieval), etc. These tasks require a higher level of understanding and reasoning for image vision tasks. The scene graph is just such a powerful tool for scene understanding. Therefore, scene graphs have attracted the attention of a large number of researchers, and related research is often cross-modal, complex, and rapidly developing. However, no relatively systematic survey of scene graphs exists at present. To this end, this survey conducts a comprehensive investigation of the current scene graph research. More specifically, we first summarized the general definition of the scene graph, then conducted a comprehensive and systematic discussion on the generation method of the scene graph (SGG) and the SGG with the aid of prior knowledge. We then investigated the main applications of scene graphs and summarized the most commonly used datasets. Finally, we provide some insights into the future development of scene graphs. We believe this will be a very helpful foundation for future research on scene graphs.

연구 동기 및 목표

  • 컴퓨터 시각 분야에서 장면 그래프에 대한 종합적이고 체계적인 서베이가 부족한 점을 보완하기 위해.
  • 사전 지식을 통합한 방법을 포함하여 장면 그래프 생성(SGG) 방법을 분석하고 분류하기 위해.
  • 시각적 관계 탐지, 이미지 캡션 생성, VQA, 이미지 편집과 같은 작업에서 장면 그래프의 다양한 응용을 조사하기 위해.
  • 장면 그래프 연구를 위한 가장 널리 사용되는 벤치마크 데이터셋을 요약하기 위해.
  • 장면 그래프 기술을 발전시키기 위한 향후 연구 방향에 통찰을 제공하기 위해.

제안 방법

  • 이 서베이는 장면 그래프 연구에 대한 체계적인 문헌 검토를 수행하며, 생성 기법과 응용에 중점을 둔다.
  • SGG 방법을 시각 데이터에만 의존하는 것과 외부 지식(예: 사전 학습된 모델 또는 지식 기반 시스템)을 통합하는 것으로 분류한다.
  • 이 논문은 시각적 관계, 객체 속성, 그리고 관계 추론이 장면 그래프 구축에 미치는 역할을 분석한다.
  • 표준 벤치마크를 사용하여 다양한 SGG 프레임워크의 성능 및 설계 선택 사항을 평가한다.
  • annotation 방식, 규모, 작업 호환성 기준으로 데이터셋을 정리하고 비교한다.
  • 응용 분야 간의 추세와 과제를 통합하여, 다중 모odal 및 추론 중심 작업에 주목한다.

실험 결과

연구 질문

  • RQ1컴퓨터 시각에서 장면 그래프의 핵심 구성 요소와 정의는 무엇인가요?
  • RQ2사전 지식을 사용하는지 여부에 따라 장면 그래프 생성 방법은 어떻게 다릅니까?
  • RQ3VQA 및 이미지 편집과 같은 고급 시각 작업에서 장면 그래프의 주요 응용은 무엇인가요?
  • RQ4장면 그래프 모델의 훈련 및 평가에 가장 널리 사용되는 데이터셋은 무엇인가요?
  • RQ5장면 그래프 연구의 핵심 과제와 향후 연구 방향은 무엇인가요?

주요 결과

  • 장면 그래프는 객체, 속성, 그리고 그들 간의 관계를 명시적으로 모델링함으로써 고차원 시각 이해를 가능하게 한다.
  • 외부 지식을 통합한 SGG 방법은 순수하게 데이터 기반 접근 방식에 비해 복잡하거나 희귀한 관계에서 향상된 성능을 보인다.
  • 시각적 질의 응답 및 이미지 편집과 같은 응용은 구조적 장면 그래프 표현으로부터 크게 이점을 얻는다.
  • 이 서베이는 다중 모odal 및 추론 중심의 장면 그래프 응용으로 향하는 증가하는 추세를 식별한다.
  • VG 및 COCO-SceneGraph와 같은 여러 벤치마크 데이터셋이 훈련 및 평가에 널리 사용되지만, 규모와 annotation 품질 측면에서 다양성이 있다.
  • 표준화된 평가 프rotocol의 부족과 관계 추론의 복잡성은 여전히 분야의 주요 과제로 남아 있다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.