[论文解读] Scene Graphs: A Survey of Generations and Applications.
本综述对场景图生成(SGG)及其在计算机视觉中的应用进行了全面、系统的回顾,涵盖无先验知识与有先验知识的SGG方法、关键数据集,以及视觉问答和图像编辑等新兴应用。该综述为未来结构化场景理解研究奠定了基础性参考。
Scene graph is a structured representation of a scene that can clearly express the objects, attributes, and relationships between objects in the scene. As computer vision technology continues to develop, people are no longer satisfied with simply detecting and recognizing objects in images; instead, people look forward to a higher level of understanding and reasoning about visual scenes. For example, given an image, we want to not only detect and recognize objects in the image, but also know the relationship between objects (visual relationship detection), and generate a text description (image captioning) based on the image content. Alternatively, we might want the machine to tell us what the little girl in the image is doing (Visual Question Answering (VQA)), or even remove the dog from the image and find similar images (image editing and retrieval), etc. These tasks require a higher level of understanding and reasoning for image vision tasks. The scene graph is just such a powerful tool for scene understanding. Therefore, scene graphs have attracted the attention of a large number of researchers, and related research is often cross-modal, complex, and rapidly developing. However, no relatively systematic survey of scene graphs exists at present. To this end, this survey conducts a comprehensive investigation of the current scene graph research. More specifically, we first summarized the general definition of the scene graph, then conducted a comprehensive and systematic discussion on the generation method of the scene graph (SGG) and the SGG with the aid of prior knowledge. We then investigated the main applications of scene graphs and summarized the most commonly used datasets. Finally, we provide some insights into the future development of scene graphs. We believe this will be a very helpful foundation for future research on scene graphs.
研究动机与目标
- 为解决计算机视觉中场景图缺乏全面、系统性综述的问题。
- 分析并分类生成场景图(SGG)的方法,包括借助先验知识增强的方法。
- 研究场景图在视觉关系检测、图像字幕生成、视觉问答(VQA)和图像编辑等任务中的多样化应用。
- 总结用于场景图研究的最广泛使用的基准数据集。
- 为推进场景图技术的未来研究方向提供洞见。
提出的方法
- 本综述对场景图研究进行了系统性文献回顾,重点关注生成技术与应用。
- 将SGG方法分类为仅依赖视觉数据的方法与结合外部知识(如预训练模型或知识库)的方法。
- 分析视觉关系、物体属性和关系推理在场景图构建中的作用。
- 利用标准基准评估各种SGG框架的性能与设计选择。
- 根据注释风格、规模和任务兼容性对数据集进行组织与比较。
- 综合分析各类应用中的趋势与挑战,突出跨模态和推理密集型任务。
实验结果
研究问题
- RQ1在计算机视觉中,场景图的核心组成部分与定义是什么?
- RQ2使用或不使用先验知识的场景图生成方法有何不同?
- RQ3场景图在视觉问答(VQA)和图像编辑等高级视觉任务中的主要应用是什么?
- RQ4哪些数据集最常用于训练和评估场景图模型?
- RQ5场景图研究中的关键挑战与未来研究方向是什么?
主要发现
- 场景图通过显式建模对象、属性及其关系,实现了更高层次的视觉理解。
- 结合先验知识的SGG方法在复杂或罕见关系上的表现优于纯数据驱动方法。
- 视觉问答和图像编辑等应用显著受益于结构化的场景图表示。
- 综述识别出向多模态和推理密集型场景图应用发展的日益增长趋势。
- 多个基准数据集(如VG和COCO-SceneGraph)被广泛用于训练与评估,尽管其规模和注释质量存在差异。
- 缺乏标准化的评估协议以及关系推理的复杂性仍是该领域的主要挑战。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。