Skip to main content
QUICK REVIEW

[论文解读] Accessible Visualization via Natural Language Descriptions: A Four-Level Model of Semantic Content

Alan Lundgard, Arvind Satyanarayan|arXiv (Cornell University)|Oct 8, 2021
Subtitles and Audiovisual Media被引用 11
一句话总结

本文提出了一种四层次模型,用于分类可视化中自然语言描述的语义内容,涵盖从低层次的可视化构建到高层次的领域洞察。通过对2,147个句子及120名参与者的混合方法研究发现,盲人和视力正常读者在语义层次的优先级上存在差异,表明有效的可访问性需要根据读者的具体需求和偏好定制描述。

ABSTRACT

Natural language descriptions sometimes accompany visualizations to better communicate and contextualize their insights, and to improve their accessibility for readers with disabilities. However, it is difficult to evaluate the usefulness of these descriptions, and how effectively they improve access to meaningful information, because we have little understanding of the semantic content they convey, and how different readers receive this content. In response, we introduce a conceptual model for the semantic content conveyed by natural language descriptions of visualizations. Developed through a grounded theory analysis of 2,147 sentences, our model spans four levels of semantic content: enumerating visualization construction properties (e.g., marks and encodings); reporting statistical concepts and relations (e.g., extrema and correlations); identifying perceptual and cognitive phenomena (e.g., complex trends and patterns); and elucidating domain-specific insights (e.g., social and political context). To demonstrate how our model can be applied to evaluate the effectiveness of visualization descriptions, we conduct a mixed-methods evaluation with 30 blind and 90 sighted readers, and find that these reader groups differ significantly on which semantic content they rank as most useful. Together, our model and findings suggest that access to meaningful information is strongly reader-specific, and that research in automatic visualization captioning should orient toward descriptions that more richly communicate overall trends and statistics, sensitive to reader preferences. Our work further opens a space of research on natural language as a data interface coequal with visualization.

研究动机与目标

  • 为理解可视化中自然语言描述所传达的语义内容缺乏系统性框架的问题提供解决方案。
  • 识别并分类描述所传达的信息类型,特别是针对视觉障碍读者。
  • 评估不同读者群体(盲人与视力正常者)对可视化描述中不同语义内容层次的有用性感知。
  • 为设计更高效、具备读者意识的自动字幕系统提供依据,以实现可访问的数据可视化。
  • 在可访问性理论和实证验证的基础上,将自然语言定位为与可视化并列的数据接口。

提出的方法

  • 基于包含120多名参与者的在线研究中收集的2,147条自然语言句子,开展基础理论分析,识别可视化描述中反复出现的语义模式。
  • 构建了一个四层次概念模型:(1) 可视化构建属性(例如,图形标记、编码方式),(2) 统计概念与关系(例如,极值、相关性),(3) 可感知与认知现象(例如,趋势、模式),(4) 领域特定洞察(例如,社会、政治背景)。
  • 将该模型应用于评估跨多个领域的120个可视化描述,使用语义编码将内容分类至四个层次。
  • 通过混合方法评估,邀请30名盲人和90名视力正常者参与,以评估不同语义层次描述的感知有用性。
  • 采用统计分析(例如,非参数检验)比较不同读者群体的偏好,识别语义内容有用性排名中的显著差异。
  • 将研究发现整合进一个框架,用于设计具备可访问性、读者敏感性的自然语言可视化接口,强调上下文感知的字幕生成。
Figure 1: A visual “fingerprint” [ 49 ] of our corpus, faceted by chart type and difficulty. Each row corresponds to a single chart. Each column shows a participant-authored description for that chart, color coded according to our model. The first column shows the provided Level 1 prompt.
Figure 1: A visual “fingerprint” [ 49 ] of our corpus, faceted by chart type and difficulty. Each row corresponds to a single chart. Each column shows a participant-authored description for that chart, color coded according to our model. The first column shows the provided Level 1 prompt.

实验结果

研究问题

  • RQ1自然语言描述可视化所传达的语义内容具有哪些明确的层次?
  • RQ2盲人和视力正常读者在对可视化描述中不同语义内容层次的有用性感知上存在哪些差异?
  • RQ3不同语义层次(例如,统计性与情境性)在实现对可视化中意义信息的有效访问方面,其贡献程度如何?
  • RQ4语义内容的概念模型在改进可访问性可视化描述的设计与评估方面有何作用?
  • RQ5读者对语义内容的特定偏好对可视化可访问性中自动字幕系统的设计有何启示?

主要发现

  • 四层次语义模型能够成功对可视化描述中的语义范围进行分类,涵盖从技术构建到高层次领域洞察的全部语义层次。
  • 盲人读者将领域特定洞察(第4级)视为最有用,其次为统计关系(第2级);而视力正常读者则优先考虑感知与认知现象(第3级)和统计关系(第2级)。
  • 在有用性排名上,盲人与视力正常读者之间存在显著差异(p < 0.01),表明可访问性具有高度的读者特异性。
  • 强调整体趋势和统计关系的描述被认为比仅关注视觉构建或低层次编码细节的描述更具用性。
  • 本研究证实,当前的自动字幕系统往往无法捕捉对盲人用户实现有意义访问至关重要的高层语义内容。
  • 研究结果表明,未来的自动字幕系统应考虑读者偏好,优先生成丰富且具备上下文意识的描述,而非机械或结构化的描述。
Accessible Visualization via Natural Language Descriptions: A Four-Level Model of Semantic Content

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。