Skip to main content
QUICK REVIEW

[论文解读] The Hebrew Bible as Data: Laboratory - Sharing - Experiences

Dirk Roorda|arXiv (Cornell University)|Jan 8, 2015
Digital Humanities and Scholarship参考文献 3被引用 4
一句话总结

本文介绍了 SHEBANQ,这是一个公开可访问的基于网络的查询系统,用于 ETCBC 的希伯来语圣经文本数据库,基于语言学标注框架(LAF)构建,并由基于 Python 的科学计算技术驱动。该系统使学者能够使用领域特定查询语言(MQL)创建、保存、共享并重现语言学查询,从而在希伯来语圣经的数字人文研究中促进可重现性与跨学科合作。

ABSTRACT

The systematic study of ancient texts including their production, transmission and interpretation is greatly aided by the digital methods that started taking off in the 1970s. But how is that research in turn transmitted to new generations of researchers? We tell a story of Bible and computer across the decades and then point out the current challenges: (1) finding a stable data representation for changing methods of computation; (2) sharing results in inter- and intra-disciplinary ways, for reproducibility and cross-fertilization. We report recent developments in meeting these challenges. The scene is the text database of the Hebrew Bible, constructed by the Eep Talstra Centre for Bible and Computer (ETCBC), which is still growing in detail and sophistication. We show how a subtle mix of computational ingredients enable scholars to research the transmission and interpretation of the Hebrew Bible in new ways: (1) a standard data format, Linguistic Annotation Framework (LAF); (2) the methods of scientific computing, made accessible by (interactive) Python and its associated ecosystem. Additionally, we show how these efforts have culminated in the construction of a new, publicly accessible search engine SHEBANQ, where the text of the Hebrew Bible and its underlying data can be queried in a simple, yet powerful query language MQL, and where those queries can be saved and shared.

研究动机与目标

  • 为解决跨学科领域在希伯来语圣经计算研究中共享与重现的挑战。
  • 开发一个可持续的开放基础设施,用于存储、查询和共享希伯来语圣经的语言学标注。
  • 通过可访问、可共享的查询工作流,弥合神学学者、计算语言学家与学生之间的差距。
  • 利用开放标准与科学计算工具,建立一个现代的数据实验室,用于历史语言学研究。
  • 通过机构归档与版本控制,确保 ETCBC 语言学数据与方法的长期保存与可访问性。

提出的方法

  • 采用语言学标注框架(LAF)作为希伯来语圣经语言学标注的稳定、标准化数据格式。
  • 构建一个名为 SHEBANQ 的网络应用,支持使用领域特定查询语言(MQL)对文本中的语言学模式进行交互式查询。
  • 实施查询保存与共享机制,以促进研究人员之间的可重现性与协作性知识构建。
  • 开发 LAF-Fabric 工具,用于分析和操作 LAF 资源,以支持数据整理与验证。
  • 利用开源软件生态系统(包括 Python 和 GitHub),确保非专业程序员也能访问与扩展。
  • 在 DANS(一个可信的研究数据存储库)中归档 ETCBC 的数据与文档,以实现长期保存与访问。

实验结果

研究问题

  • RQ1如何以稳定、机器可处理的格式表示希伯来语圣经的语言学标注,以支持不断发展的计算方法?
  • RQ2哪些机制能够有效支持不同学术群体之间语言学查询的共享与可重现性?
  • RQ3如何设计一个基于网络的系统,以支持圣经学者与计算语言学家对复杂文本数据进行查询?
  • RQ4开放标准与开源软件在实现可持续、协作性数字人文研究中扮演何种角色?
  • RQ5机构存储库与版本控制系统如何提升数字圣经学术研究的长期可访问性与可信度?

主要发现

  • SHEBANQ 系统成功通过用户友好的、基于查询的 MQL 语言界面,实现了公众对希伯来语圣经语言学数据的访问。
  • 采用 LAF 标准确保了稳定且可互操作的数据模型,能够支持不断演进的计算需求。
  • LAF-Fabric 的集成使得特征频率的自动化分析与标注一致性的验证成为可能。
  • 项目已实现初步采用,表现为非零访问日志,且在神学与计算语言学界均显示出日益增长的兴趣。
  • 通过使用 GitHub 进行代码共享,以及通过 DANS 进行数据归档,实现了软件与数据的透明、版本控制与持久访问。
  • 本项目表明,研究人员无需依赖商业软件供应商,即可成功构建并维持共享的数字研究基础设施。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。