Skip to main content
QUICK REVIEW

[论文解读] Information Carriers and Identification of Information Objects: An Ontological Approach

Martin Doerr, Yannis Tzitzikas|arXiv (Cornell University)|Jan 1, 2012
Digital and Traditional Archives Management参考文献 6被引用 6
一句话总结

本文提出了一种本体论框架,用于独立于其物理载体或编码格式,唯一识别信息对象。通过将信息对象建模为受正式标准(例如期刊格式指南)约束的离散符号排列,该方法能够在迁移过程中实现可靠的信息识别、保存和真实性验证,为数字保存决策提供正式基础。

ABSTRACT

Even though library and archival practice, as well as Digital Preservation, have a long tradition in identifying information objects, the question of their precise identity under change of carrier or migration is still a riddle to science. The objective of this paper is to provide criteria for the unique identification of some important kinds of information objects, independent from the kind of carrier or specific encoding. Our approach is based on the idea that the substance of some kinds of information objects can completely be described in terms of discrete arrangements of finite numbers of known kinds of symbols, such as those implied by style guides for scientific journal submissions. Our theory is also useful for selecting or describing what has to be preserved. This is a fundamental problem since curators and archivists would like to formally record the decisions of what has to be preserved over time and to decide (or verify) whether a migration (transformation) preserves the intended information content. Furthermore, it is important for reasoning about the authenticity of digital objects, as well as for reducing the cost of digital preservation.

研究动机与目标

  • 解决信息对象在迁移至新载体或格式时持续存在的识别难题。
  • 建立独立于其物理或数字载体的信息对象唯一识别标准。
  • 通过正式指定必须保存的内容以维持预期信息内容,支持数字保存。
  • 在转换或迁移过程中,支持对数字对象的真实性与语义保真度进行推理。
  • 通过聚焦信息对象中本质且语义稳定的组件,降低数字保存成本。

提出的方法

  • 将信息对象定义为有限符号类型的离散排列,使用诸如科学期刊格式指南等正式标准进行建模。
  • 开发一种本体论模型,区分信息对象与其载体,将载体视为物理或数字媒体。
  • 使用符号组合规则以独立于编码或表示方式的方式描述信息对象的本质。
  • 应用正式标准,通过比较迁移前后符号排列的异同,评估迁移是否保留了核心信息内容。
  • 将该模型集成到数字图书馆和档案系统中,以支持持久识别和保存决策。
  • 利用本体论验证变换过程不会改变信息对象的预期语义内容。

实验结果

研究问题

  • RQ1如何在不考虑载体或编码格式的情况下唯一识别信息对象?
  • RQ2何种正式标准可确保迁移过程保留对象的核心信息内容?
  • RQ3如何验证不同表示或载体下的数字对象的真实性?
  • RQ4为维持信息对象的身份与意义随时间不变,必须保存其哪些组成部分?
  • RQ5本体论模型在何种方式上可减少数字保存工作流中的模糊性与成本?

主要发现

  • 信息对象可通过其离散符号排列被唯一识别,且独立于载体或编码格式。
  • 所提出的模型能够通过比较符号级结构,正式验证迁移是否保留了预期信息内容。
  • 该框架通过隔离信息对象中本质且语义稳定的组成部分,支持一致的保存决策。
  • 该方法通过提供正式基础以识别必须保存的内容,减少了数字管理中的模糊性。
  • 该模型通过确保变换不改变信息对象的核心符号结构,增强了真实性评估。
  • 该理论适用于数字图书馆和档案系统,可实现长期保存,并降低语义漂移的风险。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。