Skip to main content
QUICK REVIEW

[论文解读] Reconstructing a website's lost past: Methodological issues concerning the history of www.unibo.it

Federico Nanni|arXiv (Cornell University)|Apr 20, 2016
Digital Humanities and Scholarship被引用 3
一句话总结

本文提出了一套方法论框架,用于重建博洛尼亚大学官网(www.unibo.it)的数字历史,该网站在互联网档案馆的时光机(Wayback Machine)中被排除长达13年。通过利用替代性网络存档技术并分析原始数字资料,本研究展示了即使在官方档案存在空白的情况下,仍可恢复和分析机构网络历史,为学术机构的数字遗产研究提供了可复用的模型。

ABSTRACT

This paper describes how born digital primary sources could be used to reconstruct the recent history of scientific institutions. The case study is an analysis of the first 25 years online of the University of Bologna. The focus of this work is primarily methodological: several different issues are presented, starting with the fact that the University of Bologna website has been excluded for thirteen years from the Internet Archive's Wayback Machine, and possible solutions are proposed and applied. The article is organised in three parts: in the first one, some of the fundamental concepts on web archives and the preservation of born digital sources are introduced. Then the reconstruction of the University of Bologna web's past is presented. Finally the future of this research is described, presenting a specific case study in which the historian's craft will be challenged by a completely different issue, namely the large amount of data available in the university digital library.

研究动机与目标

  • 解决当官方网络档案不完整或缺失时,重建学术机构数字历史的挑战。
  • 开发并应用方法论方法,从替代性数字来源恢复丢失的网站内容。
  • 证明利用原始数字一手资料进行科学机构历史研究的可行性。
  • 探讨在未来的数字档案研究中,数据体量与复杂性带来的影响。

提出的方法

  • 利用替代性网络存档工具与技术,从博洛尼亚大学网站恢复内容,绕过互联网档案馆时光机中缺失的问题。
  • 应用数字图书馆数据与元数据,重建机构网络存在的时间线与演变过程。
  • 结合网络爬取、元数据提取与档案分析,跨多个时间点重建网站内容。
  • 采用混合方法,整合数字图书馆资源与网络存档重建,以验证历史数据。
  • 以www.unibo.it为案例,作为测试平台,完善机构网络历史研究的方法论协议。
  • 建立一个框架,使未来的历史学家能够系统性地恢复并分析机构网络档案,即使在官方收藏存在空白的情况下。

实验结果

研究问题

  • RQ1当学术机构的官方网络存在因主要网络存档缺失而无法获取时,如何重建其数字历史?
  • RQ2可采用哪些方法论策略,从替代性数字来源恢复丢失的网站内容?
  • RQ3网络存档的空白如何影响机构数字记录的历史准确性与完整性?
  • RQ4原始数字一手资料在重建机构网络历史中扮演何种角色?
  • RQ5在面对拥有海量数据的大型机构数字图书馆时,扩展数字档案研究会面临哪些挑战?

主要发现

  • 尽管www.unibo.it在互联网档案馆时光机中被排除13年,本研究仍成功重建了该校网站超过25年的历史。
  • 替代性网络存档技术实现了对网站历史内容与结构演变的局部但可靠的恢复。
  • 本研究证明,当与元数据及数字图书馆资源结合时,原始数字资料可作为缺失网络存档的可行替代品。
  • 所开发的方法论框架可适用于其他面临类似档案空白的学术机构。
  • 本研究揭示,机构数字图书馆中的数据体量与复杂性,正成为未来数字历史研究中一个新且重大的挑战。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。