Skip to main content
QUICK REVIEW

[论文解读] The "Collections as ML Data" Checklist for Machine Learning & Cultural Heritage

Benjamin Charles Germain Lee|arXiv (Cornell University)|Jul 6, 2022
Generative Adversarial Networks and Image Synthesis被引用 6
一句话总结

本文提出了「藏品作為機器學習數據」檢查表——一項全面且以實務為導向的框架,旨在責任地將機器學習應用於文化遺產藏品。該框架整合了專案生命週期各階段的社會技術考量,結合機器學習與文化遺產的最佳實踐,並提供具體問題,以確保在GLAM(美術館、圖書館、檔案館、博物館)環境中,機器學習專案具備倫理、透明與永續性。

ABSTRACT

Within the cultural heritage sector, there has been a growing and concerted effort to consider a critical sociotechnical lens when applying machine learning techniques to digital collections. Though the cultural heritage community has collectively developed an emerging body of work detailing responsible operations for machine learning in libraries and other cultural heritage institutions at the organizational level, there remains a paucity of guidelines created specifically for practitioners embarking on machine learning projects. The manifold stakes and sensitivities involved in applying machine learning to cultural heritage underscore the importance of developing such guidelines. This paper contributes to this need by formulating a detailed checklist with guiding questions and practices that can be employed while developing a machine learning project that utilizes cultural heritage data. I call the resulting checklist the "Collections as ML Data" checklist, which, when completed, can be published with the deliverables of the project. By surveying existing projects, including my own project, Newspaper Navigator, I justify the "Collections as ML Data" checklist and demonstrate how the formulated guiding questions can be employed and operationalized.

研究动机与目标

  • 解決文化遺產數據應用機器學習時缺乏實務、專案層級指南的問題。
  • 協助實務工作者應對機器學習在文化遺產領域所帶來的社會技術風險,包括偏見、永續性與隱私問題。
  • 提供一個結構化、可重複使用的檢查表,可與專案成果一同發佈,以提升透明度與責任感。
  • 將批判性數據研究、數位人文学與責任型人工智慧的既有原則,在特定領域情境中具體化。
  • 確保文化遺產領域的機器學習專案具備倫理基礎,可重現且長期可維護。

提出的方法

  • 透過整合現有的機器學習檢查表(如Model Cards、Data Cards)與文化遺產最佳實踐(如Responsible Operations、Critical Cataloging)而開發。
  • 分為四個核心部分:資料與模型、環境影響、組織考量,以及著作權/透明度/維護。
  • 包含105個引導性問題,分屬10個子類別,設計為在專案開發期間完成,並與輸出成果一同發佈。
  • 透過對實際專案(包括作者自身的Newspaper Navigator專案)的分析進行驗證,以證明其實際適用性。
  • 強調利益相關者參與、可審計性、授權與發布後反饋迴圈,以確保專案的長期完整性。
  • 設計用於整合至機構工作流程,著重於文件記錄、可重現性與公開責任感。

实验结果

研究问题

  • RQ1文化遺產機構的實務工作者如何系統性評估在數位化藏品上使用機器學習的倫理與技術影響?
  • RQ2在將文化遺產藏品視為數據時,有哪些具體且可執行的問題可引導責任型機器學習的發展?
  • RQ3如何使GLAM環境中的機器學習專案在初始部署後仍具備透明度、可審計性與永續性?
  • RQ4利益相關者意見、著作權與環境影響在塑造文化遺產領域責任型機器學習實踐中扮演何種角色?
  • RQ5標準化檢查表在多大程度上能提升涉及文化遺產數據的機器學習專案的可重現性與責任感?

主要发现

  • 「藏品作為機器學習數據」檢查表提供了一套全面且經實證測試的框架,可將倫理與技術嚴謹性嵌入文化遺產數據的機器學習專案中。
  • 該檢查表成功地將批判性數據研究與責任型人工智慧的抽象原則,轉化為具體的專案層級問題。
  • 完成檢查表的專案在案例研究(如Newspaper Navigator)中展現出更高的透明度、利益相關者參與度與永續性規劃。
  • 透過要求記錄資料來源、模型限制與授權條款,檢查表支持可重現性與可審計性。
  • 機構採用該檢查表可促進GLAM環境中長期的數據素養與責任型創新。
  • 檢查表的發布後反饋機制(如受眾覆蓋範圍、利益相關者意見)提升了專案的責任感與迭代改進能力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。