[论文解读] There Is a Digital Art History
本文主张,大规模视觉模型——尤其是基于Transformer的基座模型——已促成向乔安娜·德鲁克所定义的‘数字艺术史’范式转变。通过分析超越再现的视觉逻辑(如非透视性、抽象性),并借助基于CLIP的模型展开两项案例研究,作者表明,数字艺术史如今必须批判性地审视模型与其训练数据之间的认识论纠缠,从而将该领域重新定义为内在跨学科且方法论上具有自我反思性的研究领域。
In this paper, we revisit Johanna Drucker's question, "Is there a digital art history?" -- posed exactly a decade ago -- in the light of the emergence of large-scale, transformer-based vision models. While more traditional types of neural networks have long been part of digital art history, and digital humanities projects have recently begun to use transformer models, their epistemic implications and methodological affordances have not yet been systematically analyzed. We focus our analysis on two main aspects that, together, seem to suggest a coming paradigm shift towards a "digital" art history in Drucker's sense. On the one hand, the visual-cultural repertoire newly encoded in large-scale vision models has an outsized effect on digital art history. The inclusion of significant numbers of non-photographic images allows for the extraction and automation of different forms of visual logics. Large-scale vision models have "seen" large parts of the Western visual canon mediated by Net visual culture, and they continuously solidify and concretize this canon through their already widespread application in all aspects of digital life. On the other hand, based on two technical case studies of utilizing a contemporary large-scale visual model to investigate basic questions from the fields of art history and urbanism, we suggest that such systems require a new critical methodology that takes into account the epistemic entanglement of a model and its applications. This new methodology reads its corpora through a neural model's training data, and vice versa: the visual ideologies of research datasets and training datasets become entangled.
研究动机与目标
- 在视觉基座模型近年进展的背景下,重新审视乔安娜·德鲁克2013年提出的‘是否存在一种‘数字’艺术史?’这一问题。
- 探究大规模视觉模型对艺术史研究的认识论影响。
- 证明数字艺术史如今必须包含对模型训练数据与视觉意识形态的批判性分析。
- 提出一种新方法论,将研究语料库与模型训练数据视为相互纠缠的意义来源。
- 倡导在基座模型时代,艺术史、媒体研究与计算人文学科之间开展跨学科合作。
提出的方法
- 将基于CLIP的视觉模型用作分析工具,从大规模图像语料库中提取视觉逻辑。
- 应用零样本提示(zero-shot prompting)与对比学习,分析非形象性与非再现性视觉形式。
- 开展两项技术性案例研究:一项使用2D-CLIP研究城市空间感知,另一项使用CLIP-MAP研究艺术风格演变。
- 采用批判性方法论,将模型训练数据视为文化档案,分析其对视觉推理的认知印记。
- 开展逆向解读:利用模型输出,质询其预训练数据中编码的视觉文化。
- 公开提供CLIP-MAP与2D-CLIP工具,以确保可复现性与方法论透明度。
实验结果
研究问题
- RQ1大规模视觉模型在多大程度上编码并再现了西方视觉传统中的非再现性与非透视性视觉逻辑?
- RQ2基座模型的训练数据在多大程度上塑造了数字艺术史的认识论框架?
- RQ3数字艺术史如何超越对象识别,通过神经网络模型整合细读与远观?
- RQ4为批判性分析模型行为与训练数据之间的纠缠关系,需要哪些方法论转变?
- RQ5艺术史在何种方式上可利用视觉模型的多义性、基于向量的推理,作为视觉分析的新形式?
主要发现
- 大规模视觉模型编码了广泛的视觉文化语汇,包括非形象性与非透视性逻辑,从而为艺术史中的新型视觉分析提供了可能。
- 这些模型在数字生活中的广泛部署,持续强化并具体化了西方视觉传统作为主导文化框架的作用。
- 使用CLIP-MAP与2D-CLIP的案例研究表明,这些模型能够大规模揭示艺术与城市形态中的空间与风格模式。
- 视觉模型的训练数据并非中立;它们承载着视觉意识形态,影响模型输出,从而影响艺术史的解读。
- 需要一种新批判性方法论,将研究语料库与模型训练数据视为共同构成意义的来源。
- 数字艺术史领域如今必须将批判性机器视觉作为核心组成部分,整合媒体研究与计算人文学科的洞见。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。