[论文解读] Exploring the Carbon Footprint of Hugging Face's ML Models: A Repository Mining Study
本研究通过仓库挖掘分析了 Hugging Face 上 1,417 个机器学习模型的碳足迹,揭示了碳排放报告率停滞不前,且排放量与模型大小、数据集大小及自然语言处理(NLP)重点存在强烈相关性。研究提出了两种分类体系——碳排放报告实践与碳效率——以促进机器学习社区的透明度与可持续模型开发。
The rise of machine learning (ML) systems has exacerbated their carbon footprint due to increased capabilities and model sizes. However, there is scarce knowledge on how the carbon footprint of ML models is actually measured, reported, and evaluated. In light of this, the paper aims to analyze the measurement of the carbon footprint of 1,417 ML models and associated datasets on Hugging Face, which is the most popular repository for pretrained ML models. The goal is to provide insights and recommendations on how to report and optimize the carbon efficiency of ML models. The study includes the first repository mining study on the Hugging Face Hub API on carbon emissions. This study seeks to answer two research questions: (1) how do ML model creators measure and report carbon emissions on Hugging Face Hub?, and (2) what aspects impact the carbon emissions of training ML models? The study yielded several key findings. These include a stalled proportion of carbon emissions-reporting models, a slight decrease in reported carbon footprint on Hugging Face over the past 2 years, and a continued dominance of NLP as the main application domain. Furthermore, the study uncovers correlations between carbon emissions and various attributes such as model size, dataset size, and ML application domains. These results highlight the need for software measurements to improve energy reporting practices and promote carbon-efficient model development within the Hugging Face community. In response to this issue, two classifications are proposed: one for categorizing models based on their carbon emission reporting practices and another for their carbon efficiency. The aim of these classification proposals is to foster transparency and sustainable model development within the ML community.
研究动机与目标
- 调查 Hugging Face 上机器学习模型开发者如何测量并报告其模型的碳排放,以填补可持续性透明度方面的关键空白。
- 识别影响机器学习模型碳足迹的关键因素,特别是模型大小、数据集大小和应用领域。
- 通过提出可操作的分类体系,解决机器学习社区中缺乏标准化碳排放报告实践的问题。
- 基于实证洞察与对从业者和研究人员的建议,推动碳效率更高的模型开发。
提出的方法
- 利用 Hugging Face Hub API 开展大规模仓库挖掘研究,提取 1,417 个模型的元数据和自报碳排放数据。
- 应用数据预处理与标准化技术,统一不同模型之间不一致的自报碳排放值。
- 进行统计分析,识别碳排放与模型属性(如模型大小、数据集大小及性能指标)之间的相关性。
- 开发并提出两种分类体系:一种用于根据碳排放报告实践对模型进行分类,另一种用于评估碳效率。
- 在 Zenodo 上提供全面的可复现包,包含代码、数据集和 Jupyter 笔记本,以支持可复现性与未来扩展。
实验结果
研究问题
- RQ1Hugging Face 上的机器学习模型开发者如何测量并报告其模型的碳排放?
- RQ2哪些模型和数据集属性与 Hugging Face 模型的碳排放具有最强相关性?
- RQ3过去两年中,Hugging Face 上报告碳排放的模型比例有何变化?
- RQ4在自然语言处理(NLP)、计算机视觉或音频处理等应用领域中,碳足迹特征有何差异?
- RQ5在 Hugging Face 机器学习生态系统中,哪些系统性障碍阻碍了碳排放报告与效率提升?
主要发现
- 过去两年中,Hugging Face 上报告碳排放的模型比例保持停滞,表明可持续性透明度方面进展甚微。
- 模型大小与碳排放之间存在显著相关性,模型越大,碳排放越高。
- 数据集大小也与碳排放呈现强烈正相关,凸显数据密集型训练是主要排放来源。
- 自然语言处理(NLP)仍是碳排放报告的主导应用领域,占自报碳足迹模型的大多数。
- 尽管意识日益增强,但未观察到模型性能与排放之间存在明确的权衡,表明其相互依赖关系复杂。
- 本研究发现缺乏标准化的报告实践,许多模型缺少信息丰富的模型卡片,限制了碳排放评估的可靠性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。