[论文解读] ESGReveal: An LLM-based approach for extracting structured data from ESG reports
ESGReveal 是一种基于检索增强的大型语言模型框架,通过三段式架构(ESG元数据、报告预处理和大型语言模型代理模块)从企业报告中提取结构化 ESG 数据。该框架在使用 GPT-4 时,数据提取准确率达到 76.9%,披露分析准确率达到 83.7%,显著优于基线模型,并揭示了 166 家香港上市公司中环境类披露率低至 69.5%,社会类披露率低至 57.2% 的现状。
ESGReveal is an innovative method proposed for efficiently extracting and analyzing Environmental, Social, and Governance (ESG) data from corporate reports, catering to the critical need for reliable ESG information retrieval. This approach utilizes Large Language Models (LLM) enhanced with Retrieval Augmented Generation (RAG) techniques. The ESGReveal system includes an ESG metadata module for targeted queries, a preprocessing module for assembling databases, and an LLM agent for data extraction. Its efficacy was appraised using ESG reports from 166 companies across various sectors listed on the Hong Kong Stock Exchange in 2022, ensuring comprehensive industry and market capitalization representation. Utilizing ESGReveal unearthed significant insights into ESG reporting with GPT-4, demonstrating an accuracy of 76.9% in data extraction and 83.7% in disclosure analysis, which is an improvement over baseline models. This highlights the framework's capacity to refine ESG data analysis precision. Moreover, it revealed a demand for reinforced ESG disclosures, with environmental and social data disclosures standing at 69.5% and 57.2%, respectively, suggesting a pursuit for more corporate transparency. While current iterations of ESGReveal do not process pictorial information, a functionality intended for future enhancement, the study calls for continued research to further develop and compare the analytical capabilities of various LLMs. In summary, ESGReveal is a stride forward in ESG data processing, offering stakeholders a sophisticated tool to better evaluate and advance corporate sustainability efforts. Its evolution is promising in promoting transparency in corporate reporting and aligning with broader sustainable development aims.
研究动机与目标
- 解决缺乏公开、详尽的 ESG 披露数据库的问题,以整合跨公司和跨行业的指标。
- 通过先进的大型语言模型和检索增强生成(RAG)技术,提升从非结构化报告中提取 ESG 数据的准确性和一致性。
- 评估并基准化不同大型语言模型在 ESG 专项数据提取与分析任务中的表现。
- 通过分析香港证券交易所 12 个行业中的披露情况,识别 ESG 报告透明度的差距。
- 建立一种可扩展、可适应的系统化 ESG 数据提取框架,以支持利益相关者决策和监管监督。
提出的方法
- ESGReveal 框架将大型语言模型(LLMs)与检索增强生成(RAG)相结合,以增强从 ESG 报告中提取上下文相关数据的能力。
- 其采用 ESG 元数据模块,利用大型语言模型驱动的语义理解,将多样化的 ESG 报告标准(如 GRI、SASB)映射为结构化提取指令。
- 报告预处理模块将原始 ESG 报告解析并结构化为机器可读格式,以支持高效检索与索引。
- LLM 代理模块利用基于领域特定知识的 RAG 增强提示,执行上下文感知的数值与文本 ESG 指标提取。
- 系统使用 GPT-4 及其他 LLM 执行数据提取与披露分析,性能通过与人工标注基准对比进行评估。
- 该框架通过可扩展的元数据模式和从参考数据库动态检索相关 ESG 标准,支持多标准适应性。

实验结果
研究问题
- RQ1基于 LLM 的 RAG 框架在从非结构化企业报告中提取结构化 ESG 数据方面效果如何?
- RQ2不同 LLM(如 GPT-3.5、GPT-4)在 ESG 数据提取与披露分析任务中的表现有何差异?
- RQ3不同行业中公司对环境、社会和治理指标的披露程度如何?是否存在可识别的模式?
- RQ4ESG 元数据模块如何提升 LLM 在多标准 ESG 报告环境下的性能?
- RQ5当前基于 LLM 的 ESG 提取系统在非文本数据(如图表和图像)处理方面存在哪些关键局限?
主要发现
- GPT-4 在结构化数据提取中达到 76.9% 的准确率,在披露分析中达到 83.7%,优于基线模型超过 20 个百分点。
- ESG 元数据模块使 GPT-4 在 ESG 分析任务中的性能提升 2.5 个百分点,使 GPT-3.5 提升 9.9 个百分点。
- 环境类数据披露的平均比例为 69.5%,社会类数据披露的平均比例为 57.2%,表明企业在透明度方面仍有显著提升空间。
- 各行业整体 ESG 披露率未超过 80%,凸显了报告完整性方面存在系统性差距。
- 在各行业中识别出常见 ESG 行动,如“减少排放”和“推动数字健康”,以及行业特有行动,如金融行业的“绿色金融”和工业行业的“可持续供应链”。
- 该框架在 12 个行业和 166 家香港联合交易所上市公司中表现出强大的适应能力,支持跨行业基准比较。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。