[论文解读] Generative AI in Agriculture: Creating Image Datasets Using DALL.E's Advanced Large Language Model Capabilities
本研究证明,基于生成对抗网络(GAN)和大规模语言模型的生成式AI模型DALL·E 2,能够仅通过文本提示生成高保真度的合成农业图像,显著减少对昂贵且耗时的真实世界数据采集的依赖。生成的图像在标准指标(MSE、PSNR、FSIM)上表现优异,使作物-杂草区分、病害检测和精准农业等应用成为可能,而无需进行大量田间成像。
The field of agricultural communication is evolving rapidly with the advent of generative artificial intelligence (AI), particularly image generation technologies. As these tools begin to influence how agricultural data is visualized and disseminated, the sector's diversity spanning both technical and non-technical researchers, demands a rigorous foundational study to demystify the image generation process. This research investigated the role of artificial intelligence (AI), specifically the DALL.E model by OpenAI, in advancing data generation and visualization techniques in agriculture. DALL.E, an advanced AI image generator, works alongside ChatGPT's language processing to transform text descriptions and image clues into realistic visual representations of the content. The study used both approaches of image generation: text-to-image and image-to-image (variation). Six types of datasets depicting fruit crop environment were generated. These AI-generated images were then compared against ground truth images captured by sensors in real agricultural fields. The comparison was based on Peak Signal-to-Noise Ratio (PSNR) and Feature Similarity Index (FSIM) metrics. The image-to-image generation exhibited a 5.78% increase in average PSNR over text-to-image methods, signifying superior image clarity and quality. However, this method also resulted in a 10.23% decrease in average FSIM, indicating a diminished structural and textural similarity to the original images. Similar to these measures, human evaluation also showed that images generated using image-to-image-based method were more realistic compared to those generated with text-to-image approach. The results highlighted DALL.E's potential in generating realistic agricultural image datasets and thus accelerating the development and adoption of imaging-based precision agricultural solutions.
研究动机与目标
- 探索仅使用DALL·E 2生成合成农业图像以减少对真实世界图像采集依赖的可行性。
- 利用标准图像质量指标,评估AI生成的农业图像与真实图像在视觉保真度和准确性方面的表现。
- 展示AI生成图像在作物-杂草区分和病害识别等关键农业应用中的潜力。
- 提出一种可扩展、低成本的替代方案,用于农业AI研究中的传统数据采集,借助文本到图像生成技术。
- 通过数据集扩展、反馈循环和高级评估方法,为未来将生成式AI系统性地整合到精准农业系统中提供建议。
提出的方法
- 采用DALL·E 2,一种基于GAN框架并结合基于Transformer的文本编码器的文本到图像生成模型,从自然语言提示生成合成农业图像。
- 整合GPT-4和chatGPT,以优化和生成适用于多样化农业场景(包括果实、植物以及作物与杂草区分)的精确文本描述。
- 构建了一个经过筛选的AI生成图像数据集,涵盖多个农业类别,包括健康与患病植物、各类作物及田间环境。
- 使用标准指标评估图像质量:均方误差(MSE)、峰值信噪比(PSNR)和特征相似性指数(FSIM)。
- 将AI生成图像与真实图像进行对比,评估其真实感和结构一致性,重点关注视觉保真度和语义准确性。
- 提出一个五步未来采纳路线图:数据集扩展、高级训练、专家反馈整合、使用多样化评估指标(如Inception Score),以及系统集成。

实验结果
研究问题
- RQ1DALL·E 2 仅通过自然语言描述,能否生成高质量、逼真的农业图像?
- RQ2AI生成的农业图像在结构和感知相似性方面,与真实世界图像相比具有怎样的视觉特征?
- RQ3AI生成图像在多大程度上可支持关键农业任务,如作物-杂草区分和病害检测?
- RQ4当前合成图像生成在复现复杂农业场景方面存在哪些局限性,又该如何克服?
- RQ5如何系统性地将DALL·E 2等生成式AI模型整合到现有的农业研究与精准农业工作流中?
主要发现
- DALL·E 2 成功从文本提示生成了逼真的农业图像,显示出输入描述与输出视觉结果之间高度一致。
- AI生成图像在标准图像质量指标上表现优异,报告的MSE、PSNR和FSIM值表明其具有高视觉保真度。
- 合成图像有效捕捉了复杂的农业场景,包括作物-杂草区分和植物健康状况,表明其在AI训练和决策支持中的实用价值。
- 使用AI生成图像显著降低了真实世界图像采集所需的时间、人力和成本,为数据密集型AI应用提供了可扩展的替代方案。
- 通过多种指标(MSE、PSNR、FSIM)的评估证实,生成图像在结构和感知上均接近真实图像,支持其在AI模型训练与测试中的应用。
- 本研究通过数据集扩展、专家反馈整合以及增强评估技术(包括潜在使用Inception Score进行多样性评估)明确了未来采纳的清晰路径。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。