Skip to main content
QUICK REVIEW

[论文解读] An Efficient Data Warehouse for Crop Yield Prediction

Vuong M. Ngo, Nhien‐An Le‐Khac|arXiv (Cornell University)|Jun 26, 2018
Data Mining Algorithms and Applications参考文献 16被引用 11
一句话总结

本文提出了一种高效的数仓架构,采用星象模式整合异构的、大规模的农业大数据,用于精准农业中的作物产量预测。通过在各利益相关方之间实现实时、安全且隐私保护的数据访问,该系统通过可扩展的数据集成与分析,增强了决策支持能力,显著提升了处理时空作物数据以支持预测建模的效率。

ABSTRACT

Nowadays, precision agriculture combined with modern information and communications technologies, is becoming more common in agricultural activities such as automated irrigation systems, precision planting, variable rate applications of nutrients and pesticides, and agricultural decision support systems. In the latter, crop management data analysis, based on machine learning and data mining, focuses mainly on how to efficiently forecast and improve crop yield. In recent years, raw and semi-processed agricultural data are usually collected using sensors, robots, satellites, weather stations, farm equipment, farmers and agribusinesses while the Internet of Things (IoT) should deliver the promise of wirelessly connecting objects and devices in the agricultural ecosystem. Agricultural data typically captures information about farming entities and operations. Every farming entity encapsulates an individual farming concept, such as field, crop, seed, soil, temperature, humidity, pest, and weed. Agricultural datasets are spatial, temporal, complex, heterogeneous, non-standardized, and very large. In particular, agricultural data is considered as Big Data in terms of volume, variety, velocity and veracity. Designing and developing a data warehouse for precision agriculture is a key foundation for establishing a crop intelligence platform, which will enable resource efficient agronomy decision making and recommendations. Some of the requirements for such an agricultural data warehouse are privacy, security, and real-time access among its stakeholders (e.g., farmers, farm equipment manufacturers, agribusinesses, co-operative societies, customers and possibly Government agencies). However, currently there are very few reports in the literature that focus on the design of efficient data warehouses with the view of enabling Agricultural Big Data analysis and data mining. In this paper ...

研究动机与目标

  • 解决精准农业中针对农业大数据缺乏高效数仓设计的问题。
  • 支持农民、农业企业及政府机构等多样化利益相关方实现实时、安全且隐私保护的数据访问。
  • 实现来自物联网传感器、卫星和农用设备等异构数据源的可扩展数据集成,以提升作物产量预测能力。
  • 设计一种支持机器学习与数据挖掘工作负载的数仓,以支持农学决策。

提出的方法

  • 采用星象模式设计数仓,以建模复杂的农业实体,如农田、作物、土壤和气象条件。
  • 整合多源异构的农业数据,包括物联网传感器数据流、卫星图像和农场管理记录。
  • 实施数据建模技术,以应对农业大数据在体量、多样性、速度和真实性方面的挑战。
  • 通过适用于多利益相关方环境的访问控制机制,确保数据安全与隐私。
  • 优化数仓模式以支持OLAP操作和高效查询,从而支持预测分析。
  • 实现端到端的实时数据摄取与处理流水线,以支持动态决策支持系统。

实验结果

研究问题

  • RQ1如何高效地整合多样化的、大规模的农业数据源,以支持作物产量预测?
  • RQ2何种模式设计能够有效建模复杂农业实体及其时空关系?
  • RQ3在多利益相关方的农业数仓中,如何确保隐私与实时访问?
  • RQ4哪些架构选择能够提升精准农业数仓的可扩展性与性能?
  • RQ5所提出的数仓在多大程度上提升了作物产量预测模型的准确性与效率?

主要发现

  • 所提出的星象模式能有效建模复杂农业实体及其相互关系,显著提升数据集成效率与查询性能。
  • 数仓支持来自物联网与卫星源的实时数据摄取,实现实时分析以支持决策。
  • 系统确保了数据隐私与访问控制,适用于在多样化农业利益相关方中的部署。
  • 该架构在处理高吞吐量、异构农业数据方面表现出良好的可扩展性,显著缩短了预测建模前的数据准备时间。
  • 多源海量数据的整合显著提升了作物产量预测模型的准确性与可靠性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。