Skip to main content
QUICK REVIEW

[Paper Review] Harnessing Large Vision and Language Models in Agriculture: A Review

Hongyan Zhu, Shuai Qin|arXiv (Cornell University)|Jul 29, 2024
Smart Agriculture and AI5 citations
TL;DR

This review surveys how large vision, language, and vision-language models (LVLM/MLLM) can address agricultural tasks—from pest/disease detection to soil and seed quality, and farmer decision support.

ABSTRACT

Large models can play important roles in many domains. Agriculture is another key factor affecting the lives of people around the world. It provides food, fabric, and coal for humanity. However, facing many challenges such as pests and diseases, soil degradation, global warming, and food security, how to steadily increase the yield in the agricultural sector is a problem that humans still need to solve. Large models can help farmers improve production efficiency and harvest by detecting a series of agricultural production tasks such as pests and diseases, soil quality, and seed quality. It can also help farmers make wise decisions through a variety of information, such as images, text, etc. Herein, we delve into the potential applications of large models in agriculture, from large language model (LLM) and large vision model (LVM) to large vision-language models (LVLM). After gaining a deeper understanding of multimodal large language models (MLLM), it can be recognized that problems such as agricultural image processing, agricultural question answering systems, and agricultural machine automation can all be solved by large models. Large models have great potential in the field of agriculture. We outline the current applications of agricultural large models, and aims to emphasize the importance of large models in the domain of agriculture. In the end, we envisage a future in which famers use MLLM to accomplish many tasks in agriculture, which can greatly improve agricultural production efficiency and yield.

Motivation & Objective

  • Explore how LVLM/MLLM can transform agricultural tasks and decision making.
  • Categorize current and potential applications across pest/disease detection, soil and seed quality, and automation.
  • Identify challenges, limitations, and data requirements for deploying large models in agriculture.
  • Highlight future directions and farmer-facing benefits of multimodal agricultural AI.

Proposed method

  • Survey existing literature on large language models (LLMs), large vision models (LVMs), and LVLMs in agriculture.
  • Analyze areas such as agricultural image processing, agricultural question answering systems, and automation.
  • Discuss multimodal data fusion, model architectures, and information requirements for agricultural tasks.
  • Outline practical considerations including data, reliability, and deployment in farming contexts.

Experimental results

Research questions

  • RQ1What agricultural tasks can be addressed by LVLMs and MLLMs?
  • RQ2What are the main challenges and limitations for deploying large models in agriculture?
  • RQ3How can multimodal models improve farmer decision-making and farm management?
  • RQ4What future directions and requirements are needed to realize widespread adoption in agriculture?

Key findings

  • Large models have potential to improve production efficiency and yield in agriculture.
  • LVLMs/MLLMs can support agricultural image processing, QA systems, and automated workflows.
  • Multimodal information (images, text, etc.) can enhance farmer decision-making and task automation.
  • Challenges include domain adaptation, data availability, reliability, and deployment in real-world farming conditions.
  • Future work envisions farmer-facing MLLMs enabling a broader range of agricultural tasks.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.