Skip to main content
QUICK REVIEW

[Paper Review] Before Smelling the Video: A Two-Stage Pipeline for Interpretable Video-to-Scent Plans

Kaicheng Wang, Keyong Shao|arXiv (Cornell University)|Jan 27, 2026
Olfactory and Sensory Function Studies0 citations
TL;DR

The paper presents a two-stage video-to-scent planning pipeline that uses a vision–language model to extract visual semantics and a large language model to generate structured scent plans, then reports user studies where system-generated plans are preferred over baselines.

ABSTRACT

Olfactory cues can enhance immersion in interactive media, yet smell remains rare because it is difficult to author and synchronize with dynamic video. Prior olfactory interfaces rely on designer triggers and fixed event-to-odor mappings that do not scale to unconstrained content. This work examines whether semantic planning for smell is intelligible to people before physical scent delivery. We present a video-to-scent planning pipeline that separates visual semantic extraction using a vision-language model from semantic-to-olfactory inference using a large language model. Two survey studies compare system-generated scent plans with over-inclusive and naive baselines. Results show consistent preference for plans that prioritize perceptually salient cues and align scent changes with visible actions, supporting semantic planning as a foundation for future olfactory media systems.

Motivation & Objective

  • Motivate olfactory augmentation for videos by separating semantic extraction from olfactory inference.
  • Investigate whether semantic scent planning can be intelligible to people before physical scent delivery.
  • Evaluate the alignment of system-generated scent plans with human expectations for relevance and temporal coherence.

Proposed method

  • Stage 1 uses a vision–language model (Gemini 3 Pro) to extract time-aligned visual semantics from sampled video frames.
  • Stage 2 uses a large language model (GPT-5.2) to transform the visual timeline into a structured scent plan under a fixed odor schema.
  • The output is a temporally organized scent plan intended for future olfactory interfaces, not physical scent generation.
  • Two online survey studies compare system-generated scent plans with an over-inclusive baseline and a naive baseline.
  • Participants evaluated perceived olfactory relevance, temporal coherence, immersion, and coherence with video progression.
Figure 1. We introduce a two-stage video-to-scent planning pipeline that translates visual events in video into structured, human-interpretable scent plans, without generating physical scents. (A) A vision–language model (Gemini 3 Pro) processes uniformly sampled video frames to extract time-aligned
Figure 1. We introduce a two-stage video-to-scent planning pipeline that translates visual events in video into structured, human-interpretable scent plans, without generating physical scents. (A) A vision–language model (Gemini 3 Pro) processes uniformly sampled video frames to extract time-aligned

Experimental results

Research questions

  • RQ1RQ1: How well can a computational system generate scent plans that users perceive as temporally coherent and aligned with dynamic video content?
  • RQ2RQ2: Are system-generated scent plans perceived as plausible and non-disruptive when imagined as part of a video viewing experience?

Key findings

  • Study 1 found the system-generated plans had the lowest mean rank (1.586) compared with the over-inclusive (1.871) and naive (2.543) baselines.
  • System plans were ranked first in 54.3% of trials, outperforming both baselines.
  • A Friedman test showed a significant difference in aggregated ranks across conditions (χ²=19.36, p=6.26×10⁻⁵).
  • Pairwise tests showed System > Over and System > Naive, while Over > Naive.
  • Qualitative responses indicated participants favored focusing on a dominant olfactory source rather than exhaustive coverage of all visible elements, and emphasized timing of scent changes with moments of action.
  • Study 2 reported that participants preferred system-generated plans over the over-inclusive baseline for immersion, coherence, and lower distraction, with timing and evolution described as appropriate; concerns centered on descriptive choices rather than the concept of olfactory augmentation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.