Skip to main content
QUICK REVIEW

[Paper Review] Manipulating Highly Deformable Materials Using a Visual Feedback Dictionary

Biao Jia, Zhe Hu|arXiv (Cornell University)|Oct 18, 2017
Advanced Vision and Imaging35 references3 citations
TL;DR

This paper proposes a visual feedback dictionary method for manipulating highly deformable materials like cloth using a novel histogram of oriented wrinkles (HOW) feature representation derived from RGB video. By combining HOW-features with sparse linear representation over a precomputed visual feedback dictionary, the system enables real-time, robust control for complex tasks—including human-robot collaboration—achieving 94.53% success rate in folding tasks using HOG+HOW features.

ABSTRACT

The complex physical properties of highly deformable materials such as clothes pose significant challenges fanipulation systems. We present a novel visual feedback dictionary-based method for manipulating defoor autonomous robotic mrmable objects towards a desired configuration. Our approach is based on visual servoing and we use an efficient technique to extract key features from the RGB sensor stream in the form of a histogram of deformable model features. These histogram features serve as high-level representations of the state of the deformable material. Next, we collect manipulation data and use a visual feedback dictionary that maps the velocity in the high-dimensional feature space to the velocity of the robotic end-effectors for manipulation. We have evaluated our approach on a set of complex manipulation tasks and human-robot manipulation tasks on different cloth pieces with varying material characteristics.

Motivation & Objective

  • To address the challenge of manipulating highly deformable materials such as cloth, which have high-dimensional configuration spaces and are sensitive to perturbations.
  • To develop a robust, real-time visual servoing approach that does not rely on extensive training data or fiducial markers.
  • To enable human-robot collaborative manipulation of deformable objects under dynamic, unpredictable conditions.
  • To create a compact, efficient mapping from visual features to robotic end-effector velocities using sparse representation.

Proposed method

  • Proposes a novel histogram of oriented wrinkles (HOW) feature representation computed from RGB video streams using Gabor filters to extract high- and low-frequency components.
  • Employs an offline training phase to build a visual feedback dictionary that maps visual features (HOW-features) to end-effector velocities via sampling and clustering.
  • Uses sparse linear representation (basis pursuit denoising) to compute real-time control velocities from the visual feedback dictionary, balancing data fitting and sparsity.
  • Combines HOW-features with HOG and color histograms to enhance feature representation, with fusion shown to improve performance.
  • Integrates the system with a 12-DOF ABB YuMi dual-arm robot and RGB camera for real-time execution on complex manipulation tasks.
  • Defines goal configurations based on task-specific demonstrations and uses visual feedback to guide the robot toward the desired state.

Experimental results

Research questions

  • RQ1Can a learned visual feedback dictionary based on histogram features enable accurate and real-time control of robotic manipulation for highly deformable materials?
  • RQ2How effective is the proposed histogram of oriented wrinkles (HOW) feature in capturing shape variation and deformation in cloth under large deformations?
  • RQ3To what extent does sparse representation over the visual feedback dictionary improve control accuracy and robustness in dynamic manipulation tasks?
  • RQ4Can the system handle human-robot collaborative manipulation where external perturbations occur due to human interaction?

Key findings

  • The HOG+HOW feature combination achieved the highest success rate of 94.53% in benchmark 4, outperforming individual features and other combinations.
  • The HOW feature alone achieved 92.21% success rate in benchmark 4, demonstrating its effectiveness in capturing deformation and wrinkles in cloth.
  • The visual feedback dictionary with sparse representation reduced velocity error, and performance saturated beyond a certain dictionary size, indicating redundancy.
  • The slack variable α in sparse representation governs the trade-off between data fitting and sparsity, with optimal values improving convergence and reducing error.
  • The system successfully performed real-time manipulation on a dual-arm robot for stretching, folding, and placement tasks, including human-robot collaboration.
  • The method achieved high performance without requiring fiducial markers or extensive training data, relying instead on visual features and precomputed dictionaries.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.