[Paper Review] Robot Vision Architecture for Autonomous Clothes Manipulation
This paper presents a novel hierarchical robot vision architecture for autonomous 3D perception of clothing configurations, enabling dexterous manipulation via a custom stereo vision system and dual-arm robot. It achieves superior performance in autonomous grasping and, for the first time, demonstrates a fully autonomous dual-arm flattening system that reduces manipulation iterations by over 50% compared to single-arm or Kinect-based methods.
This paper presents a novel robot vision architecture for perceiving generic 3D clothes configurations. Our architecture is hierarchically structured, starting from low-level curvatures, across mid-level geometric shapes \& topology descriptions; and finally approaching high-level semantic surface structure descriptions. We demonstrate our robot vision architecture in a customised dual-arm industrial robot with our self-designed, off-the-self stereo vision system, carrying out autonomous grasping and dual-arm flattening. It is worth noting that the proposed dual-arm flattening approach is unique among the state-of-the-art robot autonomous system, which is the major contribution of this paper. The experimental results show that the proposed dual-arm flattening using stereo vision system remarkably outperforms the single-arm flattening and widely-cited Kinect-based sensing system for dexterous manipulation tasks. In addition, the proposed grasping approach achieves satisfactory performance on grasping various kind of garments, verifying the capability of proposed visual perception architecture to be adapted to more than one clothing manipulation tasks.
Motivation & Objective
- Address the lack of generic, reusable vision architectures for deformable object manipulation in robotics.
- Overcome limitations of ad-hoc, sensor-specific solutions for garment perception and manipulation.
- Develop a perception system capable of parsing complex 3D garment configurations for multiple manipulation tasks.
- Enable efficient, autonomous dual-arm flattening of wrinkled garments using high-quality depth sensing.
- Demonstrate the superiority of a custom stereo vision system over commercial RGB-D sensors like Kinect for dexterous manipulation tasks.
Proposed method
- Design a hierarchical visual perception architecture that processes 3D visual features from low-level curvatures to high-level semantic surface structures.
- Integrate a high-resolution, off-the-shelf active binocular stereo head with a GPU-accelerated stereo matcher tuned for clothing textures and depth accuracy.
- Perform hand-eye calibration between the stereo vision system and dual-arm robot to enable precise spatial alignment for manipulation.
- Use 3D surface shape and topology analysis to detect key garment features such as wrinkles and grasping triplets.
- Implement perception-manipulation cycles where visual feedback guides iterative flattening actions using dual-arm coordination.
- Apply a dual-arm strategy that applies symmetric, coordinated forces to flatten long wrinkles more efficiently than single-arm approaches.
Experimental results
Research questions
- RQ1Can a generic, hierarchical vision architecture be designed to parse 3D garment configurations for multiple manipulation tasks?
- RQ2How does the performance of a custom high-resolution stereo vision system compare to commercial RGB-D sensors like Kinect in dexterous clothing manipulation?
- RQ3To what extent does a dual-arm flattening strategy reduce the number of required manipulation iterations compared to single-arm or table-based methods?
- RQ4Can the proposed vision architecture robustly detect grasping points and structural features (e.g., wrinkles) across diverse garment types?
- RQ5Does the integration of high-quality depth sensing with semantic 3D parsing significantly improve the success rate and efficiency of autonomous clothing manipulation?
Key findings
- The proposed dual-arm flattening system reduced the average number of iterations (RNI) to 4.7 for towels, 6.3 for shorts, and 8.9 for t-shirts, demonstrating adaptability across garment types.
- Dual-arm flattening required only 4.7 iterations on average, compared to 9.5 iterations with single-arm flattening, showing a 50% improvement in efficiency.
- The custom stereo vision system outperformed the Kinect-like Xtion sensor by reducing RNI by over half, due to higher depth resolution and lower noise (less than 0.5cm high-frequency noise).
- The stereo head enabled accurate detection of long wrinkles and precise displacement estimation, avoiding the splitting of long wrinkles into short ones that plagues lower-quality depth sensors.
- The grasping approach achieved robust performance across various garment types, confirming the adaptability of the perception architecture to multiple tasks.
- The hierarchical visual architecture successfully parsed complex 3D configurations by detecting and quantifying landmarks such as wrinkles and grasping triplets, enabling reliable manipulation planning.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.