Abstract
Effective dexterous manipulation hinges on the dynamic integration of visual context with fine-grained tactile feedback. This remains a significant challenge, as existing methods often rely on static fusion strategies and struggle to learn from isolated tactile features. To address this, we propose a hardware-decoupled tactile representation that learns cross-finger spatial features by training a sparse Transformer on unified tactile images, enabling cross-device generalization. Furthermore, we introduce a tactile-activity-guided adaptive visual-tactile fusion mechanism that dynamically adjusts the influence of vision and touch, showing the contribution of touch feedback upon physical contact. We evaluate our method on a series of contact-rich manipulation tasks requiring fine force control. Experimental results show that our method has an average success rate of over 90%, demonstrating effectiveness compared to other methods. Further analysis shows that adaptive multimodal fusion is essential to complete the dexterous manipulation tasks.
| Original language | English |
|---|---|
| Pages (from-to) | 6584-6591 |
| Number of pages | 8 |
| Journal | IEEE Robotics and Automation Letters |
| Volume | 11 |
| Issue number | 6 |
| DOIs | |
| State | Published - 1 Jun 2026 |
| Externally published | Yes |
Keywords
- Dexterous manipulation
- force and tactile sensing
- representation learning
- sensor fusion
Fingerprint
Dive into the research topics of 'Adaptive Visual-Tactile Fusion for Contact-Rich Dexterous Manipulation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver