Skip to main navigation Skip to search Skip to main content

MossVLN: Memory-Observation Synergistic System for Continuous Vision-Language Navigation

  • Ting Yu*
  • , Yifei Wu
  • , Qiongjie Cui
  • , Qingming Huang
  • , Jun Yu
  • *Corresponding author for this work
  • Hangzhou Normal University
  • Singapore University of Technology and Design
  • University of Chinese Academy of Sciences
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Navigating in continuous environments with vision-language cues presents critical challenges, particularly in the accuracy of waypoint prediction and the quality of navigation decision-making. Traditional methods, which predominantly rely on spatial data from depth images or straightforward RGB-depth integrations, frequently encounter difficulties in environments where waypoints share similar spatial characteristics, leading to erroneous navigational outcomes. Additionally, the capacity for effective navigation decisions is often hindered by the inadequacies of traditional topological maps and the issue of uneven data sampling. In response, this paper introduces a robust memory-observation synergistic vision-language navigation framework to substantially enhance the navigation capabilities of agents operating in continuous environments. We present an advanced observation-driven waypoint predictor that effectively utilizes spatial data and integrates aligned visual and textual cues to significantly improve the accuracy of waypoint predictions within complex real-world scenarios. Additionally, we develop a strategic memory-observation planning approach that leverages memory panoramic environmental data and detailed current observation information, enabling more informed and precise navigation decisions. Our framework sets new performance benchmarks on the VLN-CE dataset, achieving a 60.25% Success Rate (SR) and a 50.89% Path Length Score (SPL) on the R2R-CE dataset’s unseen validation splits. Furthermore, when adapted to a discrete environment, our model also shows exceptional performance on the R2R dataset, achieving a 74% SR and a 64% SPL on the unseen validation split.

Original languageEnglish
Pages (from-to)6690-6704
Number of pages15
JournalIEEE Transactions on Multimedia
Volume27
DOIs
StatePublished - 2025
Externally publishedYes

Keywords

  • Continuous environment
  • memory observation synergistic navigation
  • navigation
  • navigation decision
  • vision-language
  • waypoints predictor

Fingerprint

Dive into the research topics of 'MossVLN: Memory-Observation Synergistic System for Continuous Vision-Language Navigation'. Together they form a unique fingerprint.

Cite this