Skip to main navigation Skip to search Skip to main content

HVLF: A Holistic Visual Localization Framework Across Diverse Scenes

  • Harbin Institute of Technology
  • Yangtze River Delta HIT Robot Technology Research Institute

Research output: Contribution to journalArticlepeer-review

Abstract

Recently, integrating the multitask learning (MTL) paradigm into scene coordinate regression (SCoRe) techniques has achieved significant success in visual localization tasks. However, the feature extraction ability of existing frameworks is inherently constrained by the rigid weight activation strategy, which prevents each layer from concurrently capturing scene-universal features across diverse scenes and scene-particular attributes unique to each individual scene. In addition, the straightforward network architecture further exacerbates the issue of insufficient feature representation. To address these limitations, we introduce HVLF, a holistic framework that ensures flexible identification of both scene-universal and scene-particular attributes while integrating various attention mechanisms to enhance feature representation effectively. Technically, for the first issue, HVLF proposes a soft weight activation strategy (SWAS) equipped with polyhedral convolution to concurrently optimize scene-shared and scene-specific weights within each layer, which facilitates sufficient discernment of both scene-universal features and scene-particular attributes, thereby boosting the network's capability for comprehensive scene perception. For the second issue, HVLF introduces a mixed attention perception module (MAPM) that incorporates channelwise, spatialwise, and elementwise attention mechanisms to perform multilevel feature fusion, hence extracting discriminative features to regress precise scene coordinates. Extensive experiments on indoor and outdoor datasets prove that HVLF realizes impressive localization performance. In addition, experiments conducted on 3-D object detection and feature matching tasks prove that the two proposed techniques are universal and can be seamlessly inserted into other methods.

Original languageEnglish
Pages (from-to)18859-18873
Number of pages15
JournalIEEE Transactions on Neural Networks and Learning Systems
Volume36
Issue number10
DOIs
StatePublished - 2025

Keywords

  • Multitask learning (MTL)
  • scene coordinate regression (SCoRe)
  • visual localization

Fingerprint

Dive into the research topics of 'HVLF: A Holistic Visual Localization Framework Across Diverse Scenes'. Together they form a unique fingerprint.

Cite this