Skip to main navigation Skip to search Skip to main content

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video

  • Weili Guan
  • , Haoyu Zhang*
  • , Meng Liu
  • , Qianlong Xiang
  • , Yaowei Wang
  • , Liqiang Nie
  • *Corresponding author for this work
  • School of Information Science and Technology, Harbin Institute of Technology Shenzhen
  • Harbin Institute of Technology
  • Pengcheng Laboratory
  • Shandong Jianzhu University
  • Zhongguancun Academy
  • City University of Hong Kong
  • Shenzhen Loop Area Institute

Research output: Contribution to journalArticlepeer-review

Abstract

Visual-spatial understanding, defined as the ability to infer object relationships and scene layouts from visual inputs, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, pre-trained vision-language models (VLMs) remain constrained by spatial uncertainty stemming from inherently 2D observations and by the scarcity of data for 3D spatial understanding. To address these limitations, we proposed a novel framework, SpaceEra, in the NeurIPS 2025 Spotlight paper. Although it achieved significant performance gains, we further observed that its effectiveness is hindered by insufficient input from scanning videos and weak reasoning constraints. To tackle these newly emerged challenges, we extend the original framework into a comprehensive system, termed SpaceEra++, which spans data construction, model design, training optimization, and prompting inference. Specifically, to alleviate input insufficiency, we introduce ScenePick, a frame sampling strategy that balances spatial coverage with object semantics to produce compact yet comprehensive scene representations. In addition, to enhance spatial reasoning, we develop SpaceAlign, which enforces pairwise object constraints by jointly exploiting absolute coordinates and relative spatial relations, thereby aligning optimization with spatial accuracy. Extensive experiments across multiple benchmarks demonstrate consistent improvements over strong baselines, while ablation studies validate both the individual and joint contributions of each component, and further analyses provide guidance for future research.

Original languageEnglish
JournalIEEE Transactions on Pattern Analysis and Machine Intelligence
DOIs
StateAccepted/In press - 2026
Externally publishedYes

Keywords

  • Frame Sampling
  • Reinforcement Learning
  • Spatial Reasoning
  • Vision-Language Model

Fingerprint

Dive into the research topics of 'SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video'. Together they form a unique fingerprint.

Cite this