Skip to main navigation Skip to search Skip to main content

GEOMR: Integrating image geographic features and human reasoning knowledge for image geolocalization

  • Jian Fang
  • , Siyi Qian
  • , Shaohui Liu*
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • Peking University

Research output: Contribution to journalArticlepeer-review

Abstract

Worldwide image geolocalization aims to accurately predict the geographic location where a given image was captured. Due to the vast scale of the Earth and the uneven distribution of geographic features, this task remains highly challenging. Traditional methods exhibit clear limitations when handling global-scale data. To address these challenges, we propose GEOMR, an effective and adaptive framework that integrates image geographic features and human reasoning knowledge to enhance global geolocalization accuracy. GEOMR consists of two modules. The first module extracts geographic features from images by jointly learning multimodal features. The second module involves training a multimodal large language model in a two-phase process to enhance its geolocalization reasoning capabilities. The first phase learns human geolocalization reasoning knowledge, enabling the model to utilize geographic cues present in images effectively. The second phase focuses on learning how to use reference information to infer the correct geographic coordinates. Extensive experiments conducted on the IM2GPS3K, YFCC4K, and YFCC26K datasets demonstrate that GEOMR significantly outperforms state-of-the-art methods.

Original languageEnglish
Article number115391
JournalKnowledge-Based Systems
Volume337
DOIs
StatePublished - 25 Mar 2026

Keywords

  • Geographic feature extraction
  • Human reasoning knowledge
  • Multimodal large language model

Fingerprint

Dive into the research topics of 'GEOMR: Integrating image geographic features and human reasoning knowledge for image geolocalization'. Together they form a unique fingerprint.

Cite this