Skip to main navigation Skip to search Skip to main content

Enhancing Hyperbolic Vision–Language Representations via Adaptive Apertures and Contrastive Entailment

  • Changli Wang
  • , Fang Yin
  • , Suo Gao
  • , Rui Wu*
  • *Corresponding author for this work
  • Faculty of Computing, Harbin Institute of Technology
  • Harbin University of Science and Technology
  • Dalian Polytechnic University

Research output: Contribution to journalArticlepeer-review

Abstract

Hyperbolic representations have shown great promise for modeling visual semantic hierarchies, owing to their exponential representational capacity and inherent ability to encode hierarchical geometric structures. However, the existing hyperbolic vision–language model, MERU, suffers from two fundamental limitations, which we identify and substantiate through both theoretical analysis and empirical evaluation. First, the half-aperture angles frequently exceed the valid domain of the arcsin function, causing the entailment cones to degenerate into hemispheres and constraining the model’s representational expressiveness. Second, the entailment loss considers only positive samples, lacking explicit constraints to exclude negative examples; empirically, we observe that approximately 50% of negative samples fall within each entailment cone. To overcome these issues, we propose HyperConeNet, a novel framework that integrates a contrastive learning module, an entailment module, and an aperture diversity module. HyperConeNet decouples half-aperture prediction from curvature, reformulating it as a learnable parameterization, and incorporates both positive and negative samples into the entailment loss to strengthen the model’s discriminative power. Extensive experiments on 14 image classification and retrieval benchmarks demonstrate that HyperConeNet consistently outperforms baseline models while maintaining geometrically valid cone structures. Moreover, the predicted half-aperture angles remain strictly within the valid range (0, π/2), ensuring stable and interpretable geometry. Further empirical analyses confirm that no negative samples fall within any entailment cone, indicating that HyperConeNet effectively addresses the two principal limitations of MERU. The code will be made publicly available soon.

Original languageEnglish
JournalIEEE Transactions on Circuits and Systems for Video Technology
DOIs
StateAccepted/In press - 2026
Externally publishedYes

Keywords

  • Contrastive Learning
  • Entailment Cone
  • Hierarchical Modeling
  • Hyperbolic Embeddings

Fingerprint

Dive into the research topics of 'Enhancing Hyperbolic Vision–Language Representations via Adaptive Apertures and Contrastive Entailment'. Together they form a unique fingerprint.

Cite this