Abstract
Hyperbolic representations have shown great promise for modeling visual semantic hierarchies, owing to their exponential representational capacity and inherent ability to encode hierarchical geometric structures. However, the existing hyperbolic vision–language model, MERU, suffers from two fundamental limitations, which we identify and substantiate through both theoretical analysis and empirical evaluation. First, the half-aperture angles frequently exceed the valid domain of the arcsin function, causing the entailment cones to degenerate into hemispheres and constraining the model’s representational expressiveness. Second, the entailment loss considers only positive samples, lacking explicit constraints to exclude negative examples; empirically, we observe that approximately 50% of negative samples fall within each entailment cone. To overcome these issues, we propose HyperConeNet, a novel framework that integrates a contrastive learning module, an entailment module, and an aperture diversity module. HyperConeNet decouples half-aperture prediction from curvature, reformulating it as a learnable parameterization, and incorporates both positive and negative samples into the entailment loss to strengthen the model’s discriminative power. Extensive experiments on 14 image classification and retrieval benchmarks demonstrate that HyperConeNet consistently outperforms baseline models while maintaining geometrically valid cone structures. Moreover, the predicted half-aperture angles remain strictly within the valid range (0, π/2), ensuring stable and interpretable geometry. Further empirical analyses confirm that no negative samples fall within any entailment cone, indicating that HyperConeNet effectively addresses the two principal limitations of MERU. The code will be made publicly available soon.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Circuits and Systems for Video Technology |
| DOIs | |
| State | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- Contrastive Learning
- Entailment Cone
- Hierarchical Modeling
- Hyperbolic Embeddings
Fingerprint
Dive into the research topics of 'Enhancing Hyperbolic Vision–Language Representations via Adaptive Apertures and Contrastive Entailment'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver