Skip to main navigation Skip to search Skip to main content

EMMA: EMPOWERING MULTI-MODAL MAMBA WITH STRUCTURAL AND HIERARCHICAL ALIGNMENT

  • Yifei Xing
  • , Xiangyuan Lan*
  • , Ruiping Wang*
  • , Dongmei Jiang
  • , Wenjun Huang
  • , Qingfang Zheng
  • , Yaowei Wang
  • *Corresponding author for this work
  • CAS - Institute of Computing Technology
  • Pengcheng Laboratory
  • University of Chinese Academy of Sciences
  • Pazhou Laboratory (Huangpu)
  • Sun Yat-Sen University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Mamba-based architectures have shown to be a promising new direction for deep learning models owing to their competitive performance and sub-quadratic deployment speed. However, current Mamba multi-modal large language models (MLLM) are insufficient in extracting visual features, leading to imbalanced cross-modal alignment between visual and textural latents, negatively impacting performance on multi-modal tasks. In this work, we propose Empowering Multi-modal Mamba with Structural and Hierarchical Alignment (EMMA), which enables the MLLM to extract fine-grained visual information. Specifically, we propose a pixel-wise alignment module to autoregressively optimize the learning and processing of spatial image-level features along with textual tokens, enabling structural alignment at the image level. In addition, to prevent the degradation of visual information during the cross-model alignment process, we propose a multi-scale feature fusion (MFF) module to combine multi-scale visual features from intermediate layers, enabling hierarchical alignment at the feature level. Extensive experiments are conducted across a variety of multi-modal benchmarks. Our model shows lower latency than other Mamba-based MLLMs and is nearly four times faster than transformer-based MLLMs of similar scale during inference. Due to better cross-modal alignment, our model exhibits lower degrees of hallucination and enhanced sensitivity to visual details, which manifests in superior performance across diverse multi-modal benchmarks. Code provided at https://github.com/xingyifei2016/EMMA.

Original languageEnglish
Title of host publication13th International Conference on Learning Representations, ICLR 2025
PublisherInternational Conference on Learning Representations, ICLR
Pages33369-33397
Number of pages29
ISBN (Electronic)9798331320850
StatePublished - 2025
Externally publishedYes
Event13th International Conference on Learning Representations, ICLR 2025 - Singapore, Singapore
Duration: 24 Apr 202528 Apr 2025

Publication series

Name13th International Conference on Learning Representations, ICLR 2025

Conference

Conference13th International Conference on Learning Representations, ICLR 2025
Country/TerritorySingapore
CitySingapore
Period24/04/2528/04/25

Fingerprint

Dive into the research topics of 'EMMA: EMPOWERING MULTI-MODAL MAMBA WITH STRUCTURAL AND HIERARCHICAL ALIGNMENT'. Together they form a unique fingerprint.

Cite this