Repository logo
  • Research Outputs
  • Researchers
  • Schools
    Felizberta Lo Padilla Tong School of Social SciencesIp Ying To Lee Yu Yee School of Humanities and LanguagesRita Tong Liu School of Business and Hospitality ManagementS.K. Yee School of Health SciencesYam Pak Charitable Foundation School of Computing and Information Sciences
  • Help
Repository logo
  1. Home
  2. Humanities and Languages
  3. HL Publication
  4. A cognitive-inspired multimodal information fusion algorithm
 
  • Details

A cognitive-inspired multimodal information fusion algorithm

Author(s)
Lin, Qiuhan  
Author(s)
Yao, Z.
Date Issued
2025
Publisher
World Scientific Publishing Company
Journal
International Journal of Pattern Recognition and Artificial Intelligence
Abstract
Multimodal information fusion plays a critical role in enabling intelligent systems to process and reason over heterogeneous data sources such as vision, language, and audio. Traditional fusion methods often struggle with challenges related to semantic alignment, cross-modal dependencies, and robustness in complex environments. This study proposes a cognitive- and psychology-inspired multimodal information fusion algorithm that explicitly incorporates principles such as selective attention, working memory integration, and hierarchical reasoning into a novel optimal transport-based semantic alignment and adaptive attention fusion framework, distinguishing it from existing multimodal fusion methods. The proposed algorithm first encodes and aligns heterogeneous modalities — images, text, and audio — within a shared latent space using a 2-Wasserstein optimal transport alignment strategy, then applies a cognitive-inspired semantic attention mechanism to dynamically weigh modality contributions for downstream tasks such as classification and retrieval. Extensive experiments on four public benchmarks (LUMA, WIT, VGGSound, and MusicTM) show consistent gains, e.g. achieving up to +2.4% accuracy improvement and +2.1% F1-score over the best baseline on cross-modal classification tasks. The method’s robustness is demonstrated across diverse modality combinations and domain complexities, highlighting its practical relevance for real-world applications such as cross-modal retrieval, intelligent decision support, and multi-sensor integration. Future work will explore large-scale cross-modal adaptation and low-resource settings to further enhance scalability and generalization.
URI
https://repository.sfu.edu.hk/handle/sfu/5309
DOI
10.1142/S0218001425590219
SFU Affiliated Publication
No
Availability at SFU Library

No database links found.

Responsible Use of E‑Resources | Privacy Policy | Disclaimer
© SFU Library. All Rights Reserved.
SFU Library