Lin, QiuhanQiuhanLinYao, Z.2025-12-232025-12-232025https://repository.sfu.edu.hk/handle/sfu/5309Multimodal information fusion plays a critical role in enabling intelligent systems to process and reason over heterogeneous data sources such as vision, language, and audio. Traditional fusion methods often struggle with challenges related to semantic alignment, cross-modal dependencies, and robustness in complex environments. This study proposes a cognitive- and psychology-inspired multimodal information fusion algorithm that explicitly incorporates principles such as selective attention, working memory integration, and hierarchical reasoning into a novel optimal transport-based semantic alignment and adaptive attention fusion framework, distinguishing it from existing multimodal fusion methods. The proposed algorithm first encodes and aligns heterogeneous modalities — images, text, and audio — within a shared latent space using a 2-Wasserstein optimal transport alignment strategy, then applies a cognitive-inspired semantic attention mechanism to dynamically weigh modality contributions for downstream tasks such as classification and retrieval. Extensive experiments on four public benchmarks (LUMA, WIT, VGGSound, and MusicTM) show consistent gains, e.g. achieving up to +2.4% accuracy improvement and +2.1% F1-score over the best baseline on cross-modal classification tasks. The method’s robustness is demonstrated across diverse modality combinations and domain complexities, highlighting its practical relevance for real-world applications such as cross-modal retrieval, intelligent decision support, and multi-sensor integration. Future work will explore large-scale cross-modal adaptation and low-resource settings to further enhance scalability and generalization.enA cognitive-inspired multimodal information fusion algorithmjournal article10.1142/S0218001425590219