Repository logo
  • Research Outputs
  • Researchers
  • Schools
    Felizberta Lo Padilla Tong School of Social SciencesIp Ying To Lee Yu Yee School of Humanities and LanguagesRita Tong Liu School of Business and Hospitality ManagementS.K. Yee School of Health SciencesYam Pak Charitable Foundation School of Computing and Information Sciences
  • Help
Repository logo
  1. Home
  2. Computing and Information Sciences
  3. CIS Publication
  4. Test-to-image generation using NLP: Transformer vs. hybrid ML approaches
 
  • Details

Test-to-image generation using NLP: Transformer vs. hybrid ML approaches

Author(s)
Arul, Joseph Maria
Cheung, William Chi Keung  
Chan, Anthony Hing-Hung  
Date Issued
2025
Publisher
Society of Photo-Optical Instrumentation Engineers (SPIE)
Related Publication(s)
Proceedings of International Conference on AI-Generated Content (AIGC 2025)
Volume
14183
Abstract
Recent breakthroughs in natural language processing (NLP) have enabled more natural and intuitive communication between humans and machines. While many current state-of-the-art methods rely heavily on Generative architectures, this study explores alternative strategies for image-text retrieval. Specifically, it investigates the use of CLIP (Contrastive Language-Image Pretraining) for encoding both image and text data, as well as a hybrid model that incorporates additional shape and color features via DeepLabV3. In the transformer-based approach, the textual prompts are encoded using CLIP’s text encoder. Cosine similarity is computed between these vectors and CLIP-encoded image features to retrieve the most relevant image. The second hybrid model integrates vision-language-aligned CLIP features with CNN-based DeepLabV3 features using a learned fusion module, named Prompt-Aware Fusion, which attends to both modalities using the text prompt as a query. This fusion mechanism enables the model to combine semantic, shape, and texture cues for improved retrieval accuracy. Experimental results on the Oxford 102 flower dataset demonstrates that the hybrid approach outperforms the standalone CLIP model in top-1 precision, highlighting the benefits of combining transformer and CNNbased representations in a unified framework. To evaluate the performance of both models and present the results, 50 random text queries were generated. The retrieved images were analyzed, and the number of correctly retrieved images by each model was recorded. The CLIP model successfully retrieved 37 out of 50 images, while the hybrid model achieved correct retrieval for all 50 images. This shows that the hybrid model performance greatly improved in this text-to-image retrieval.
URI
https://repository.sfu.edu.hk/handle/sfu/5396
DOI
10.1117/12.3109371
SFU Affiliated Publication
Yes
Availability at SFU Library

No database links found.

Responsible Use of E‑Resources | Privacy Policy | Disclaimer
© SFU Library. All Rights Reserved.
SFU Library