Automated Caption-Guided Image Retrieval: A Multi-Positive Contrastive Learning Paradigm

Authors

  • Danyang Cao School of Artificial Intelligence and Computer Science, North China University of Technology, Beijing 100144, China
  • Liang Cheng School of Artificial Intelligence and Computer Science, North China University of Technology, Beijing 100144, China
  • Hongbo Zhou School of Artificial Intelligence and Computer Science, North China University of Technology, Beijing 100144, China

DOI:

https://doi.org/10.54691/g45e5263

Keywords:

Image-to-Image Retrieval; Semantic Representation; Multi-Positive Contrastive Loss; Contrastive Learning.

Abstract

[Objective] This study aims to improve textual similarity measurement for image retrieval while mitigating the anisotropy of sentence embeddings. [Methods] We propose an image-caption semantic encoder trained with a multi-positive contrastive loss. The conventional contrastive objective is extended to accommodate multiple positive samples, and image captioning is used to generate training data automatically. On this basis, we construct a caption-based image-to-image retrieval framework. [Results] Experiments show that the proposed model outperforms baseline methods on semantic textual similarity (STS) benchmarks and improves the agreement between retrieval results and human semantic judgments. [Limitations] Short captions cannot fully represent the complex semantics, ambiguity, and fine-grained details of an image. [Conclusion] The proposed Multi-Positive Example Contrastive Learning (MPC) model provides more discriminative sentence embeddings and improves semantic similarity measurement in image retrieval.

Downloads

Download data is not yet available.

References

[1] Dubey, S. R. (2021). A decade survey of content based image retrieval using deep learning. IEEE Transactions on Circuits and Systems for Video Technology, 32(5), 2687–2704. https://doi.org/10.1109/tcsvt.2021.3080920.

[2] Huang, K. (2018). Content based image retrieval using generated textual meta data. In Proceedings of the 2nd International Conference on Advances in Artificial Intelligence (pp. 16–19).

[3] Yoon, S., Kang, W. Y., Jeon, S., Lee, S., Han, C., Park, J., & Kim, E. S. (2021). Image to image retrieval by learning similarity between scene graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, 35(12), 10718–10726. https://doi.org/10.1609/aaai.v35i12.17281AAAI Publi...

[4] Li, J., Li, D., Xiong, C., & Hoi, S. (2023). BLIP: Bootstrapping language image pre training for unified vision language understanding and generation. International Journal of Computer Vision, 131, 1795–1814.

[5] Reimers, N., & Gurevych, I. (2019). Sentence BERT: Sentence embeddings using Siamese BERT Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP IJCNLP) (pp. 3982–3992). https://doi.org/10.18653/v1/D19 1410

[6] Li, B., Zhou, H., He, J., Wang, M., Yang, Y., & Li, L. (2020). On the sentence embeddings from pre trained language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 9119–9130).

[7] Ismail, F., Ramzan, N., Sajid, M., & Khan, H. (2022). Deep learning for image retrieval: Recent progress and challenges. Information Fusion, 82, 59–83.

[8] Yang, H., & Shi, S. (2023). A review of content based image retrieval technology. Software Guide, 22(4), 229–244.

[9] Wang, Y., Zhang, J., Li, X., & Wang, S. (2023). Multimodal image retrieval: A survey of methods and applications. Information Fusion, 93, 24–45.

[10] Li, Y., Mo, H., Wang, S., & Liu, Q. (2018). A person retrieval method based on image descriptions. Journal of System Simulation, 7, 2794–2800.

[11] Wei, X., Qi, Y., Liu, J., & Chen, J. (2017). Image retrieval by dense caption reasoning. In 2017 IEEE Visual Communications and Image Processing (VCIP) (pp. 1–4).

[12] Yoon, S., Kang, W. Y., Jeon, S., Lee, S., Han, C., Park, J., & Kim, E. S. (2021). Image to image retrieval by learning similarity between scene graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, 35(12), 10718–10726. https://doi.org/10.1609/aaai.v35i12.17281AAAI Publi...

[13] Mikolov, T., Sutskever, I., Chen, K., Corrado, G., & Dean, J. (2013). Distributed representations of words and phrases and their compositionality. Advances in Neural Information Processing Systems, 26.

[14] Pennington, J., Socher, R., & Manning, C. D. (2014). Glove: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 1532–1543).

[15] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). Bert: Pre training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 4171–4186). https://doi.org/10.18653/v1/N19 1423

[16] Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.

[17] Su, J., Cao, J., Liu, W., & Li, Y. (2021). Whitening sentence representations for better semantics and faster retrieval. arXiv preprint arXiv:2103.15316.

[18] Musgrave, K., Belongie, S., & Lim, S. N. (2023). A metric learning reality check. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(3), 3272–3289.

[19] He, K., Fan, H., Wu, Y., Xie, S., & Girshick, R. (2020). Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9729–9738).

[20] Lin, T. Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., & Zitnick, C. L. (2014). Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6 12, 2014, Proceedings, Part V 13 (pp. 740–755). Springer International Publishing.

[21] Young, P., Lai, A., Hodosh, M., & Hockenmaier, J. (2014). From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2, 67–78.

[22] Suhr, A., Zhou, S., Zhang, A., Zhang, I., Bai, H., & Artzi, Y. (2019). A corpus for reasoning about natural language grounded in photographs. In Proceedings of the Annual Meeting of the Association for Computational Linguistics.

[23] Cao, D., Zhou, H., & Wang, Y. (2025). Improve the image caption generation on out domain dataset by external knowledge augmented. Multimedia Systems, 31(1), 1–18.

Downloads

Published

20-08-2026

Issue

Section

Articles