The style transformer with common knowledge optimization for image-text retrieval

Li, Wenrui; Ma, Zhengyu; Shi, Jinqiao; Fan, Xiaopeng

Computer Science > Computer Vision and Pattern Recognition

arXiv:2303.00448 (cs)

[Submitted on 1 Mar 2023 (v1), last revised 3 Apr 2023 (this version, v2)]

Title:The style transformer with common knowledge optimization for image-text retrieval

Authors:Wenrui Li, Zhengyu Ma, Jinqiao Shi, Xiaopeng Fan

View PDF

Abstract:Image-text retrieval which associates different modalities has drawn broad attention due to its excellent research value and broad real-world application. However, most of the existing methods haven't taken the high-level semantic relationships ("style embedding") and common knowledge from multi-modalities into full consideration. To this end, we introduce a novel style transformer network with common knowledge optimization (CKSTN) for image-text retrieval. The main module is the common knowledge adaptor (CKA) with both the style embedding extractor (SEE) and the common knowledge optimization (CKO) modules. Specifically, the SEE uses the sequential update strategy to effectively connect the features of different stages in SEE. The CKO module is introduced to dynamically capture the latent concepts of common knowledge from different modalities. Besides, to get generalized temporal common knowledge, we propose a sequential update strategy to effectively integrate the features of different layers in SEE with previous common feature units. CKSTN demonstrates the superiorities of the state-of-the-art methods in image-text retrieval on MSCOCO and Flickr30K datasets. Moreover, CKSTN is constructed based on the lightweight transformer which is more convenient and practical for the application of real scenes, due to the better performance and lower parameters.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
Cite as:	arXiv:2303.00448 [cs.CV]
	(or arXiv:2303.00448v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2303.00448

Submission history

From: Wenrui Li [view email]
[v1] Wed, 1 Mar 2023 12:17:33 UTC (300 KB)
[v2] Mon, 3 Apr 2023 11:17:11 UTC (300 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:The style transformer with common knowledge optimization for image-text retrieval

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:The style transformer with common knowledge optimization for image-text retrieval

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators