IP Library › Granted Patent US 12,670,236
Granted Patent B2
US 12,670,236 · App. 17/502,385 · Granted Jun 30, 2026

Method for training cross-modal retrieval model, electronic device and storage medium

Inventors: Feng He (Beijing, CN); Qi Wang (Beijing, CN); Zhifan Feng (Beijing, CN); Hu Yang (Beijing, CN); Chunguang Chai (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
G06F18/256G06F18/2178G06F18/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,236
App. No.
17/502,385
Filed
Oct 15, 2021
Granted
Jun 30, 2026
Kind
B2
Art Unit
2144
USPC
706/45
Abstract

The present disclosure discloses a method for training a cross-modal retrieval model, an electronic device and a storage medium, and relates to the field of computer technologies, and particularly to the field of artificial intelligence technologies, such as knowledge graph technologies, computer vision technologies, deep learning technologies, or the like. The method for training a cross-modal retrieval model includes: determining similarity of a cross-modal sample pair according to the cross-modal sample pair, the cross-modal sample pair including a sample of a first modal and a sample of a second modal, and the first modal being different from the second modal; determining a soft margin based on the similarity, and determining a soft margin loss function based on the soft margin; and determining a total loss function based on the soft margin loss function, and training a cross-modal retrieval model according to the total loss function.

Claims (64)

1 . A computer-implemented method for training a cross-modal retrieval model which is used by a cross-modal retrieval system, wherein a cross-modal retrieval is used for a retrieval of data of one modal using data of another modal, the method comprising:

determining a similarity of a cross-modal sample pair according to the cross-modal sample pair, the cross-modal sample pair comprising a sample of a first modal and a sample of a second modal, and the first modal being different from the second modal, wherein the first modal is a text, and the second modal is a video, and the cross-modal sample pair comprises a positive sample pair and a negative sample pair, the positive sample pair comprises an anchor sample and a positive sample, the negative sample pair comprises the anchor sample and a negative sample, the anchor sample has the first modal, and the positive sample and the negative sample have the second modal, and wherein the anchor sample is a text in a sample set, the positive sample is a video related to the text in the sample set, and the negative sample is a randomly selected video which is related or not related to the text in the sample set;

determining a soft margin based on the similarity, and determining a soft margin loss function based on the soft margin, wherein the soft margin is a non-fixed value; and

determining a total loss function based on the soft margin loss function, and

training a cross-modal retrieval model according to the total loss function,

using the trained cross-modal retrieval model to receive a text input by a user, determine a video matched with the text using the cross-modal retrieval model, and feed the matched video back to the user,

wherein the determining a soft margin based on the similarity comprises:

calculating a distance between similarity of the positive sample pair and similarity of the negative sample pair to obtain a similarity distance; and

normalizing the similarity distance to obtain a normalized similarity distance, and determining the normalized similarity distance as the soft margin, and

wherein the similarity distance comprises a similarity distance in the first modal and a similarity distance in the second modal, the similarity distance in the first modal being obtained by processing the sample pair in the first modal using a semantic representation model in the first modal and the similarity distance in the second modal being obtained by processing the sample pair in the second modal using a semantic representation model in the second modal;

the determining a soft margin based on the similarity and determining a soft margin loss function based on the soft margin comprises:

determining a soft margin in the first modal based on the similarity distance in the first modal, and calculating a contrastive loss function in the first modal based on the soft margin in the first modal;

determining a soft margin in the second modal based on the similarity distance in the second modal, and calculating a contrastive loss function in the second modal based on the soft margin in the second modal; and

calculating the soft margin loss function according to the contrastive loss function in the first modal and the contrastive loss function in the second modal.

2 . The method according to claim 1 , wherein the cross-modal sample pair corresponds to at least one contrastive sample set, and the determining a total loss function based on the soft margin loss function comprises:

calculating a loss function of the corresponding contrastive sample set based on the soft margin loss function; and

calculating the total loss function based on the loss function of each corresponding sample set of the at least one contrastive sample set.

3 . The method according to claim 2 , wherein the soft margin loss function comprises a soft margin loss function for at least one state, and the calculating a loss function of the corresponding contrastive sample set based on the soft margin loss function comprises:

performing weighted summation on the soft margin loss function for the at least one state to obtain a weighted summation function; and

adding the weighted summation function and a hard margin loss function, and calculating the loss function of the corresponding contrastive sample set based on the added function.

4 . The method according to claim 1 , wherein the similarity distance comprises a static similarity distance and a dynamic similarity distance, the soft margin loss function comprises a static soft margin loss function and a dynamic soft margin loss function, the static soft margin loss function is calculated based on the static similarity distance, the dynamic soft margin loss function is calculated based on the dynamic similarity distance, and the calculating a distance between similarity of the positive sample pair and similarity of the negative sample pair to obtain a similarity distance comprises:

calculating the distance between the similarity of the positive sample pair and the similarity of the negative sample pair using a pre-training model, so as to obtain the static similarity distance; and/or

calculating the distance between the similarity of the positive sample pair and the similarity of the negative sample pair using current parameters of the cross-modal retrieval model to obtain the dynamic similarity distance.

5 . The method according to claim 1 , wherein the cross-modal sample pair corresponds to at least one contrastive sample set, and the determining a total loss function based on the soft margin loss function comprises:

calculating a loss function of the corresponding contrastive sample set based on the soft margin loss function; and

calculating the total loss function based on the loss function of each corresponding sample set of the at least one contrastive sample set.

6 . An electronic device, comprising:

at least one processor; and

a memory communicatively connected with the at least one processor;

wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform a method for training a cross-modal retrieval model which is used by a cross-modal retrieval system, wherein a cross-modal retrieval is used for a retrieval of data of one modal using data of another modal, wherein the method comprises:

determining a similarity of a cross-modal sample pair according to the cross-modal sample pair, the cross-modal sample pair comprising a sample of a first modal and a sample of a second modal, and the first modal being different from the second modal, wherein the first modal is a text, and the second modal is a video, and the cross-modal sample pair comprises a positive sample pair and a negative sample pair, the positive sample pair comprises an anchor sample and a positive sample, the negative sample pair comprises the anchor sample and a negative sample, the anchor sample has the first modal, and the positive sample and the negative sample have the second modal, and wherein the anchor sample is a text in a sample set, the positive sample is a video related to the text in the sample set, and the negative sample is a randomly selected video which is related or not related to the text in the sample set;

determining a soft margin based on the similarity, and determining a soft margin loss function based on the soft margin, wherein the soft margin is a non-fixed value; and

determining a total loss function based on the soft margin loss function, and training a cross-modal retrieval model according to the total loss function,

using the trained cross-modal retrieval model to receive a text input by a user, determine a video matched with the text using the cross-modal retrieval model, and feed the matched video back to the user,

wherein the determining a soft margin based on the similarity comprises:

calculating a distance between similarity of the positive sample pair and similarity of the negative sample pair to obtain a similarity distance; and

normalizing the similarity distance to obtain a normalized similarity distance, and determining the normalized similarity distance as the soft margin, and

wherein the similarity distance comprises a similarity distance in the first modal and a similarity distance in the second modal, the similarity distance in the first modal being obtained by processing the sample pair in the first modal using a semantic representation model in the first modal and the similarity distance in the second modal being obtained by processing the sample pair in the second modal using a semantic representation model in the second modal;

the determining a soft margin based on the similarity and determining a soft margin loss function based on the soft margin comprises:

determining a soft margin in the first modal based on the similarity distance in the first modal, and calculating a contrastive loss function in the first modal based on the soft margin in the first modal;

determining a soft margin in the second modal based on the similarity distance in the second modal, and calculating a contrastive loss function in the second modal based on the soft margin in the second modal; and

calculating the soft margin loss function according to the contrastive loss function in the first modal and the contrastive loss function in the second modal.

7 . The electronic device according to claim 6 , wherein the cross-modal sample pair corresponds to at least one contrastive sample set, and the determining a total loss function based on the soft margin loss function comprises:

calculating a loss function of the corresponding contrastive sample set based on the soft margin loss function; and

calculating the total loss function based on the loss function of each corresponding sample set of the at least one contrastive sample set.

8 . The electronic device according to claim 7 , wherein the soft margin loss function comprises a soft margin loss function for at least one state, and the calculating a loss function of the corresponding contrastive sample set based on the soft margin loss function comprises:

performing weighted summation on the soft margin loss function for the at least one state to obtain a weighted summation function; and

adding the weighted summation function and a hard margin loss function, and calculating the loss function of the corresponding contrastive sample set based on the added function.

9 . The electronic device according to claim 6 , wherein the similarity distance comprises a static similarity distance and a dynamic similarity distance, the soft margin loss function comprises a static soft margin loss function and a dynamic soft margin loss function, the static soft margin loss function is calculated based on the static similarity distance, the dynamic soft margin loss function is calculated based on the dynamic similarity distance, and the calculating a distance between similarity of the positive sample pair and similarity of the negative sample pair to obtain a similarity distance comprises:

calculating the distance between the similarity of the positive sample pair and the similarity of the negative sample pair using a pre-training model, so as to obtain the static similarity distance; and/or

calculating the distance between the similarity of the positive sample pair and the similarity of the negative sample pair using current parameters of the cross-modal retrieval model to obtain the dynamic similarity distance.

10 . A non-transitory computer readable storage medium with computer instructions stored thereon, wherein the computer instructions are used for causing a computer to perform a method for training a cross-modal retrieval model which is used by a cross-modal retrieval system, wherein a cross-modal retrieval is used for a retrieval of data of one modal using data of another modal, wherein the method comprises:

determining a similarity of a cross-modal sample pair according to the cross-modal sample pair, the cross-modal sample pair comprising a sample of a first modal and a sample of a second modal, and the first modal being different from the second modal, wherein the first modal is a text, and the second modal is a video, and the cross-modal sample pair comprises a positive sample pair and a negative sample pair, the positive sample pair comprises an anchor sample and a positive sample, the negative sample pair comprises the anchor sample and a negative sample, the anchor sample has the first modal, and the positive sample and the negative sample have the second modal, and wherein the anchor sample is a text in a sample set, the positive sample is a video related to the text in the sample set, and the negative sample is a randomly selected video which is related or not related to the text in the sample set;

determining a soft margin based on the similarity, and determining a soft margin loss function based on the soft margin, wherein the soft margin is a non-fixed value; and

determining a total loss function based on the soft margin loss function, and training a cross-modal retrieval model according to the total loss function,

using the trained cross-modal retrieval model to receive a text input by a user, determine a video matched with the text using the cross-modal retrieval model, and feed the matched video back to the user,

wherein the determining a soft margin based on the similarity comprises:

calculating a distance between similarity of the positive sample pair and similarity of the negative sample pair to obtain a similarity distance; and

normalizing the similarity distance to obtain a normalized similarity distance, and determining the normalized similarity distance as the soft margin, and

wherein the similarity distance comprises a similarity distance in the first modal and a similarity distance in the second modal, the similarity distance in the first modal being obtained by processing the sample pair in the first modal using a semantic representation model in the first modal and the similarity distance in the second modal being obtained by processing the sample pair in the second modal using a semantic representation model in the second modal;

the determining a soft margin based on the similarity and determining a soft margin loss function based on the soft margin comprises:

determining a soft margin in the first modal based on the similarity distance in the first modal, and calculating a contrastive loss function in the first modal based on the soft margin in the first modal;

determining a soft margin in the second modal based on the similarity distance in the second modal, and calculating a contrastive loss function in the second modal based on the soft margin in the second modal; and

calculating the soft margin loss function according to the contrastive loss function in the first modal and the contrastive loss function in the second modal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2021
From: HE, FENG; WANG, QI; FENG, ZHIFAN; YANG, HU; CHAI, CHUNGUANG
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 057807/0423 →
Priority Claims (1)
CN 202110244645.X · Mar 5, 2021 · national
Continuity (1)
Related Publication 20220284246A1 · Sep 8, 2022
References Cited (29)
US 11308353B2 · Singh · 2022 [cited by examiner]
US 11586927B2 · Li · 2023 [cited by examiner]
US 20200250537A1 · Li et al. · 2020 [cited by applicant]
US 20200302340A1 · Durand et al. · 2020 [cited by applicant]
US 20220012530A1 · Singh · 2022 [cited by examiner]
CN 108182279A · 2018 [cited by applicant]
CN 109522850A · 2019 [cited by applicant]
CN 111091010A · 2020 [cited by applicant]
CN 111325223A · 2020 [cited by applicant]
CN 111507218A · 2020 [cited by applicant]
CN 111862175A · 2020 [cited by applicant]
CN 112119411A · 2020 [cited by applicant]
CN 112148916A · 2020 [cited by applicant]
JP 2020177465A · 2020 [cited by applicant]
Yang, Zhao, et al. “A novel soft margin loss function for deep discriminative embedding learning.” IEEE Access 8 (2020): 202785-202794. (Year: 2020). [cited by examiner]
Zhang, Linguang, and Szymon Rusinkiewicz. “Learning local descriptors with a CDF-based dynamic soft margin.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019. (Year: 2019). [cited by examiner]
Wang, Bokun, et al. “Adversarial cross-modal retrieval.” Proceedings of the 25th ACM international conference on Multimedia. 2017. (Year: 2017). [cited by examiner]
Hermans, Alexander, Lucas Beyer, and Bastian Leibe. “In defense of the triplet loss for person re-identification.” arXiv preprint arXiv:1703.07737 (2017). (Year: 2017). [cited by examiner]
Deng, Cheng, et al. “Triplet-based deep hashing network for cross-modal retrieval.” IEEE Transactions on Image Processing 27.8 (2018): 3893-3903. (Year: 2018). [cited by examiner]
Semedo, David, and João Magalhães. “Cross-modal subspace learning with scheduled adaptive margin constraints.” Proceedings of the 27th ACM International Conference on Multimedia. 2019. (Year: 2019). [cited by examiner]
He, Feng, et al. “Improving video retrieval by adaptive margin.” Proceedings of the 44th international ACM SIGIR conference on research and development in information retrieval. 2021. (Year: 2021). [cited by examiner]
Yasuda et al. “Cross-Mode Speech Search based on Specific Correlation of Weak Tags”, Acoustical Society of Japan 2020 Fall Conference Lecture Proceedings CD-ROM [CD-ROM], Japan, Acoustical Society of Japan, Aug. 26, 202… [cited by applicant]
Wei et al., Universal Weighting Metric Learning for Cross-Modal Matching, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13002-13011. [cited by applicant]
Semedo, et al., Adaptive Temporal Triplet-loss for Cross-modal Embedding Learning, MM '20: Proceedings of the 28th ACM International Conference on Multimedia, ACM, Oct. 12, 2020, pp. 1152-1161, Internet <URL:https://doi… [cited by applicant]
Extended European Search Report of European application No. 21201915.2 dated Mar. 11, 2022, 11 pages. [cited by applicant]
Semedo et al., “Adaptive Temporal Triplet-loss for Cross-modal Embedding Learning”, Proceedings of the 28th ACM International Conference on Multimedia, ACMPUB27, New York, NY, USA, Oct. 12, 2020, pp. 1152-1161, XP058478… [cited by applicant]
Aggarwal et al., “Text-based Person Search via Attribute-aided Matching”, 2020 IEEE Winter Conference on Applications of Computer Vision (WACV), IEEE, Mar. 1, 2020, pp. 2606-2614, XP033771285. [cited by applicant]
Wu et al., “Online Fast Adaptive Low-Rank Similarity Learning for Cross-Modal Retrieval”, IEEE Transactions on Multimedia, IEEE, USA, vol. 22, No. 5, Sep. 20, 2019, pp. 1310-13322, XP011784986. [cited by applicant]
Liu, Triplet Loss and Manifold Dimensionally Reduction Based Method for Text-Independent Speaker Recognition, Harbin Institute of Technology, School of Computer Science and Technology, Jan. 17, 2020, 75 pages. [cited by applicant]