IP Library › Granted Patent US 12,386,891
Granted Patent B2
US 12,386,891 · App. 17/932,598 · Granted Aug 12, 2025

Information search method and device, electronic device, and storage medium

Inventors: Wenbin Jiang (Beijing, CN); Yajuan Lyu (Beijing, CN); Yong Zhu (Beijing, CN); Hua Wu (Beijing, CN); Haifeng Wang (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
G06F16/735
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,386,891
App. No.
17/932,598
Granted
Aug 12, 2025
Kind
B2
Abstract

An information search method includes: obtaining search words at least including a question to be searched and obtaining an initial text vector representation of the search words; obtaining a video corresponding to the search words, and obtaining multi-modality vector representations of the video; starting from the initial text vector representation, performing N rounds of interaction between the video and the search words based on the multi-modality vector representations and a text vector representation of the search words of a current round, to generate a target fusion vector representation, where N is an integer greater than or equal to 1; and obtaining target video frames matching the question to be searched by annotating the video based on the target fusion vector representation.

Claims (75)

1. An information search method, comprising:

obtaining search words at least comprising a question to be searched, and obtaining an initial text vector representation of the search words;

obtaining a video corresponding to the search words, and obtaining multi-modality vector representations of the video;

starting from the initial text vector representation, generating a target fusion vector representation by performing N rounds of interaction between the video and the search words based on the multi-modality vector representations and a text vector representation of the search term of an i-th round of interaction, where N is an integer greater than or equal to 1 and i is an integer greater than or equal to 1 and less than or equal to N; and

obtaining target video frames matching the question to be searched by annotating the video based on the target fusion vector representation;

wherein the search words further comprise candidate answers, and obtaining the initial text vector representation of the search words comprises:

obtaining a first character string corresponding to the question to be searched and obtaining a second character string corresponding to the candidate answers;

obtaining a target character string by splicing the first character string and the second character string; and

obtaining the initial text vector representation by performing a word embedding processing on the target character string;

wherein the method further comprises:

obtaining a text vector representation obtained after an N-th round of interaction; and

obtaining text annotation results of the search words and text annotation results of the candidate answers by annotating the text vector representation.

2. The method of claim 1 , wherein obtaining the multi-modality vector representation of the video comprises:

obtaining at least two feature extraction dimensions; and

obtaining the multi-modality vector representations by performing feature extraction on each image frame of the video based on all of the at least two feature extraction dimensions.

3. The method of claim 1 , wherein generating the target fusion vector representation comprises:

for the i-th round of interaction, obtaining a mono-modality fusion vector representation corresponding to each modality separately by fusing modality vector representations corresponding to respective modalities contained in the multi-modality vector representations based on a text vector representation obtained after an (i−1)-th round of interaction; and

generating a fusion vector representation corresponding to the i-th round of interaction by fusing mono-modality fusion vector representations of all modalities based on the text vector representations obtained after the (i−1)-th round of interaction.

4. The method of claim 3 , further comprising:

obtaining a text vector representation corresponding to the i-th round by fusing the text vector representation obtained after the (i−1)-th round of interaction based on the fusion vector representation of the i-th round.

5. The method of claim 3 , wherein obtaining the mono-modality fusion vector representation corresponding to each modality comprises: for each modality,

obtaining a first weight matrix corresponding to the modality based on the modality vector representations corresponding to respective modalities and the text vector representation obtained after the (i−1)-th round of interaction; and

performing a weighted sum on the modality vector representation corresponding to the modality based on the corresponding first weight matrix and determining a weighted sum result as the mono-modality fusion vector representation of the modality.

6. The method of claim 5 , wherein obtaining the fusion vector representation corresponding to the i-th round of interaction comprises:

obtaining a target parameter matrix, and obtaining a second weight matrix based on the target parameter matrix, the mono-modality fusion vector representations corresponding to respective modalities, and the text vector representation obtained after the (i−1)-th round of interaction; and

performing a weighted sum on the mono-modality fusion vector representations corresponding to all modalities based on the second weight matrix, and determining a weighted sum result as the fusion vector representation corresponding to the i-th round of interaction.

7. The method of claim 6 , wherein obtaining the text vector representation of the i-th round comprises:

obtaining a third weight matrix based on the fusion vector representation corresponding to the i-th round of interaction, the target parameter matrix, and the text vector representation obtained after the (i−1)-th round of interaction; and

performing a weighted sum on the text vector representation obtained after the (i−1)-th round of interaction and the fusion vector representation corresponding to the i-th round of interaction based on the third weight matrix, and determining a weighted sum result as the text vector representation corresponding to the i-th round.

8. The method of claim 2 , wherein the feature extraction dimensions comprise at least one of a video dimension, a text dimension and a sound dimension.

9. An electronic device, comprising:

at least one processor; and

a memory, communicatively coupled to the at least one processor;

wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is configured to:

obtain search words at least comprising a question to be searched, and obtain an initial text vector representation of the search words;

obtain a video corresponding to the search words, and obtain multi-modality vector representations of the video;

starting from the initial text vector representation, generate a target fusion vector representation by performing N rounds of interaction between the video and the search words based on the multi-modality vector representations and a text vector representation of the search term of an i-th round, where N is an integer greater than or equal to 1 and i is an integer greater than or equal to 1 and less than or equal to N; and

obtain target video frames matching the question to be searched by annotating the video based on the target fusion vector representation;

wherein the search words further comprise candidate answers, and the at least one processor is configured to:

obtain a first character string corresponding to the question to be searched and obtain a second character string corresponding to the candidate answers;

obtain a target character string by splicing the first character string and the second character string; and

obtain the initial text vector representation by performing a word embedding processing on the target character string;

wherein the at least one processor is configured to:

obtain a text vector representation obtained after an N-th round of interaction; and

obtain text annotation results of the search words and text annotation results of the candidate answers by annotating the text vector representation.

10. The electronic device of claim 9 , wherein the at least one processor is configured to:

obtain at least two feature extraction dimensions; and

obtain the multi-modality vector representations by performing feature extraction on each image frame of the video based on all of the at least two feature extraction dimensions.

11. The electronic device of claim 9 , wherein the at least one processor is configured to:

for the i-th round of interaction, obtain a mono-modality fusion vector representation corresponding to the modality separately by fusing modality vector representation corresponding to the modality contained in the multi-modality vector representations based on a text vector representation obtained after an (i−1)-th round of interaction, where i is an integer greater than 1; and

generate a fusion vector representation corresponding to the i-th round of interaction by fusing mono-modality fusion vector representations of all modalities based on the text vector representations obtained after the (i−1)-th round of interaction.

12. The electronic device of claim 11 , wherein the at least one processor is configured to:

obtain a text vector representation corresponding to the i-th round by fusing the text vector representation obtained after the (i−1)-th round of interaction based on the fusion vector representation of the i-th round.

13. The electronic device of claim 11 , wherein the at least one processor is configured to:

for each modality,

obtain a first weight matrix corresponding to the modality based on the modality vector representation corresponding to the modality and the text vector representations obtained after the (i−1)-th round of interaction; and

perform a weighted sum on the modality vector representation corresponding to the modality based on the first weight matrix and determine a weighted sum result as the mono-modality fusion vector representation of the modality.

14. The electronic device of claim 13 , wherein the at least one processor is configured to:

obtain a target parameter matrix, and obtain a second weight matrix based on the target parameter matrix, the mono-modality fusion vector representations corresponding to respective modalities, and the text vector representation obtained after the (i−1)-th round of interaction; and

perform a weighted sum on the mono-modality fusion vector representations corresponding to all modalities based on the second weight matrix, and determine a weighted sum result as the fusion vector representation corresponding to the i-th round of interaction.

15. The electronic device of claim 14 , wherein the at least one processor is configured to:

obtain a third weight matrix based on the fusion vector representation corresponding to the i-th round of interaction, the target parameter matrix, and the text vector representation obtained after the (i−1)-th round of interaction; and

perform a weighted sum on the text vector representation obtained after the (i−1)-th round of interaction and the fusion vector representation corresponding to the i-th round of interaction based on the third weight matrix, and determining a weighted sum result as the text vector representation corresponding to the i-th round.

16. A non-transitory computer readable storage medium, having computer instructions stored thereon, wherein the computer instructions are configured to cause a computer to perform an information search method, the method comprising:

obtaining search words at least comprising a question to be searched, and obtaining an initial text vector representation of the search words;

obtaining a video corresponding to the search words, and obtaining multi-modality vector representations of the video;

starting from the initial text vector representation, generating a target fusion vector representation by performing N rounds of interaction between the video and the search words based on the multi-modality vector representations and a text vector representation of the search term of an i-th round, where N is an integer greater than or equal to 1 and i is an integer greater than or equal to 1 and less than or equal to N; and

obtaining target video frames matching the question to be searched by annotating the video based on the target fusion vector representation;

wherein the search words further comprise candidate answers, and obtaining the initial text vector representation of the search words comprises:

obtaining a first character string corresponding to the question to be searched and obtaining a second character string corresponding to the candidate answers;

obtaining a target character string by splicing the first character string and the second character string; and

obtaining the initial text vector representation by performing a word embedding processing on the target character string;

wherein the method further comprises:

obtaining a text vector representation obtained after an N-th round of interaction; and

obtaining text annotation results of the search words and text annotation results of the candidate answers by annotating the text vector representation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2022
From: JIANG, WENBIN; LYU, YAJUAN; ZHU, YONG; WU, HUA; WANG, HAIFENG
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 061182/0673 →
Priority Claims (1)
CN 202111101827.8 · Sep 18, 2021 · national
Continuity (1)
Related Publication 20230008897A1 · Jan 12, 2023
References Cited (26)
US 6970860B1 · Liu · 2005 [cited by examiner]
US 7263671B2 · Hull · 2007 [cited by examiner]
US 7315857B2 · Dettinger · 2008 [cited by examiner]
US 7446803B2 · Leow · 2008 [cited by examiner]
US 20040123231A1 · Adams, Jr. · 2004 [cited by examiner]
US 20040205482A1 · Basu · 2004 [cited by examiner]
US 20060195858A1 · Takahashi · 2006 [cited by examiner]
US 20090319883A1 · Mei · 2009 [cited by examiner]
US 20120278337A1 · Acharya · 2012 [cited by examiner]
US 20130060784A1 · Acharya · 2013 [cited by examiner]
US 20150296228A1 · Chen · 2015 [cited by examiner]
US 20200104318A1 · Ponjou Tasse · 2020 [cited by examiner]
US 20210200802A1 · Lyu · 2021 [cited by examiner]
US 20210248375A1 · Geng · 2021 [cited by examiner]
US 20220222920A1 · Huang · 2022 [cited by examiner]
CN 110225368A · 2019 [cited by applicant]
CN 110866184A · 2020 [cited by applicant]
CN 111382309A · 2020 [cited by applicant]
CN 112364204A · 2021 [cited by applicant]
CN 112668559A · 2021 [cited by applicant]
CN 113010740A · 2021 [cited by applicant]
CN 113392265A · 2021 [cited by applicant]
Tadas Baltrusaitis et al, “Multimodal Machine Learning: A Survey and Taxonomy”, IEEE, vol. 41, No. 2, Feb. 2019, pp. 423-443 (Year: 2019). [cited by examiner]
Ameen Ali et al, “Video and Text Matching with Conditioned Embeddings”, IEEE, pp. 478-487 (Year: 2022). [cited by examiner]
EPO, Extended European Search Report for EP Application No. 22196154.3, Feb. 8, 2023. [cited by applicant]
CNIPA, First Office Action for CN Application No. 202111101827.8, Jun. 16, 2023. [cited by applicant]
Cited By (1)
US 12,620,225