IP Library › Granted Patent US 12,327,084
Granted Patent B2
US 12,327,084 · App. 17/954,767 · Granted Jun 10, 2025

Video question answering method, electronic device and storage medium

Inventors: Bohao Feng (Beijing, CN); Yuxin Liu (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
G06F40/279G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,327,084
App. No.
17/954,767
Granted
Jun 10, 2025
Kind
B2
Abstract

There is provided a video question answering method and apparatus, an electronic device and a storage medium, which relates to the field of artificial intelligence, such as natural language processing technologies, deep learning technologies, voice recognition technologies, knowledge graph technologies, computer vision technologies, or the like. The method includes: determining M key frames for a video corresponding to a to-be-answered question, M being a positive integer greater than 1 and less than or equal to a number of video frames in the video; and determining an answer corresponding to the question according to the M key frames.

Claims (47)

1. A computer-implemented method for video question answering, comprising:

determining M key frames for a video corresponding to a to-be-answered question, M being a positive integer greater than 1 and less than or equal to a number of video frames in the video, including: for any video frame in the video, extracting object information in the video frame respectively, and generating a sentence describing the objects in the video frame using conjunctions in a corpus and the object information; in response to the generated sentence conforming to a grammatical rule, outputting the generated sentence as description information of the video frame, and otherwise, updating the conjunctions until a sentence conforming to the grammatical rule is generated and outputting the generated sentence as description information of the video frame; acquiring scores of correlation between the description information of respective video frames and the question; and sorting the video frames in the descending order of the scores of correlation, and using the first M video frames after the sort as the key frames; and

acquiring vector representations of the key frames; and determining the answer corresponding to the question according to the vector representations of the key frames, including: determining a question type of the question, if the question is an intuitive question, determining the corresponding answer using the vector representation of each key frame and the question, wherein the intuitive question is a question which can be answered directly using the information in the video; and if the question is a non-intuitive question, determining the corresponding answer using the vector representation of each key frame, the question and a corresponding knowledge graph, wherein the non-intuitive question is a question which may not be answered directly using the information in the video, and the knowledge graph comprises a generic knowledge graph and/or a special knowledge graph constructed from the video,

the method further comprising:

for any key frame, acquiring a product of the vector representation of a previous key frame and a predetermined fourth weight parameter and a product of the vector representation of the key frame and a predetermined fifth weight parameter; adding the obtained two products and a predetermined second bias vector, and performing a hyperbolic tangent neural network activation function operation on the added sum; and multiplying an operation result by a predetermined sixth weight parameter, and acquiring the timing attention weight corresponding to the key frame based on the obtained product, the previous key frame being located before the key frame and closest to the key frame in time,

wherein the determining the answer corresponding to the question according to the vector representations of the key frames comprises: determining the answer corresponding to the question according to the updated vector representations of the key frames.

2. The method according to claim 1 , wherein the acquiring vector representations of the key frames comprises:

performing the following processing operations on any key frame:

extracting a target region of the key frame;

extracting features of the key frame to obtain a feature vector corresponding to the key frame;

extracting features of any extracted target region to obtain a feature vector corresponding to the target region, and generating a vector representation of the target region according to the feature vector corresponding to the target region; and

generating the vector representation of the key frame according to the vector representation of each extracted target region and the feature vector corresponding to the key frame.

3. The method according to claim 2 , wherein the generating a vector representation of the target region according to the feature vector corresponding to the target region comprises:

splicing the feature vector corresponding to the target region and at least one of the following vectors: a feature vector corresponding to the key frame where the target region is located, a text vector corresponding to the key frame where the target region is located, and an audio vector corresponding to the video; and acquiring the vector representation of the target region based on a splicing result;

the text vector is obtained after text information extracted from the key frame where the target region is located is transformed to a vector, and the audio vector is obtained after text information corresponding to an audio of the video is transformed to a vector.

4. The method according to claim 3 , wherein the acquiring the vector representation of the target region based on a splicing result comprises:

acquiring a spatial attention weight corresponding to the target region, multiplying the spatial attention weight by the splicing result corresponding to the target region, and using a multiplying result as the vector representation of the target region.

5. The method according to claim 4 , wherein the acquiring a spatial attention weight corresponding to the target region comprises:

determining the spatial attention weight corresponding to the target region according to a question vector and the splicing result corresponding to the target region, the question vector being obtained after text information corresponding to the question is transformed to a vector.

6. The method according to claim 5 , wherein the determining the spatial attention weight corresponding to the target region according to a question vector and the splicing result corresponding to the target region comprises:

acquiring a product of the question vector and a predetermined first weight parameter and a product of the splicing result and a predetermined second weight parameter;

adding the obtained two products and a predetermined first bias vector, and performing a hyperbolic tangent neural network activation function operation on an added sum; and

multiplying an operation result by a predetermined third weight parameter, and acquiring the spatial attention weight corresponding to the target region based on the obtained product.

7. The method according to claim 2 , wherein the generating the vector representation of the key frame according to the vector representation of each extracted target region and the feature vector corresponding to the key frame comprises:

splicing the vector representation of each target region and the feature vector corresponding to the key frame, and using a splicing result as the vector representation of the key frame.

8. An electronic device, comprising:

at least one processor; and

a memory communicatively connected with the at least one processor;

wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform a method for video question answering, wherein the method comprises:

determining M key frames for a video corresponding to a to-be-answered question, M being a positive integer greater than 1 and less than or equal to a number of video frames in the video, including: for any video frame in the video, extracting object information in the video frame respectively, and generating a sentence describing the objects in the video frame using conjunctions in a corpus and the object information; in response to the generated sentence conforming to a grammatical rule, outputting the generated sentence as description information of the video frame, and otherwise, updating the conjunctions until a sentence conforming to the grammatical rule is generated and outputting the generated sentence as description information of the video frame; acquiring scores of correlation between the description information of respective video frames and the question; and sorting the video frames in the descending order of the scores of correlation, and using the first M video frames after the sort as the key frames; and

acquiring vector representations of the key frames; and determining the answer corresponding to the question according to the vector representations of the key frames, including: determining a question type of the question, if the question is an intuitive question, determining the corresponding answer using the vector representation of each key frame and the question, wherein the intuitive question is a question which can be answered directly using the information in the video; and if the question is a non-intuitive question, determining the corresponding answer using the vector representation of each key frame, the question and a corresponding knowledge graph, wherein the non-intuitive question is a question which may not be answered directly using the information in the video, and the knowledge graph comprises a generic knowledge graph and/or a special knowledge graph constructed from the video,

the method further comprising:

for any key frame, acquiring a product of the vector representation of a previous key frame and a predetermined fourth weight parameter and a product of the vector representation of the key frame and a predetermined fifth weight parameter; adding the obtained two products and a predetermined second bias vector, and performing a hyperbolic tangent neural network activation function operation on the added sum; and multiplying an operation result by a predetermined sixth weight parameter, and acquiring the timing attention weight corresponding to the key frame based on the obtained product, the previous key frame being located before the key frame and closest to the key frame in time,

wherein the determining the answer corresponding to the question according to the vector representations of the key frames comprises: determining the answer corresponding to the question according to the updated vector representations of the key frames.

9. The electronic device according to claim 8 ,

wherein the acquiring vector representations of the key frames comprises: performing the following processing operations on any key frame: extracting a target region of the key frame; extracting features of the key frame to obtain a feature vector corresponding to the key frame; extracting features of any extracted target region to obtain a feature vector corresponding to the target region, and generating a vector representation of the target region according to the feature vector corresponding to the target region; and generating the vector representation of the key frame according to the vector representation of each extracted target region and the feature vector corresponding to the key frame.

10. The electronic device according to claim 9 ,

wherein the generating a vector representation of the target region according to the feature vector corresponding to the target region comprises: splicing the feature vector corresponding to the target region and at least one of the following vectors: a feature vector corresponding to the key frame where the target region is located, a text vector corresponding to the key frame where the target region is located, and an audio vector corresponding to the video; and acquire the vector representation of the target region based on a splicing result;

the text vector is obtained after text information extracted from the key frame where the target region is located is transformed to a vector, and the audio vector is obtained after text information corresponding to an audio of the video is transformed to a vector.

11. The electronic device according to claim 10 ,

wherein the acquiring the vector representation of the target region based on a splicing result comprises: acquiring a spatial attention weight corresponding to the target region, multiplying the spatial attention weight by the splicing result corresponding to the target region, and using a multiplying result as the vector representation of the target region.

12. A non-transitory computer readable storage medium with computer instructions stored thereon, wherein the computer instructions are used for causing a method for video question answering, wherein the method comprises:

determining M key frames for a video corresponding to a to-be-answered question, M being a positive integer greater than 1 and less than or equal to a number of video frames in the video, including: for any video frame in the video, extracting object information in the video frame respectively, and generating a sentence describing the objects in the video frame using conjunctions in a corpus and the object information; in response to the generated sentence conforming to a grammatical rule, outputting the generated sentence as description information of the video frame, and otherwise, updating the conjunctions until a sentence conforming to the grammatical rule is generated and outputting the generated sentence as description information of the video frame; acquiring scores of correlation between the description information of respective video frames and the question; and sorting the video frames in the descending order of the scores of correlation, and using the first M video frames after the sort as the key frames; and

acquiring vector representations of the key frames; and determining the answer corresponding to the question according to the vector representations of the key frames, including: determining a question type of the question, if the question is an intuitive question, determining the corresponding answer using the vector representation of each key frame and the question, wherein the intuitive question is a question which can be answered directly using the information in the video; and if the question is a non-intuitive question, determining the corresponding answer using the vector representation of each key frame, the question and a corresponding knowledge graph, wherein the non-intuitive question is a question which may not be answered directly using the information in the video, and the knowledge graph comprises a generic knowledge graph and/or a special knowledge graph constructed from the video,

the method further comprising:

for any key frame, acquiring a product of the vector representation of a previous key frame and a predetermined fourth weight parameter and a product of the vector representation of the key frame and a predetermined fifth weight parameter; adding the obtained two products and a predetermined second bias vector, and performing a hyperbolic tangent neural network activation function operation on the added sum; and multiplying an operation result by a predetermined sixth weight parameter, and acquiring the timing attention weight corresponding to the key frame based on the obtained product, the previous key frame being located before the key frame and closest to the key frame in time,

wherein the determining the answer corresponding to the question according to the vector representations of the key frames comprises: determining the answer corresponding to the question according to the updated vector representations of the key frames.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2022
From: FENG, BOHAO; LIU, YUXIN
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 061244/0087 →
Priority Claims (1)
CN 202111196420.8 · Oct 14, 2021 · national
Continuity (1)
Related Publication 20230121838A1 · Apr 20, 2023
References Cited (43)
US 11244167B2 · Zhao · 2022 [cited by examiner]
US 12067759B2 · Zhang · 2024 [cited by examiner]
US 20180084023A1 · Stoop · 2018 [cited by examiner]
US 20180144208A1 · Lu · 2018 [cited by examiner]
US 20180285456A1 · Nichkawde · 2018 [cited by examiner]
US 20190266409A1 · He et al. · 2019 [cited by applicant]
US 20200139973A1 · Palanisamy · 2020 [cited by examiner]
US 20210012222A1 · Kim · 2021 [cited by examiner]
US 20210103615A1 · Jindal · 2021 [cited by examiner]
US 20210248375A1 · Geng · 2021 [cited by examiner]
US 20210248376A1 · Zhao · 2021 [cited by examiner]
US 20220230628A1 · Zhu · 2022 [cited by examiner]
US 20220350826A1 · Zhang · 2022 [cited by examiner]
US 20230017614A1 · Srinivasa · 2023 [cited by examiner]
US 20230027713A1 · Wu · 2023 [cited by examiner]
US 20230121838A1 · Feng · 2023 [cited by examiner]
US 20240037896A1 · Zhang · 2024 [cited by examiner]
US 20250022458A1 · Audhkhasi · 2025 [cited by examiner]
CN 107463609A · 2017 [cited by applicant]
CN 107832724A · 2018 [cited by applicant]
CN 109271506A · 2019 [cited by applicant]
CN 109508642A · 2019 [cited by applicant]
CN 109587581A · 2019 [cited by applicant]
CN 109784280A · 2019 [cited by applicant]
CN 109829049A · 2019 [cited by applicant]
CN 110072142A · 2019 [cited by applicant]
CN 110399457A · 2019 [cited by applicant]
CN 110781347A · 2020 [cited by applicant]
CN 111310655A · 2020 [cited by applicant]
CN 111368656A · 2020 [cited by applicant]
CN 112131364A · 2020 [cited by applicant]
CN 112288142A · 2021 [cited by applicant]
CN 113392288A · 2021 [cited by applicant]
WO 2021045434A1 · 2021 [cited by applicant]
Yao, Li, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Pal, Hugo Larochelle, and Aaron Courville. “Describing videos by exploiting temporal structure.” In Proceedings of the IEEE international conference on … [cited by examiner]
M. Heilman and N. A. Smith, “Question generation via overgenerating transformations and ranking,” Lang. Technol. Inst., School Comput. Sci., Carnegie Mellon Univ., Pittsburgh, PA, USA, Tech. Rep., 2009. (Year: 2009). [cited by examiner]
Zhao, Zhou, Qifan Yang, Deng Cai, Xiaofei He, Yueting Zhuang, Zhou Zhao, Qifan Yang, Deng Cai, Xiaofei He, and Yueting Zhuang. “Video Question Answering via Hierarchical Spatio-Temporal Attention Networks.” In IJCAI, vo… [cited by examiner]
J. Li, X. Liu, W. Zhang, M. Zhang, J. Song and N. Sebe, “Spatio-Temporal Attention Networks for Action Recognition and Detection,” in IEEE Transactions on Multimedia, vol. 22, No. 11, pp. 2990-3001, Nov. 2020, doi: 10.1… [cited by examiner]
Zhang, et al., Frame Augmented Alternating Attention Network for Video Question Answering, IEEE Transactions on Multimedia, vol. X, No. X, X, X, 1-9, 2019. [cited by applicant]
Extended European Search Report of European patent application No. 22198021.2 dated Feb. 28, 2023, 8 pages. [cited by applicant]
Wang Weining et al., “Long video question answering: A Matching-guided Attention Model”, Pattern Recognition, Elsevier, GB, vol. 102, Jan. 30, 2020 (Jan. 30, 2020), pp. 1-11, XP086066543, ISSN: 0031-3203, DOI: 10.1016/J… [cited by applicant]
Zhou Zhao et al., “Video Question Answering via Hierarchical Spatio-Temporal Attention Networks”, Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, Aug. 2017 (Aug. 2017), pp. 351… [cited by applicant]
Yu Ting et al., “Long-Term Video Question Answering via Multimodal Hierarchical Memory Attentive Networks”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, No. 3, May 20, 2020 (May 20, 2020), pp… [cited by applicant]