IP Library › Granted Patent US 12,737,414
Granted Patent B2
US 12,737,414 · App. 18/535,554 · Granted Sep 15, 2026

Converting video semantics into language for real-time query and information retrieval

Inventor: Dongeek Shin (San Jose, CA)
Assignee: GOOGLE LLC
G06F16/583G06F40/30G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,414
App. No.
18/535,554
Granted
Sep 15, 2026
Kind
B2
Abstract

Implementations utilize a LLM to generate content responsive to a user query directed to a video and cause audio data for the generated content to be rendered as a response to the user query. Implementations extract a subset of frames from all frames of the video as key frame(s) for the video, and utilize a vision-language model in generating a natural language description for the key frame(s) of the video. A prompt can be generated based on a transcription of the user query and based on the natural language description for the key frame(s) of the video. The prompt is processed as input, using the LLM, to generate the content responsive to the user query directed to the video.

Claims (27)

1 . A method implemented by one or more processors, the method comprising:

determining a plurality of key image frames from a video, wherein the plurality of key image frames includes a past key image frame that occurs earlier than a current image frame in the video and a future key image frame that occurs later than the current image frame in the video;

processing the plurality of key image frames, including the past and future key image frames, using a vision-language model,

wherein processing the plurality of key image frames using the vision-language model causes a natural language description of the past and future key image frames to be generated using the vision-language model;

storing the natural language description for the plurality of key image frames in association with the video;

receiving, from a computing device, a user query related to the video, wherein the user query is received when the current image frame of the video is being rendered; and

in response to receiving the user query,

generating a prompt based on the user query and based on the natural language description for the plurality of key image frames of the video,

processing the prompt as input using a generative model, to generate a generative model output, wherein the generative model output is operable to cause a response responsive to the user query to be rendered by an output device, and

providing the generative model output to the computing device.

2 . The method of claim 1 , wherein processing the plurality of key image frames using the vision-language model comprises, for each of multiple key image frames:

processing a respective key image frame as input using the vision-language model, to generate a respective model output from which a respective text is determined for the respective key image frame, and

assembling the natural language description based on a combination of the respective text for each of the multiple key image frames.

3 . The method of claim 1 , wherein determining the plurality of key image frames from the video comprises evaluating a plurality of frames of the video to select, as the plurality of key video frames that comprise less than all of the frames of the video, one or more of the plurality of frames that satisfy one or more criteria.

4 . The method of claim 3 , wherein the one or more criteria include a measure of visual difference between two adjacent frames of the plurality of video frames satisfying a threshold.

5 . The method of claim 3 , wherein the one or more criteria include a new object being detected in a frame of the plurality of frames of the video.

6 . The method of claim 3 , wherein the one or more criteria include a new voice being detected in an audio portion of the video that corresponds temporally with a frame of the plurality of frames of the video.

7 . A system comprising one or more processors and a memory storing instructions that, when executed on the one or more processors, cause the one or more processors to:

determine a plurality of key image frames from a video, wherein the plurality of key image frames includes a past key image frame that occurs earlier than a current image frame in the video and a future key image frame that occurs later than the current image frame in the video;

process the plurality of key image frames, including the past and future key image frames, using a vision-language model,

wherein processing the plurality of key image frames using the vision-language model causes a natural language description of the past and future key image frames to be generated using the vision-language model;

store the natural language description for the plurality of key image frames in association with the video;

receive, from a computing device, a user query related to the video, wherein the user query is received when the current image frame of the video is being rendered; and

in response to receiving the user query,

generate a prompt based on the user query and based on the natural language description for the plurality of key image frames of the video,

process the prompt as input using a generative model, to generate a generative model output, wherein the generative model output is operable to cause a response responsive to the user query to be rendered by an output device, and

provide the generative model output to the computing device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2023
From: SHIN, DONGEEK
To: GOOGLE LLC
Reel/Frame 065861/0811 →
Continuity (1)
Related Publication 20250190488A1 · Jun 12, 2025
References Cited (19)
US 9462230B1 · Agrawal · 2016 [cited by examiner]
US 20020097380A1 · Moulton · 2002 [cited by examiner]
US 20050068581A1 · Hull · 2005 [cited by examiner]
US 20130343721A1 · Abecassis · 2013 [cited by applicant]
US 20140316785A1 · Bennett et al. · 2014 [cited by applicant]
US 20180352195A1 · Segal · 2018 [cited by applicant]
US 20180359530A1 · Marlow et al. · 2018 [cited by applicant]
US 20210250404A1 · Xia · 2021 [cited by examiner]
US 20230306056A1 · Lee · 2023 [cited by examiner]
US 20240346102A1 · Miller · 2024 [cited by examiner]
US 20240394404A1 · O'Neal · 2024 [cited by examiner]
US 20250190488A1 · Shin · 2025 [cited by examiner]
US 20250232123A1 · Tow · 2025 [cited by examiner]
US 20250315283A1 · Volyn · 2025 [cited by examiner]
US 20250355905A1 · Brenner · 2025 [cited by examiner]
Zeng et al, Leveraging Video Descriptions to Learn Video Question Answering, Dec. 16, 2016 (Year: 2016). [cited by examiner]
Khurana et al., “Video Question-Answering Techniques, Benchmark Datasets and Evaluation Metrics Leveraging Video Captioning: A Comprehensive Survey” DOI: 10.1109/ACCESS.2017, 25 pages, vol. 4, dated 2016. [cited by applicant]
Zeng, A. et al., “Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language”; arXiv.org, Cornell University; arXiv:2204.00598; 20 pages; dated Apr. 1, 2022. [cited by applicant]
European Patent Office; International Search Report and Written Opinion issued in Application No. PCT/US2024/054390; 10 pages; dated Feb. 27, 2025. [cited by applicant]