Converting video semantics into language for real-time query and information retrieval
Implementations utilize a LLM to generate content responsive to a user query directed to a video and cause audio data for the generated content to be rendered as a response to the user query. Implementations extract a subset of frames from all frames of the video as key frame(s) for the video, and utilize a vision-language model in generating a natural language description for the key frame(s) of the video. A prompt can be generated based on a transcription of the user query and based on the natural language description for the key frame(s) of the video. The prompt is processed as input, using the LLM, to generate the content responsive to the user query directed to the video.
1 . A method implemented by one or more processors, the method comprising:
determining a plurality of key image frames from a video, wherein the plurality of key image frames includes a past key image frame that occurs earlier than a current image frame in the video and a future key image frame that occurs later than the current image frame in the video;
processing the plurality of key image frames, including the past and future key image frames, using a vision-language model,
wherein processing the plurality of key image frames using the vision-language model causes a natural language description of the past and future key image frames to be generated using the vision-language model;
storing the natural language description for the plurality of key image frames in association with the video;
receiving, from a computing device, a user query related to the video, wherein the user query is received when the current image frame of the video is being rendered; and
in response to receiving the user query,
generating a prompt based on the user query and based on the natural language description for the plurality of key image frames of the video,
processing the prompt as input using a generative model, to generate a generative model output, wherein the generative model output is operable to cause a response responsive to the user query to be rendered by an output device, and
providing the generative model output to the computing device.
2 . The method of claim 1 , wherein processing the plurality of key image frames using the vision-language model comprises, for each of multiple key image frames:
processing a respective key image frame as input using the vision-language model, to generate a respective model output from which a respective text is determined for the respective key image frame, and
assembling the natural language description based on a combination of the respective text for each of the multiple key image frames.
3 . The method of claim 1 , wherein determining the plurality of key image frames from the video comprises evaluating a plurality of frames of the video to select, as the plurality of key video frames that comprise less than all of the frames of the video, one or more of the plurality of frames that satisfy one or more criteria.
4 . The method of claim 3 , wherein the one or more criteria include a measure of visual difference between two adjacent frames of the plurality of video frames satisfying a threshold.
5 . The method of claim 3 , wherein the one or more criteria include a new object being detected in a frame of the plurality of frames of the video.
6 . The method of claim 3 , wherein the one or more criteria include a new voice being detected in an audio portion of the video that corresponds temporally with a frame of the plurality of frames of the video.
7 . A system comprising one or more processors and a memory storing instructions that, when executed on the one or more processors, cause the one or more processors to:
determine a plurality of key image frames from a video, wherein the plurality of key image frames includes a past key image frame that occurs earlier than a current image frame in the video and a future key image frame that occurs later than the current image frame in the video;
process the plurality of key image frames, including the past and future key image frames, using a vision-language model,
wherein processing the plurality of key image frames using the vision-language model causes a natural language description of the past and future key image frames to be generated using the vision-language model;
store the natural language description for the plurality of key image frames in association with the video;
receive, from a computing device, a user query related to the video, wherein the user query is received when the current image frame of the video is being rendered; and
in response to receiving the user query,
generate a prompt based on the user query and based on the natural language description for the plurality of key image frames of the video,
process the prompt as input using a generative model, to generate a generative model output, wherein the generative model output is operable to cause a response responsive to the user query to be rendered by an output device, and
provide the generative model output to the computing device.