IP Library Granted Patent US 12713085
Granted Patent B2
US 12713085 · App. 18/620,998 · Granted Aug 18, 2026

Generating event commentary in videos using AI models

Inventors: Ram Rangan (Tamil Nadu, IN); Deep Shekhar (Bangalore, IN); Siddharth Sharma (San Jose, CA); Marc Seth Blackstein (Portland, OR)
Assignee: NVIDIA Corporation
H04N21/26603G06T13/205G06T13/40G06V30/10G10L13/08G10L15/00H04N21/4884
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12713085
App. No.
18/620,998
Granted
Aug 18, 2026
Kind
B2
Abstract

Disclosed are apparatuses, systems, and techniques for automatically generating commentary to videos that capture sporting activities, computer games, artistic events, political rallies, security-sensitive scenes, and/or any other actions. The techniques include processing a video segment that includes a plurality of video frames, to obtain a description of one or more objects pictured in the video segment and generating, using the obtained description, a prompt for a language model (LM). The techniques further include causing the LM to process the prompt to generate a commentary about an action performed by the one or more objects over a time interval associated with the plurality of video frames.

Claims (94)

1 . A method comprising:

sampling a plurality of video frames from a video segment;

processing, using a computer vision (CV) model, the plurality of video frames to obtain an output comprising one or more tracks for respective one or more objects pictured in the video segment, each track comprising:

an identification of a corresponding object of the one or more objects, and

a location of the corresponding object in at least a subset of the plurality of video frames;

generating, using the output, a prompt for a language model (LM); and

generating, using the LM, a commentary about an action performed by the one or more objects over a time interval associated with the plurality of video frames.

2 . The method of claim 1 , wherein the plurality of video frames are sampled from the video segment with a sampling rate set in view of a speed of the action, and wherein the output further comprises one or more of:

a description of a state of motion of the one or more objects pictured in the video segment,

a description of action performed by the one or more objects pictured in the video segment, or

a description of interaction between the one or more objects pictured in the video segment.

3 . The method of claim 1 , further comprising:

processing, using optical character recognition, the video segment to recognize one or more symbols pictured in the video segment, wherein the prompt for the LM is further generated using the one or more recognized symbols.

4 . The method of claim 1 , further comprising:

processing, using a speech recognition model, the video segment to recognize one or more speech utterances in the video segment, wherein the prompt for the LM is further generated using the one or more recognized speech utterances.

5 . The method of claim 1 , further comprising:

obtaining a representation of a type of activity captured in the video segment; and

performing at least one of:

including the obtained representation to the prompt for the LM; or

causing, prior to the processing of the prompt by the LM, the LM to process the obtained representation.

6 . The method of claim 1 , further comprising:

using the generated commentary to perform at least one of:

storing the generated commentary in a computer memory;

presenting the generated commentary on a user interface; or

causing at least a portion of the generated commentary to be attributed to one or more characters associated with an activity represented by the video segment.

7 . The method of claim 1 , further comprising:

generating a mapping of the generated commentary to one or more timestamps of the video segment.

8 . The method of claim 7 , further comprising:

generating, using the generated mapping, a closed captioning for the video segment.

9 . The method of claim 7 , further comprising:

applying the generated commentary to a text-to-speech conversion model to obtain an audio file comprising a spoken commentary about the action performed by the one or more objects.

10 . The method of claim 9 , further comprising:

generating a facial animation corresponding to the spoken commentary.

11 . The method of claim 1 , wherein the video segment is associated with at least one of:

an athletic activity,

a computer game,

an artistic event,

an activity captured by a home automation system,

an activity captured by a security surveillance system,

an activity associated with one or more vulnerable persons, or

an activity associated with an automotive environment.

12 . The method of claim 1 , wherein the prompt for the LM comprises an indication of a length limit for the commentary.

13 . The method of claim 1 , wherein the prompt for the LM comprises one or more previous instances of the commentary generated for a type of activity pictured in the video segment.

14 . A system comprising:

one or more processing units to:

sample a plurality of video frames from a video segment;

process, using a computer vision (CV) model, the plurality of video frames to obtain an output comprising:

one or more tracks for respective one or more objects pictured in the video segment, each track comprising (i) an identification of a corresponding object of the one or more objects and (ii) a location of the corresponding object in at least a subset of the plurality of video frames, and

one or more of:

a description of a state of motion of the one or more objects,

a description of action performed by the one or more objects, or

a description of interaction between the one or more objects;

generate, using the obtained output, a prompt for a language model (LM); and

cause the LM to process the prompt to generate a commentary about an action performed by the one or more objects over a time interval associated with the plurality of video frames.

15 . The system of claim 14 , wherein the system is comprised in at least one of:

an in-vehicle infotainment system for an autonomous or semi-autonomous machine;

a system for performing one or more simulation operations;

a system for performing one or more digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets;

a system for performing one or more deep learning operations;

a system implemented using an edge device;

a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content;

a system implemented using a robot;

a system for performing one or more conversational AI operations;

a system implementing one or more large language models (LLMs);

a system implementing one or more language models;

a system for performing one or more generative AI operations;

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

16 . The system of claim 14 , wherein the one or more processing units are further to:

process, using optical character recognition, the video segment to recognize one or more symbols pictured in the video segment, wherein the prompt for the LM is further generated using the one or more recognized symbols.

17 . The system of claim 14 , wherein the one or more processing units are further to:

process, using a speech recognition model, the video segment to recognize one or more speech utterances in the video segment, wherein the prompt for the LM is further generated using the one or more recognized speech utterances.

18 . The system of claim 14 , wherein the one or more processing units are further to:

use the generated commentary to perform at least one of:

storing the generated commentary in a computer memory;

presenting the generated commentary on a user interface; or

causing at least a portion of the generated commentary to be attributed to one or more characters associated with an activity represented by the video segment.

19 . The system of claim 14 , wherein the one or more processing units are further to:

generate a mapping of the generated commentary to one or more timestamps of the video segment; and

generate, using the generated mapping, a closed captioning for the video segment.

20 . A non-transitory computer-readable storage medium storing instructions thereon that, when executed by a processing device, cause the processing device to:

sample a plurality of video frames from a video segment;

process, using a computer vision (CV) model, the plurality of video frames to obtain an output comprising:

one or more tracks for respective one or more objects pictured in the video segment, each track comprising (i) an identification of a corresponding object of the one or more objects and (ii) a location of the corresponding object in at least a subset of the plurality of video frames, and

one or more of:

a description of a state of motion of the one or more objects,

a description of action performed by the one or more objects, or

a description of interaction between the one or more objects;

generate, using the obtained output, a prompt for a language model (LM); and

cause the LM to process the prompt to generate a commentary about an action performed by the one or more objects over a time interval associated with the plurality of video frames.