IP Library Granted Patent US 12,260,642
Granted Patent B2
US 12,260,642 · App. 17/560,386 · Granted Mar 25, 2025

Computerized system and method for fine-grained video frame classification and content creation therefrom

Inventors: Deven Santosh Shah (San Jose, CA); Avijit Shah (Santa Clara, CA); Topojoy Biswas (San Jose, CA); Biren Barodia (Milpitas, CA)
Assignee: YAHOO AD TECH LLC
G06V20/44G06V10/764G06V20/41G06V20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,642
App. No.
17/560,386
Granted
Mar 25, 2025
Kind
B2
Abstract

The disclosed systems and methods provide a novel framework that enables cost-effective, accurate and scalable detection and recognition of key events in sporting or live events. The framework functions by creating a domain-specific video dataset with frame level annotations (i.e., deep domain datasets) and then training a lightweight camera view classifier to detect camera views for a given video. The disclosed framework uses pre-trained pose estimation and panoptic segmentation models along with geometric rules as labeling functions to define scene types and derive frame level classification training data. According to some embodiments, disclosed frameworks may be used to identify key persons or events, select a thumbnail corresponding to a key person or event, generate personalized highlights to enhance user experience and social media promotions for a team, sport or players, and predict and select the best camera view sequence for automatic highlights generation.

Claims (51)

1. A method comprising:

identifying, by a device, a video depicting a key event, the video having a plurality of frames;

extracting, by the device, a sequence of frames from the plurality of frames;

determining, by the device, a camera view for each frame of the sequence of frames to form a sequence of camera views by applying a pre-trained panoptic segmentation model to identify relevant objects within the frame;

applying a pre-trained pose estimation model to estimate poses for any persons within the frame;

assigning a label to the frame based on a predetermined labeling function and semantic information yielded by the panoptic segmentation model and the pose estimation model; and

determining, by the device, a type of key event from the sequence of camera views by comparing the sequence of camera views to predetermined arrangements of camera views associated with different types of key events.

2. The method of claim 1 , wherein determining a camera view for each frame of the sequence of frames further comprises analyzing each frame using a lightweight classification model.

3. The method of claim 2 , wherein the lightweight classification model is trained on a domain-specific video dataset.

4. The method of claim 3 , wherein the domain-specific video dataset is created by performing steps comprising:

retrieving a training video from a database, the training video having a plurality of training video frames;

applying a previously trained feature identification model to each of the training video frames to identify at least one feature of the training video frame;

applying at least one geometric rule to the at least one feature of the training video frame to determine a label of the training video frame; and

determining from the at least one label, a camera view corresponding to the training video frame.

5. The method of claim 4 , wherein the previously trained feature identification model comprises a panoptic segmentation model.

6. The method of claim 4 , wherein the previously trained feature identification model comprises a pose estimation model.

7. The method of claim 4 , wherein the at least one geometric rule is predetermined according to a domain of the domain-specific video dataset.

8. The method of claim 4 , wherein the previously trained feature identification model was trained using a domain-agnostic training dataset.

9. The method of claim 1 , further comprising determining from the sequence of camera views a start time and an end time of the key event.

10. A computing device comprising:

a processor configured to:

identify a video depicting a key event, the video having a plurality of frames;

extract, a sequence of frames from the plurality of frames;

determine, using a lightweight classification model, a camera view for each frame of the sequence of frames, the camera views of the sequence of frames forming a sequence of camera views the lightweight classification model comprising a pre-trained panoptic segmentation model to identify relevant objects within the frame;

applying a pre-trained pose estimation model to estimate poses for any persons within the frame;

assigning a label to the frame based on a predetermined labeling function and semantic information yielded by the panoptic segmentation model and the pose estimation model; and

determine, a type of key event from the sequence of camera views by comparing the sequence of camera views to predetermined arrangements of camera views associated with different types of key events.

11. The device of claim 10 , wherein the processor is configured to train the lightweight classification model on a domain-specific video dataset.

12. The device of claim 11 , further comprising a memory containing a database; and wherein the processor is further configured to create the domain-specific video dataset by:

retrieving, a training video from the database, the training video having a plurality of training video frames;

applying, a previously trained feature identification model to each of the training video frames to identify at least one feature of the training video frame;

applying, at least one geometric rule to the at least one feature of the training video frame to determine a label of the training video frame; and

determining, from the at least one label, a camera view corresponding to the training video frame.

13. The device of claim 12 , wherein the previously trained feature identification model comprises a panoptic segmentation model.

14. The device of claim 12 , wherein the previously trained feature identification model comprises a pose estimation model.

15. A non-transitory computer-readable storage medium for tangibly storing computer program instructions capable of being executed by a computer processor, the computer program instructions defining steps of:

identifying, by a device, a video depicting a key event, the video having a plurality of frames;

extracting, by the device, a sequence of frames from the plurality of frames;

determining, by the device, a camera view for each frame of the sequence of frames to form a sequence of camera views by applying a pre-trained panoptic segmentation model to identify relevant objects within the frame;

applying, by the device, a pre-trained pose estimation model to estimate poses for any persons within the frame;

assigning, by the device, a label to the frame based on a predetermined labeling function and semantic information yielded by the panoptic segmentation model and the pose estimation model; and

determining, by the device, a type of key event from the sequence of camera views by comparing the sequence of camera views to predetermined arrangements of camera views associated with different types of key events.

16. The non-transitory computer-readable storage medium of claim 15 , wherein determining a camera view for each frame of the sequence of frames further comprises analyzing each frame using a lightweight classification model.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the lightweight classification model is trained on a domain-specific video dataset.

18. The non-transitory computer-readable storage medium of claim 17 , wherein the domain-specific video dataset is created by performing steps comprising:

retrieving a training video from a database, the training video having a plurality of training video frames;

applying a previously trained feature identification model to each of the training video frames to identify at least one feature of the training video frame;

applying at least one geometric rule to the at least one feature of the training video frame to determine a label of the training video frame; and

determining from the at least one label, a camera view corresponding to the training video frame.

19. The non-transitory computer-readable storage medium of claim 18 , wherein the previously trained feature identification model comprises one of a panoptic segmentation model and a pose estimation model.

20. The non-transitory computer-readable storage medium of claim 15 , the steps further comprising determining from the sequence of camera views a start time and an end time of the key event.

Assignments (2)
CHANGE OF NAME Recorded Mar 22, 2022
From: VERIZON MEDIA INC.
To: YAHOO AD TECH LLC
Reel/Frame 059472/0328 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2021
From: SHAH, DEVEN SANTOSH; SHAH, AVIJIT; BISWAS, TOPOJOY; BARODIA, BIREN
To: VERIZON MEDIA INC.
Reel/Frame 058467/0935 →
Continuity (1)
Related Publication 20230206632A1 · Jun 29, 2023
References Cited (4)
US 20190377957A1 · Johnston · 2019 [cited by examiner]
US 20200023262A1 · Young · 2020 [cited by examiner]
US 20210004589A1 · Turkelson · 2021 [cited by examiner]
US 20220319016A1 · Graber · 2022 [cited by examiner]