IP Library Granted Patent US 12664779
Granted Patent B2
US 12664779 · App. 18/584,401 · Granted Jun 23, 2026

System and method for automatic identification of spatial/temporal attention regions and training data generation using the same

Inventor: Avijit Shah (Santa Clara, CA)
Assignee: YAHOO ASSETS LLC
G06V20/42G06V10/774G06V20/44G06V20/46G06V20/49
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12664779
App. No.
18/584,401
Granted
Jun 23, 2026
Kind
B2
Abstract

The present teaching relates to identify events of interests. Given each of video clips, each capturing an event of interest, spatial attention regions are identified therefrom, each of which includes objects that meet a first condition. A temporal attention region is determined in each video clip according to a second condition. An action that causes an event of interest in the temporal attention region is labeled. The video clips, the respective spatial/temporal attention regions, and the action labels are then used to generate training data for machine learning of models for automatically determining, from an input video clip, a temporal attention zone for an event of interest and an action that causes the event of interest.

Claims (76)

1 . A method, comprising:

receiving a plurality of historic video clips, each of which captures an event of interest;

with respect to each of the historic video clips,

identifying spatial attention regions in a plurality of frames of the historic video clip, wherein each of the spatial attention regions includes one or more objects that satisfy a first predetermined condition,

determining a temporal attention region in the historic video clip based on the identified spatial attention regions in accordance with a second predetermined condition,

labeling an action occurring within the temporal attention region that causes the event of interest; and

generating, based on the historic video clips, their respective spatial and temporal attention regions, and the respective action labels, training data for machine learning to train models used in automatically determining, from an input video clip, a temporal attention zone corresponding to an event of interest and classifying an action captured in the input video clip that causes the event of interest.

2 . The method of claim 1 , wherein

an event of interest corresponds to a scoring event in a sports game; and

the scoring event occurs when an action is performed in the sports game.

3 . The method of claim 2 , wherein

the scoring event includes a basket event in a basketball game; and

an action that causes a basket event includes one of dunk, layup, hoop, and 3-pointer.

4 . The method of claim 3 , wherein the step of identifying spatial attention regions comprises:

with respect to each of the plurality of frames in the historic video clip,

detecting objects involved in the event of interest,

retrieving the first predetermined condition in an action event configuration defining a spatial relationship among the detected objects within the frame, and

identifying a spatial attention region in the frame that encompasses the detected objects when they satisfy the first predetermined condition.

5 . The method of claim 4 , wherein the first predetermined condition requires that the detected objects be within a certain distance.

6 . The method of claim 2 , wherein the step of determining at least one temporal attention region comprises:

identifying at least one key frame in the historic video clip according to the second predetermined condition defining a scoring event as the event of interest;

determining consecutive frames from the plurality of frames centering around the at least one key frame based on domain knowledge.

7 . The method of claim 6 , wherein the domain knowledge includes information on:

a frame rate of the historic video clip; and

an estimated duration of the event of interest.

8 . A machine readable and non-transitory medium having information recorded thereon, wherein the information, when read by the machine, causes the machine to perform the following steps:

receiving a plurality of historic video clips, each of which captures an event of interest;

with respect to each of the historic video clips,

identifying spatial attention regions in a plurality of frames of the historic video clip, wherein each of the spatial attention regions includes one or more objects that satisfy a first predetermined condition,

determining a temporal attention region in the historic video clip based on the identified spatial attention regions in accordance with a second predetermined condition,

labeling an action occurring within the temporal attention region that causes the event of interest; and

generating, based on the historic video clips, their respective spatial and temporal attention regions, and the respective action labels, training data for machine learning to train models used in automatically determining, from an input video clip, a temporal attention zone corresponding to an event of interest and classifying an action captured in the input video clip that causes the event of interest.

9 . The medium of claim 8 , wherein

an event of interest corresponds to a scoring event in a sports game; and

the scoring event occurs when an action is performed in the sports game.

10 . The medium of claim 9 , wherein

the scoring event includes a basket event in a basketball game; and

an action that causes a basket event includes one of dunk, layup, hoop, and 3-pointer.

11 . The medium of claim 10 , wherein the step of identifying spatial attention regions comprises:

with respect to each of the plurality of frames in the historic video clip,

detecting objects involved in the event of interest,

retrieving the first predetermined condition in an action event configuration defining a spatial relationship among the detected objects within the frame, and

identifying a spatial attention region in the frame that encompasses the detected objects when they satisfy the first predetermined condition.

12 . The medium of claim 11 , wherein the first predetermined condition requires that the detected objects be within a certain distance.

13 . The medium of claim 9 , wherein the step of determining at least one temporal attention region comprises:

identifying at least one key frame in the historic video clip according to the second predetermined condition defining a scoring event as the event of interest;

determining consecutive frames from the plurality of frames centering around the at least one key frame based on domain knowledge.

14 . The medium of claim 13 , wherein the domain knowledge includes information on:

a frame rate of the historic video clip; and

an estimated duration of the event of interest.

15 . A system, comprising:

an S/T attention segmentation unit implemented using a processor and configured for:

receiving a plurality of historic video clips, each of which captures an event of interest,

with respect to each of the historic video clips,

identifying spatial attention regions in a plurality of frames of the historic video clip, wherein each of the spatial attention regions includes one or more objects that satisfy a first predetermined condition,

determining a temporal attention region in the historic video clip based on the identified spatial attention regions in accordance with a second predetermined condition;

an action labeling unit implemented by a processor and configured for:

labeling an action occurring within the temporal attention region that causes the event of interest, and

generating, based on the historic video clips, their respective spatial and temporal attention regions, and the respective action labels, training data for machine learning to train models used in automatically determining, from an input video clip, a temporal attention zone corresponding to an event of interest and classifying an action captured in the input video clip that causes the event of interest.

16 . The system of claim 15 , wherein

an event of interest corresponds to a scoring event in a sports game; and

the scoring event occurs when an action is performed in the sports game.

17 . The system of claim 16 , wherein

the scoring event includes a basket event in a basketball game; and

an action that causes a basket event includes one of dunk, layup, hoop, and 3-pointer.

18 . The system of claim 17 , wherein the step of identifying spatial attention regions comprises:

with respect to each of the plurality of frames in the historic video clip,

detecting objects involved in the event of interest,

retrieving the first predetermined condition in an action event configuration defining a spatial relationship among the detected objects within the frame, and

identifying a spatial attention region in the frame that encompasses the detected objects when they satisfy the first predetermined condition.

19 . The system of claim 16 , wherein the step of determining at least one temporal attention region comprises:

identifying at least one key frame in the historic video clip according to the second predetermined condition defining a scoring event as the event of interest;

determining consecutive frames from the plurality of frames centering around the at least one key frame based on domain knowledge.

20 . The system of claim 19 , wherein the domain knowledge includes information on:

a frame rate of the historic video clip; and

an estimated duration of the event of interest.