IP Library Granted Patent US 12,192,543
Granted Patent B2
US 12,192,543 · App. 18/393,664 · Granted Jan 7, 2025

Video frame action detection using gated history

Inventors: Gaurav Mittal (Redmond, WA); Ye Yu (Redmond, WA); Mei Chen (Bellevue, WA); Junwen Chen (Rochester, NY)
Assignee: Microsoft Technology Licensing, LLC.
H04N21/23418G06T7/246G06V20/46G06T2207/10021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,192,543
App. No.
18/393,664
Granted
Jan 7, 2025
Kind
B2
Abstract

Example solutions for video frame action detection use a gated history and include: receiving a video stream comprising a plurality of video frames; grouping the plurality of video frames into a set of present video frames and a set of historical video frames, the set of present video frames comprising a current video frame; determining a set of attention weights for the set of historical video frames, the set of attention weights indicating how informative a video frame is for predicting action in the current video frame; weighting the set of historical video frames with the set of attention weights to produce a set of weighted historical video frames; and based on at least the set of weighted historical video frames and the set of present video frames, generating an action prediction for the current video frame.

Claims (55)

1. A system comprising:

a processor; and

a computer-readable medium storing instructions that are operative upon execution by the processor to:

receive a video stream comprising a plurality of video frames;

group the plurality of video frames into a set of present video frames and a set of historical video frames, the set of present video frames comprising a current video frame;

based on at least the set of historical video frames and the set of present video frames, generate an action prediction for the current video frame; and

perform background suppression, wherein the action prediction comprises a confidence and wherein performing the background suppression comprises:

generating a loss function that weights low confidence video frames more heavily, with separate emphasis on action and background classes, for a classifier that generates the action prediction.

2. The system of claim 1 , wherein the instructions are further operative to:

determine a set of attention weights for the set of historical video frames, the set of attention weights indicating how informative a video frame is for predicting action in the current video frame; and

weight the set of historical video frames with the set of attention weights to produce a set of weighted historical video frames, wherein generating the action prediction for the current video frame is based on at least the set of weighted historical video frames and the set of present video frames.

3. The system of claim 2 , wherein determining the set of attention weights comprises:

determining, for each video frame of the set of historical video frames, a position-guided gating score.

4. The system of claim 1 , wherein the instructions are further operative to:

based on at least the action prediction for the current video frame, generate an annotation for the current video frame; and

display the current video frame subject to the annotation for the current video frame.

5. The system of claim 1 , wherein the action prediction comprises a no action prediction or an action class prediction selected from a plurality of action classes.

6. The system of claim 1 , wherein the instructions are further operative to:

based on at least the set of historical video frames and the set of present video frames, generate a future action prediction for a video frame not yet observed.

7. The system of claim 6 , wherein the future action prediction is based on at least a predicted trajectory of an autonomous driving vehicle.

8. A computerized method comprising:

receiving a video stream comprising a plurality of video frames;

grouping the plurality of video frames into a set of present video frames and a set of historical video frames, the set of present video frames comprising a current video frame;

based on at least the set of historical video frames and the set of present video frames, generating an action prediction for the current video frame; and

performing background suppression, wherein the action prediction comprises a confidence and wherein performing the background suppression comprises:

generating a loss function that weights low confidence video frames more heavily, with separate emphasis on action and background classes, for a classifier that generates the action prediction.

9. The method of claim 8 , further comprising:

determining a set of attention weights for the set of historical video frames, the set of attention weights indicating how informative a video frame is for predicting action in the current video frame; and

weighting the set of historical video frames with the set of attention weights to produce a set of weighted historical video frames, wherein generating the action prediction for the current video frame is based on at least the set of weighted historical video frames and the set of present video frames.

10. The method of claim 9 , wherein determining the set of attention weights comprises:

determining, for each video frame of the set of historical video frames, a position-guided gating score.

11. The method of claim 8 , further comprising:

based on at least the action prediction for the current video frame, generating an annotation for the current video frame; and

displaying the current video frame subject to the annotation for the current video frame.

12. The method of claim 8 , wherein the action prediction comprises a no action prediction or an action class prediction selected from a plurality of action classes.

13. The method of claim 8 , further comprising:

based on at least the set of historical video frames and the set of present video frames, generating a future action prediction for a video frame not yet observed.

14. The method of claim 13 , wherein the future action prediction is based on at least a predicted trajectory of an autonomous driving vehicle.

15. One or more computer storage devices having computer-executable instructions stored thereon, which, on execution by a computer, cause the computer to perform operations comprising:

receiving a video stream comprising a plurality of video frames;

grouping the plurality of video frames into a set of present video frames and a set of historical video frames, the set of present video frames comprising a current video frame;

based on at least the set of historical video frames and the set of present video frames, generating an action prediction for the current video frame; and

performing background suppression, wherein the action prediction comprises a confidence and wherein performing the background suppression comprises:

generating a loss function that weights low confidence video frames more heavily, with separate emphasis on action and background classes, for a classifier that generates the action prediction.

16. The one or more computer storage devices of claim 15 , wherein the operations further comprise:

determining a set of attention weights for the set of historical video frames, the set of attention weights indicating how informative a video frame is for predicting action in the current video frame; and

weighting the set of historical video frames with the set of attention weights to produce a set of weighted historical video frames, wherein generating the action prediction for the current video frame is based on at least the set of weighted historical video frames and the set of present video frames.

17. The one or more computer storage devices of claim 16 , wherein determining the set of attention weights comprises:

determining, for each video frame of the set of historical video frames, a position-guided gating score.

18. The one or more computer storage devices of claim 16 , wherein the operations further comprise:

based on at least the action prediction for the current video frame, generating an annotation for the current video frame; and

displaying the current video frame subject to the annotation for the current video frame.

19. The one or more computer storage devices of claim 15 , wherein the action prediction comprises a no action prediction or an action class prediction selected from a plurality of action classes.

20. The one or more computer storage devices of claim 15 , wherein the operations further comprise:

based on at least the set of historical video frames and the set of present video frames, generating a future action prediction for a video frame not yet observed.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2023
From: MITTAL, GAURAV; YU, YE; CHEN, MEI; CHEN, JUNWEN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065937/0951 →
Continuity (3)
Continuation 17852310 · Jun 28, 2022
Provisional Application 63348993 · Jun 3, 2022
Related Publication 20240244279A1 · Jul 18, 2024
References Cited (1)
US 20220121853A1 · Son · 2022 [cited by examiner]