IP Library Granted Patent US 10,061,986
Granted Patent B2
US 10,061,986 · App. 15/210,708 · Granted Aug 28, 2018

Systems and methods for identifying activities in media contents based on prediction confidences

Inventors: Leonid Sigal (Pittsburgh, PA); Shugao Ma (Pittsburgh, PA)
Assignee: Disney Enterprises, Inc.
G06K9/00751G06T7/408G06T7/602H04L65/4069G06K2009/00738G06T2207/10016G06T2207/10024G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,061,986
App. No.
15/210,708
Granted
Aug 28, 2018
Kind
B2
Abstract

There is provided a system comprising a memory and a processor configured to receive a media content depicting an activity, extract a first plurality of features from a first segment of the media content, make a first prediction that the media content depicts a first activity based on the first plurality of features, wherein the first prediction has a first confidence level, extract a second plurality of features from a second segment of the media content, the second segment temporally following the first segment in the media content, make a second prediction that the media content depicts the first activity based on the second plurality of features, wherein the second prediction has a second confidence level, determine that the media content depicts the first activity based on the first prediction and the second prediction, wherein the second confidence level is at least as high as the first confidence level.

Claims (43)

1. A system comprising:

a non-transitory memory storing an executable code and an activity database; and

a hardware processor executing the executable code to:

receive a media content including a plurality of segments depicting an activity;

extract a first plurality of features from a first segment of the plurality of segments;

make a first prediction that the media content depicts a first activity from the activity database based on the first plurality of features, wherein the first prediction has a first confidence level;

extract a second plurality of features from a second segment of the plurality of segments, the second segment temporally following the first segment in the media content;

make a second prediction that the media content depicts the first activity based on the second plurality of features, wherein the second prediction has a second confidence level;

make a third prediction that the media content depicts a second activity, the third prediction having a third confidence level based on the first plurality of features and a fourth confidence level based on the second plurality of features;

compare a first difference between the first confidence level and the third confidence level with a second difference between the second confidence level and the fourth confidence level; and

determine that the media content depicts the first activity based on when the comparing indicates that the second difference is at least as much as the first difference based on, wherein the second confidence level is at least as high as the first confidence level.

2. The system of claim 1 , wherein the first segment is a first frame of the media content and the second segment is a second frame of the media content.

3. The system of claim 1 , wherein the second segment is a segment immediately following the first segment in the media content.

4. The system of claim 1 , wherein determining the media content depicts the first activity includes comparing a magnitude of the first confidence with a magnitude of the second confidence.

5. The system of claim 1 , wherein the determining that the media content depicts the first activity is further based on the first confidence level being higher than the third confidence level.

6. The system of claim 1 , wherein a prediction is made for each frame of the media content and each prediction that has a preceding prediction is compared to the preceding prediction to determine that the media content depicts the first activity.

7. The system of claim 1 , wherein the first activity includes a plurality of steps.

8. The system of claim 1 , wherein the hardware processor executes the executable code to:

determine a start time of the first activity in the media content based on the first confidence level and the second confidence level.

9. The system of claim 1 , wherein the hardware processor executes the executable code to:

determine an end time of the first activity in the media content based on the first confidence level and the second confidence level.

10. The system of claim 1 , wherein the system uses a Recurrent Neural Network (RNN) for predicting the media content depicts the first activity and determining a corresponding confidence level.

11. A method for use with a system including a non-transitory memory and a hardware processor, the method comprising:

receiving, using the hardware processor, a media content including a plurality of segments depicting an activity;

extracting, using the hardware processor, a first plurality of features from a first segment of the plurality of segments;

making, using the hardware processor, a first prediction that the media content depicts a first activity from the activity database based on the first plurality of features, wherein the first prediction has a first confidence level;

extracting, using the hardware processor, a second plurality of features from a second segment of the plurality of segments, the second segment temporally following the first segment in the media content;

making, using the hardware processor, a second prediction that the media content depicts the first activity based on the second plurality of features, wherein the second prediction has a second confidence level;

making, using the hardware processor, a third prediction that the media content depicts a second activity, the third prediction having a third confidence level based on the first plurality of features and a fourth confidence level based on the second plurality of features;

comparing, using the hardware processor, a first difference between the first confidence level and the third confidence level with a second difference between the second confidence level and the fourth confidence level; and

determining, using the hardware processor, that the media content depicts the first activity based on when the comparing indicates that the second difference is at least as much as the first difference based on, wherein the second confidence level is at least as high as the first confidence level.

12. The method of claim 11 , wherein the first segment is a first frame of the media content and the second segment is a second frame of the media content.

13. The method of claim 11 , wherein the second segment is a segment immediately following the first segment in the media content.

14. The method of claim 11 , wherein determining the media content depicts the first activity includes comparing a magnitude of the first confidence with a magnitude of the second confidence.

15. The method of claim 11 , wherein the determining that the media content depicts the first activity is further based on the first confidence level being higher than the third confidence level.

16. The method of claim 11 , wherein a prediction is made for each frame of the media content and each prediction that has a preceding prediction is compared to the preceding prediction to determine that the media content depicts the first activity.

17. The method of claim 11 , wherein the first activity includes a plurality of steps.

18. The method of claim 11 , further comprising:

determining, using the hardware processor, a start time of the first activity in the media content based on the first confidence level and the second confidence level.

19. The method of claim 11 , further comprising:

determining, using the hardware processor, an end time of the first activity in the media content based on the first confidence level and the second confidence level.

20. The method of claim 11 , further comprising:

determining, using the hardware processor, the media content depicts the first activity when only a portion of the first activity is shown.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2026
From: DISNEY ENTERPRISES, INC.
To: ADEIA MEDIA HOLDINGS INC.
Reel/Frame 075575/0145 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2016
From: SIGAL, LEONID; MA, SHUGAO
To: DISNEY ENTERPRISES, INC.
Reel/Frame 039162/0177 →
Continuity (2)
Provisional Application 62327951 · Apr 26, 2016
Related Publication 20170308756A1 · Oct 26, 2017