IP Library Granted Patent US 10,482,329
Granted Patent B2
US 10,482,329 · App. 15/908,617 · Granted Nov 19, 2019

Systems and methods for identifying activities and/or events in media contents based on object data and scene data

Inventors: Zuxuan Wu (Shanghai, CN); Yanwei Fu (Pittsburgh, PA); Leonid Sigal (Pittsburgh, PA)
Assignee: Disney Enterprises, Inc.
G06K9/00751G06K9/00335G06K9/00671G06K9/00718G06K9/00744G06K9/4628G06T7/62G06T7/90G11B27/102H04L65/4069G06K2009/00738G06K2209/21G06N3/0445G06N3/0454G06N3/084G06T2207/10016G06T2207/10024G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,482,329
App. No.
15/908,617
Granted
Nov 19, 2019
Kind
B2
Abstract

There is provided a system including a non-transitory memory storing an executable code and a hardware processor executing the executable code to receive a plurality of training contents depicting a plurality of activities, extract training object data from the plurality of training contents including a first training object data corresponding to a first activity, extract training scene data from the plurality of training contents including a first training scene data corresponding to the first activity, determine that a probability of the first activity is maximized when the first training object data and the first training scene data both exist in a sample media content.

Claims (38)

1. A system comprising:

a non-transitory memory storing an executable code including an object data module, a scene data module, an image data module, and a plurality of fusion layers of a neural network; and

a hardware processor to:

receive a media content having video frames;

extract object data, by executing the object data module, from the video frames of the media content;

generate an average object data representation of the extracted object data;

extract scene data, by executing the scene data module, from the frames of the video media content;

generate an average scene data representation of the extracted scene data;

extract image data, by executing the image data module, from the frames of the video media content;

generate an average image data representation of the extracted image data;

feed the average object data representation, the average scene data representation and the average image data representation to the plurality of fusion layers of the neural network; and

classify an action in the video frames of the media content by executing the plurality of fusion layers using the average object data representation, the average scene data representation and the average image data representation.

2. The system of claim 1 , wherein the media content includes at least one of object annotations, scene annotations, and activity annotations.

3. The system of claim 1 , wherein first object data includes at least one of a color of the object, a shape of the object, a size of the object, and a relative size of the object.

4. The system of claim 1 , wherein the scene data includes at least one of a location of the scene, a lighting of the scene and an identifiable structure of the scene.

5. The system of claim 1 , wherein the image data includes at least one of a color of the image and a texture of the image.

6. The system of claim 1 , wherein the plurality of fusion layers include a first layer having a first number of object data neurons, a second number of scene data neurons, and a third number of image data neurons.

7. The system of claim 6 , wherein each of the first number and the third number is higher than the second number.

8. The system of claim 6 , wherein the plurality of fusion layers include a second layer having a fourth number of neurons across the first layer.

9. The system of claim 8 , wherein the plurality of fusion layers include a classifier layer across the second layer.

10. A method for use with a system including a hardware processor and a non-transitory memory storing an executable code including an object data module, a scene data module, an image data module, and a plurality of fusion layers of a neural network, the method comprising:

receiving, using the hardware processor, a media content having video frames;

extracting object data, by executing the object data module using the hardware processor, from the video frames of the media content;

generating, using the hardware processor, an average object data representation of the extracted object data;

extracting scene data, by executing the scene data module using the hardware processor, from the frames of the video media content;

generating, using the hardware processor, an average scene data representation of the extracted scene data;

extracting image data, by executing the image data module using the hardware processor, from the frames of the video media content;

generating, using the hardware processor, an average image data representation of the extracted image data;

feeding, using the hardware processor, the average object data representation, the average scene data representation and the average image data representation to the plurality of fusion layers of the neural network; and

classifying, by executing the plurality of fusion layers using the hardware processor, an action in the video frames of the media content using the average object data representation, the average scene data representation and the average image data representation.

11. The method of claim 10 , wherein the media content includes at least one of object annotations, scene annotations, and activity annotations.

12. The method of claim 10 , wherein first object data includes at least one of a color of the object, a shape of the object, a size of the object, and a relative size of the object.

13. The method of claim 10 , wherein the scene data includes at least one of a location of the scene, a lighting of the scene and an identifiable structure of the scene.

14. The method of claim 10 , wherein the image data includes at least one of a color of the image and a texture of the image.

15. The method of claim 10 , wherein the plurality of fusion layers include a first layer having a first number of object data neurons, a second number of scene data neurons, and a third number of image data neurons.

16. The method of claim 15 , wherein each of the first number and the third number is higher than the second number.

17. The method of claim 15 , wherein the plurality of fusion layers include a second layer having a fourth number of neurons across the first layer.

18. The method of claim 17 , wherein the plurality of fusion layers include a classifier layer across the second layer.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2026
From: DISNEY ENTERPRISES, INC.
To: ADEIA MEDIA HOLDINGS INC.
Reel/Frame 075817/0438 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2018
From: WU, ZUXUAN; FU, YANWEI; SIGAL, LEONID
To: DISNEY ENTERPRISES, INC.
Reel/Frame 045377/0352 →
Continuity (3)
Continuation 15211403 · Jul 15, 2016
Provisional Application 62327951 · Apr 26, 2016
Related Publication 20180189569A1 · Jul 5, 2018