IP Library › Granted Patent US 11,120,293
Granted Patent B1
US 11,120,293 · App. 15/823,249 · Granted Sep 14, 2021

Automated indexing of media content

Inventors: Jesse Jerome Rosenzweig (Portland, OR); Brian Lewis (Portland, OR); Leah Siddall (Portland, OR)
Assignee: AMAZON TECHNOLOGIES, INC.
G06K9/6202G06F16/164G06F16/7837G06F16/9535G06T7/215G06T7/246H04N19/142H04N19/87
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,120,293
App. No.
15/823,249
Filed
Nov 27, 2017
Granted
Sep 14, 2021
Kind
B1
Examiner
KASSA, YOSEF
Art Unit
2665
USPC
382/100
Abstract

Various types of objects or occurrences can be automatically detected in input media being processed using a transcoder. The media content can be analyzed to determine various transitions, such as scene changes, which provide insight into useful locations for performing object recognition. Representative frames subsequent a transition are analyzed to determine whether they are appropriate for image analysis, using factors such as amount of motion, brightness, color, or pixel disparity within the frame. If a representative frame meets the various criteria, that frame is sent to an object recognition service for analysis. The output of the service can be a set of object tags that provide information identifying the object and its location in the media. The output tags can be encoded into the output video or stored to an associated metadata file, among other such options.

Claims (67)

1. A computer-implemented method, comprising:

detecting a transition in a video input media during a transcoding of the video input media;

selecting a plurality of representative frames of image data from the video input media subsequent the transition;

determining, during the transcoding, that one or more of the representative frames satisfies at least one frame selection criterion; and

sending the image data for the representative frames to an object recognition system configured to identify one or more objects represented in the representative frame and determine a confidence level for individual identified objects of the one or more objects indicating likelihood that the respective object appears in a given representative frame of the representative frames.

2. The computer-implemented method of claim 1 , further comprising:

identifying the one or more objects represented in the image data; and

storing object data associated with the video input media, the object data identifying the one or more objects.

3. The computer-implemented method of claim 2 , further comprising:

generating object tags for the one or more objects, the object tags including at least one of an object identifier, a confidence score, or location information with respect to the video input media.

4. The computer-implemented method of claim 3 , further comprising:

receiving a request pertaining to a type of object; and

identifying the video input media based at least in part upon an associated object tag, of the object tags, corresponding to the type of object.

5. The computer-implemented method of claim 1 , further comprising:

analyzing at least one section of the representative frame with respect to at least one proximate image frame in the video input media;

generating one or more motion vectors indicative of a change in position of one or more objects represented in the representative image frame and the at least one proximate image frame; and

analyzing the at least one motion vector to determine that a motion, represented by the one or more motion vectors, satisfies a motion criterion of the at least one frame selection criterion, the motion criterion relating to at least one of an amount, direction, or type of motion represented by the one or more motion vectors.

6. The computer-implemented method of claim 1 , further comprising:

analyzing pixel values for at least one section of the representative image frame to determine whether the pixel values satisfy at least one of a diversity criterion, a compression criterion, or a color criterion of the at least one frame selection criterion.

7. The computer-implemented method of claim 1 , further comprising:

receiving a request to receive the video input media for presentation;

providing the object tags with the video input media; and

causing supplemental content to be selected for playback with the video input media based at least in part upon the object tags.

8. The computer-implemented method of claim 1 , further comprising:

detecting a set of transitions represented in the video input media; and

selecting at least one representative frame between adjacent pairs of the transitions in the video input media.

9. The computer-implemented method of claim 1 , wherein an object identified by the object recognition system includes at least one of a person, an animal, an inanimate object, a location, a product, a genre, an emotion, a sentiment, an action, an expression, or an occurrence.

10. The computer-implemented method of claim 1 , further comprising:

detecting the transition by analyzing at least one of audio data, video data, image data, captioning data, or metadata contained within the video input media.

11. A system, comprising:

at least one processor; and

memory including instructions that, when executed by the system, cause the system to:

detect a transition in a video input media during a transcoding of the video input media;

select a plurality of representative frames of image data from the video input media subsequent the transition;

determine, during the transcoding, that one or more of the representative frames satisfies-at least one frame selection criterion; and

provide the image data for the representative frames to an object recognition system configured to identify one or more objects represented in the representative frames and determine a confidence level for individual identified objects of the one or more objects indicating likelihood that the respective object appears in a given representative frame of the representative frames.

12. The system of claim 11 , wherein the instructions when executed further cause the system to:

identify the one or more objects represented in the image data; and

store object data associated with the video input media, the object data identifying the one or more objects.

13. The system of claim 11 , wherein the instructions when executed further cause the system to:

generate object tags for the one or more objects, the object tags including at least one of an object identifier, a confidence score, or location information with respect to the video input media;

receive a request pertaining to a type of object; and

identify the video input media based at least in part upon an associated object tag, of the object tags, corresponding to the type of object.

14. The system of claim 11 , wherein the instructions when executed further cause the system to:

analyze at least one section of the representative frame with respect to at least one proximate image frame in the video input media;

generate one or more motion vectors indicative of a change in position of one or more objects represented in the representative image frame and the at least one proximate image frame; and

analyze the at least one motion vector to determine that the type of motion, represented by the one or more motion vectors, satisfies a motion criterion of the at least one frame selection criterion, the motion criterion relating to at least one of an amount, direction, or type of motion represented by the one or more motion vectors.

15. The system of claim 11 , wherein the instructions when executed further cause the system to:

analyze pixel values for at least one section of the representative image frame to determine whether the pixel values satisfy at least one of a diversity criterion, a compression criterion, or a color criterion of the at least one frame selection criterion.

16. A non-transient computer readable medium storing one or more sequences of instructions executable by one or more processors to perform a set of operations comprising:

detecting a transition in video input media during a transcoding of the video input media;

selecting a plurality of representative frames of image data from the video input media subsequent the transition;

determining, during the transcoding, that one or more of the representative frames satisfies at least one frame selection criterion; and

providing the image data for the representative frames to an object recognition system configured to identify one or more objects represented in the representative frames and determine a confidence level for individual identified objects of the one or more objects indicating likelihood that the respective object appears in a given representative frame of the representative frames.

17. The system of claim 16 , wherein the instructions further comprise:

identifying the one or more objects represented in the image data; and

storing object data associated with the video input media, the object data identifying the one or more objects.

18. The system of claim 16 , wherein the instructions further comprise:

generating object tags for the one or more objects, the object tags including at least one of an object identifier, a confidence score, or location information with respect to the video input media;

receiving a request pertaining to a type of object; and

identifying the video input media based at least in part upon an associated object tag, of the object tags, corresponding to the type of object.

19. The system of claim 16 , wherein the instructions further comprise:

analyzing at least one section of the representative frames with respect to at least one proximate image frame in the video input media;

generating one or more motion vectors indicative of a change in position of one or more objects represented in the representative image frames and the at least one proximate image frame; and

analyzing the at least one motion vector to determine that the type of motion, represented by the one or more motion vectors, satisfies a motion criterion of the at least one frame selection criterion, the motion criterion relating to at least one of an amount, direction, or type of motion represented by the one or more motion vectors.

20. The system of claim 16 , wherein the instructions further comprise:

analyzing pixel values for at least one section of the representative image frames to determine whether the pixel values satisfy at least one of a diversity criterion, a compression criterion, or a color criterion of the at least one frame selection criterion.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 27, 2017
From: ROSENZWEIG, JESSE JEROME; LEWIS, BRIAN; SIDDALL, LEAH
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 044228/0387 →
Cited By (3)
US 12,489,953 US 12,541,948 US 12,695,952