IP Library Granted Patent US 10,303,756
Granted Patent B2
US 10,303,756 · App. 15/669,885 · Granted May 28, 2019

Creating a narrative description of media content and applications thereof

Inventor: Hyduke Noshadi (Northridge, CA)
Assignee: Google LLC
G06F17/241G06K9/00228G06K9/00751H04N1/00167H04N1/00196H04N2201/3261H04N2201/3266H04N2201/3274
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,303,756
App. No.
15/669,885
Granted
May 28, 2019
Kind
B2
Abstract

This invention relates to creating a narrative description of media content. In an embodiment, a computer-implemented method describes content of a group of images. The group of images includes a first image and a second image. A first object in the first image is recognized to determine a first content data. A second object in the second image is recognized to determine a second content data. Finally, a narrative description of the group of images is determined according to a parameterized template and the first and second content data.

Claims (54)

1. A computer-implemented method comprising:

recognizing two or more objects in a plurality of images within an album, wherein recognizing the two or more objects comprises recognizing one or more of a face or a landmark;

based on the recognizing, assembling a single narrative text for the plurality of images according to a storyline, wherein the single narrative text includes a plurality of sentences descriptive of the plurality of images within the album and provides a summary of the plurality of images within the album according to the storyline;

storing the single narrative text with the album as a description of the album;

and

causing the single narrative text and one or more of the plurality of images to be displayed.

2. The computer-implemented method of claim 1 , further comprising determining a parameter value based at least in part on the recognized two or more objects.

3. The computer-implemented method of claim 2 , wherein assembling the single narrative text comprises inserting the parameter value into a parameterized template.

4. The computer-implemented method of claim 3 , further comprising:

determining a plurality of timestamps, each timestamp associated with a respective image of the plurality of images;

calculating a time period that encompasses the plurality of timestamps; and

inserting the time period into the parameterized template.

5. The computer-implemented method of claim 1 , further comprising determining a region that encompasses the landmark.

6. The computer-implemented method of claim 5 , further comprising determining a parameter value based at least in part on the region, and wherein assembling the single narrative text comprises inserting the parameter value into a parameterized template.

7. The computer-implemented method of claim 1 , further comprising:

determining first content data based on a first object of the two or more objects; and

determining second content data based on a second object of the two or more objects,

wherein a parameterized template specifies how to construct the single narrative text from the first content data and second content data.

8. A computer-implemented method comprising:

recognizing two or more objects in a plurality of images within a collection of images within an album, wherein recognizing the two or more objects comprises recognizing at least one face or at least one landmark;

based on the recognizing, assembling a single narrative text for the plurality of images according to a storyline, wherein the single narrative text includes a plurality of sentences descriptive of the collection of images;

storing the single narrative text with the album as a description of the album; and

causing the single narrative text to be displayed in conjunction with the collection of images as a textual summary of the collection of images.

9. The computer-implemented method of claim 8 , further comprising determining a parameter value based at least in part on the recognized at least one face.

10. The computer-implemented method of claim 9 , wherein the parameter value is a name of an individual associated with the at least one face.

11. The computer-implemented method of claim 9 , wherein assembling the single narrative text comprises inserting the parameter value into a parameterized template.

12. The computer-implemented method of claim 8 , further comprising determining a name of the landmark, and wherein assembling the single narrative text comprises inserting the name of the landmark into a parameterized template.

13. The computer-implemented method of claim 11 , further comprising:

determining a name of a city based on metadata associated with one or more of the plurality of images, and wherein assembling the single narrative text further comprises inserting the name of the city into the parameterized template.

14. The computer-implemented method of claim 8 , further comprising:

determining a plurality of timestamps, each timestamp associated with a respective image of the plurality of images;

calculating a time period that encompasses the plurality of timestamps; and

inserting the time period into a parameterized template.

15. The computer-implemented method of claim 8 , further comprising:

determining first content data based on a first object of the two or more objects; and

determining second content data based on a second object of the two or more objects,

wherein a parameterized template specifies how to construct the single narrative text from the first content data and second content data.

16. A non-transitory computer-readable medium with instructions stored thereon that, when executed by a processor, cause the processor to perform operations comprising:

recognizing two or more objects in a plurality of images within an album, wherein recognizing the two or more objects comprises recognizing one or more of a face or a landmark;

determining a parameter value based at least in part on the recognized two or more objects;

assembling a single narrative text for the plurality of images according to a storyline, wherein the assembling comprises inserting the parameter value into a parameterized template, wherein the single narrative text includes a plurality of sentences descriptive of the plurality of images within the album and provides a summary of the plurality of images within the album according to the storyline;

storing the single narrative text with the album as a description of the album;

and

causing the single narrative text and one or more of the plurality of images to be displayed.

17. The non-transitory computer-readable medium of claim 16 , wherein the parameterized template is a sentence template that includes one or more parameters.

18. The non-transitory computer-readable medium of claim 17 , with further instructions stored thereon that, when executed by the processor, cause the processor to perform further operations comprising:

determining a region that encompasses the landmark; and

determining a parameter value based at least in part on the region, and wherein inserting the parameter value into the parameterized template comprises inserting the parameter value into the parameterized template corresponding to at least one of the one or more parameters.

19. The non-transitory computer-readable medium of claim 16 , with further instructions stored thereon that, when executed by the processor, cause the processor to perform further operations comprising:

determining a parameter value based at least in part on the recognized face, wherein the parameter value is a name of an individual associated with the recognized face.

20. The non-transitory computer-readable medium of claim 16 , with further instructions stored thereon that, when executed by the processor, cause the processor to perform further operations comprising:

determining first content data based on a first object of the two or more objects; and

determining second content data based on a second object of the two or more objects,

wherein the parameterized template specifies how to construct the single narrative text from the first content data and second content data.

Assignments (2)
CHANGE OF NAME Recorded Oct 20, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044567/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2017
From: NOSHADI, HYDUKE
To: GOOGLE INC.
Reel/Frame 043746/0292 →
Continuity (3)
Continuation 13902307 · May 24, 2013
Continuation 12393787 · Feb 26, 2009
Related Publication 20170337170A1 · Nov 23, 2017