IP Library › Granted Patent US 11,567,991
Granted Patent B2
US 11,567,991 · App. 16/619,006 · Granted Jan 31, 2023

Digital image classification and annotation

Inventors: Maryam Garrett (Cambridge, MA); Wan Fen Nicole Quah (Cambridge, MA); Gordon Sims (West Newbury, MA)
Assignee: GOOGLE LLC
G06F16/5866G06F16/535G06F16/55G06F40/169G06K9/6256G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,567,991
App. No.
16/619,006
Granted
Jan 31, 2023
Kind
B2
Abstract

Implementations are described herein for automatically annotating or curating digital images using various signals generated by individual users, in addition to or instead of content of the digital images themselves, thereby to enable the digital images to be retrieved from a searchable database based on their annotations. In particular, techniques are described herein for identifying events associated with a user, e.g., based on natural language input provided by a user, and automatically classifying/annotating images inferred to be related to those events.

Claims (48)

1. A method implemented using one or more processors, comprising:

obtaining a natural language input provided by a user via one or more input components of a computing device, wherein the natural language input is directed by the user to an automated assistant executing at least in part on the computing device;

analyzing the natural language input to determine an event associated with the user, one or more tokens of the natural language input that describe the event, and a date associated with the event;

identifying a plurality of digital images captured on the date;

comparing an amount of time each of the identified plurality of digital images has been displayed via one or more graphical user interfaces relative to respective amounts of time other digital images of the identified plurality of digital images have been displayed;

based on the comparing, classifying a subset of the plurality of digital images captured on the date as being related to the event, wherein other digital images of the plurality of digital images captured on the date are not classified as being related to the event; and

storing, in a searchable database, data indicative of one or more of the tokens that describe the event in association with the subset of the plurality of digital images that were classified as being related to the event, wherein the digital images in the searchable database are searchable based on the one or more tokens that are stored in association with the digital images.

2. The method of claim 1 , wherein the obtaining comprises:

receiving, by the automated assistant, via the one or more input components of the computing device, an audio recording of an utterance by the user directed to the automated assistant; and

performing speech-to-text processing on the audio recording to generate the natural language input.

3. The method of claim 1 , wherein the amount of time any given digital image of the plurality of digital images was displayed comprises a cumulative amount of time the given digital image was displayed across the one or more graphical user interfaces.

4. The method of claim 3 , wherein the one or more graphical user interfaces comprise a plurality of graphical user interfaces rendered on a plurality of different displays, and the cumulative amount of time the given image was displayed comprises a cumulative amount of time the given image was displayed across the plurality of graphical user interfaces.

5. The method of claim 1 , wherein the comparing includes determining a measure of image manipulation applied to each of the plurality of digital images via one or more digital image manipulation applications.

6. The method of claim 1 , wherein the comparing includes determining a measure of sharing associated with each of the plurality of digital images.

7. The method of claim 6 , wherein the measure of sharing associated with a given digital image of the plurality of digital images comprises a count of shares of the given digital image by a user who captured the given digital image to a plurality of other users.

8. The method of claim 6 , wherein the measure of sharing associated with a given digital image of the plurality of digital images comprises a count of shares of the given digital image across a plurality of other users.

9. The method of claim 1 , further comprising formulating, based on the one or more tokens that describe the event, a natural language caption for each digital image of the subset of digital images.

10. The method of claim 1 , further comprising:

applying one or more digital images of the subset of digital images as input across a machine learning classifier to generate an output;

comparing the output to the one or more tokens that describe the event to generate an error; and

training the machine learning classifier based on the error, wherein the training configures the machine learning classifier to classify subsequent digital images as being related to the one or more tokens that describe the event.

11. The method of claim 1 , further comprising:

performing image recognition processing on the identified plurality of images captured on the date to identify one or more objects or entities associated with the event that are depicted in a subset of the plurality of digital images, wherein the performing includes biasing the image recognition processing towards recognition of one or more of the tokens related to the event;

wherein the classifying is further based on the identified one or more objects or entities.

12. A method implemented using one or more processors, comprising:

obtaining a natural language input provided by a user via one or more input components of a computing device;

analyzing the natural language input to determine an event associated with the user, one or more tokens that describe the event, and a date associated with the event;

identifying a plurality of digital images captured on the date;

performing image recognition processing on the identified plurality of images captured on the date to identify one or more objects or entities associated with the event that are depicted in a subset of the plurality of digital images, wherein the performing includes biasing the image recognition processing towards recognition of one or more of the tokens related to the event;

classifying the subset of the plurality of digital images as being related to the event, wherein other digital images of the plurality of digital images captured on the date that do not depict the one or more objects or entities are not classified as being related to the event; and

storing, in a searchable database, data indicative of one or more of the tokens that describe the event in association with the subset of the plurality of digital images that were classified as being related to the event.

13. The method of claim 12 , wherein the obtaining comprises:

receiving, by an automated assistant executing at least in part on the computing device, via the one or more input components of the computing device, an audio recording of an utterance by the user directed to the automated assistant; and

performing speech-to-text processing on the audio recording to generate the natural language input.

14. The method of claim 12 , further comprising formulating, based on the one or more tokens that describe the event, a natural language caption for each digital image of the subset of digital images.

15. A method implemented using one or more processors, comprising:

obtaining a natural language input provided by a user via one or more input components of a computing device, wherein the natural language input is directed by the user to an automated assistant executing at least in part on the computing device;

analyzing the natural language input to determine an event associated with the user, one or more tokens of the natural language input that describe the event, and a date associated with the event;

identifying a plurality of digital images captured on the date;

comparing a number of times each of the identified plurality of digital images has been shared to respective numbers of times other digital images of the identified plurality of digital images have been shared;

based on the comparing, classifying a subset of the plurality of digital images captured on the date as being related to the event, wherein other digital images of the plurality of digital images captured on the date are not classified as being related to the event; and

storing, in a searchable database, data indicative of one or more of the tokens that describe the event in association with the subset of the plurality of digital images that were classified as being related to the event, wherein the digital images in the searchable database are searchable based on the one or more tokens that are stored in association with the digital images.

16. The method of claim 15 , wherein the obtaining comprises:

receiving, by the automated assistant, via the one or more input components of the computing device, an audio recording of an utterance by the user directed to the automated assistant; and

performing speech-to-text processing on the audio recording to generate the natural language input.

17. The method of claim 15 , wherein the number of times a given digital image of the plurality of digital images has been shared comprises a count of shares of the given digital image by a user who captured the given digital image to a plurality of other users.

18. The method of claim 15 , wherein the number of times a given digital image of the plurality of digital images has been shared further comprises a count of shares of the given digital by a plurality of other users.

19. The method of claim 15 , further comprising formulating, based on the one or more tokens that describe the event, a natural language caption for each digital image of the subset of digital images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2020
From: GARRETT, MARYAM; QUAH, WAN FEN NICOLE; SIMS, GORDON
To: GOOGLE LLC
Reel/Frame 051872/0182 →
Continuity (2)
Provisional Application 62742866 · Oct 8, 2018
Related Publication 20210073272A1 · Mar 11, 2021
Cited By (1)
US 12,380,715