IP Library Granted Patent US 12,014,259
Granted Patent B2
US 12,014,259 · App. 17/092,837 · Granted Jun 18, 2024

Generating natural language descriptions of images

Inventors: Samy Bengio (Los Altos, CA); Oriol Vinyals (London, GB); Alexander Toshkov Toshev (San Francisco, CA); Dumitru Erhan (San Francisco, CA)
Assignee: Google LLC
G06N3/047G06F40/40G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,014,259
App. No.
17/092,837
Granted
Jun 18, 2024
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating descriptions of input images. One of the methods includes obtaining an input image; processing the input image using a first neural network to generate an alternative representation for the input image; and processing the alternative representation for the input image using a second neural network to generate a sequence of a plurality of words in a target natural language that describes the input image.

Claims (40)

1. A method performed by one or more computers, the method comprising:

obtaining an input image;

processing the input image to generate an alternative representation of the input image in an embedding space representative of features for images rather than text; and

processing the alternative representation of the input image using a neural network to generate a sequence of words in a target natural language that describes the input image, including using the neural network to select words for inclusion in the sequence of words until a stop condition occurs that indicates an end to the sequence of words, wherein each word in the sequence of words after an initial word is selected by conditioning the neural network on a preceding word in the sequence of words.

2. The method of claim 1 , comprising processing the input image with a convolutional neural network to generate the alternative representation of the input image.

3. The method of claim 2 ,

wherein the convolutional neural network comprises a plurality of core neural network layers each having a respective set of parameters,

wherein processing the input image with the convolutional neural network comprises processing the input through each of the core neural network layers, and

wherein the alternative representation of the input image is the output generated by a last core neural network layer in the plurality of core neural network layers.

4. The method of claim 3 ,

wherein current values of the respective sets of parameters are determined by training a third neural network on a plurality of training images, and

wherein the third neural network includes the plurality of core neural network layers and an output layer configured to, for each training image, receive the output generated by the last core neural network layer for the training image and generate a respective score for each of a plurality of object categories, the respective score for each of the plurality of object categories representing a predicted likelihood that the training image contains an image of an object from the object category.

5. The method of claim 1 , wherein the neural network is a long-short term memory (LSTM) neural network.

6. The method of claim 5 , wherein the LSTM neural network is configured to process a representation of a current word in the sequence to generate, in accordance with a current hidden state of the LSTM neural network and current values of a set of parameters of the LSTM neural network, a respective word score for each word in a set of words that represents a respective likelihood that the word is a next word in the sequence.

7. The method of claim 6 , wherein the set of words includes a vocabulary of words in the target natural language and a special stop word.

8. The method of claim 5 , wherein processing the alternative representation of the input image using the neural network comprises:

processing the alternative representation using the LSTM neural network using a left to right beam search decoding to generate a plurality of possible sequences and a respective sequence score for each of the possible sequences; and

selecting one or more highest-scoring possible sequences as descriptions of the input image.

9. The method of claim 1 , further comprising selecting the initial word in the sequence by initializing a hidden state of the neural network with the alternative representation of the input image.

10. The method of claim 1 , wherein conditioning the neural network on a preceding word in the sequence of words comprises conditioning the neural network on a numeric representation of the preceding word.

11. The method of claim 1 , wherein the sequence of words is arranged according to an output order, and selecting a word for a current position in the output order comprises conditioning the neural network using a word that was selected at a preceding position in the output order that precedes the current position.

12. The method of claim 1 , wherein the neural network is a first neural network, and wherein the alternative representation of the input image is generated using a second neural network that is initially trained to perform an image classification task and is subsequently trained in a process that involves backpropagating gradients from the first neural network.

13. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising: obtaining an input image;

processing the input image to generate an alternative representation of the input image in an embedding space representative of features for images rather than text; and

processing the alternative representation of the input image using a neural network to generate a sequence of words in a target natural language that describes the input image, including using the neural network to select words for inclusion in the sequence of words until a stop condition occurs that indicates an end to the sequence of words, wherein each word in the sequence of words after an initial word is selected by conditioning the neural network on a preceding word in the sequence of words.

14. The system of claim 13 , wherein the operations comprise processing the input image with a convolutional neural network to generate the alternative representation of the input image.

15. The system of claim 14 ,

wherein the convolutional neural network comprises a plurality of core neural network layers each having a respective set of parameters,

wherein processing the input image with the convolutional neural network comprises processing the input through each of the core neural network layers, and

wherein the alternative representation of the input image is the output generated by a last core neural network layer in the plurality of core neural network layers.

16. The system of claim 15 ,

wherein current values of the respective sets of parameters are determined by training a third neural network on a plurality of training images, and

wherein the third neural network includes the plurality of core neural network layers and an output layer configured to, for each training image, receive the output generated by the last core neural network layer for the training image and generate a respective score for each of a plurality of object categories, the respective score for each of the plurality of object categories representing a predicted likelihood that the training image contains an image of an object from the object category.

17. The system of claim 13 , wherein the neural network is a long-short term memory (LSTM) neural network.

18. The system of claim 17 , wherein the LSTM neural network is configured to process a representation of a current word in the sequence to generate, in accordance with a current hidden state of the LSTM neural network and current values of a set of parameters of the LSTM neural network, a respective word score for each word in a set of words that represents a respective likelihood that the word is a next word in the sequence.

19. The system of claim 13 , wherein the sequence of words is arranged according to an output order, and selecting a word for a current position in the output order comprises conditioning the neural network using a word that was selected at a preceding position in the output order that precedes the current position.

20. A computer program product encoded on one or more non-transitory computer storage media, the computer program product comprising instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

obtaining an input image;

processing the input image to generate an alternative representation of the input image in an embedding space representative of features for images rather than text; and

processing the alternative representation of the input image using a neural network to generate a sequence of words in a target natural language that describes the input image, including using the neural network to select words for inclusion in the sequence of words until a stop condition occurs that indicates an end to the sequence of words, wherein each word in the sequence of words after an initial word is selected by conditioning the neural network on a preceding word in the sequence of words.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYING PARTY EXECUTION DATE PREVIOUSLY RECORDED AT REEL: 55087 FRAME: 902. ASSIGNOR(S) HEREBY CONFIRMS THE CHANGE OF NAME. Recorded Mar 12, 2024
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 066802/0085 →
CORRECTIVE ASSIGNMENT TO CORRECT THE LEGAL NAME FOR INVENTOR SAMUEL TO SAMY PREVIOUSLY RECORDED AT REEL: 055000 FRAME: 0709. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 20, 2021
From: BENGIO, SAMY; VINYALS, ORIOL; TOSHEV, ALEXANDER TOSHKOV; ERHAN, DUMITRU
To: GOOGLE INC.
Reel/Frame 055973/0790 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2021
From: BENGIO, SAMY; VINYALS, ORIOL; TOSHEV, ALEXANDER TOSHKOV; ERHAN, DUMITRU
To: GOOGLE INC.
Reel/Frame 055000/0709 →
CHANGE OF NAME Recorded Jan 22, 2021
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 055087/0902 →
Continuity (5)
Continuation 16538712 · Aug 12, 2019
Continuation 15856453 · Dec 28, 2017
Continuation 14941454 · Nov 13, 2015
Provisional Application 62080081 · Nov 14, 2014
Related Publication 20210125038A1 · Apr 29, 2021