IP Library › Granted Patent US 12,141,236
Granted Patent B1
US 12,141,236 · App. 17/526,282 · Granted Nov 12, 2024

Vision-and-language model training

Inventors: Tarik Arici (New York, NY); Mehmet Saygin Seyfioglu (Seattle, WA); Ismail Baha Tutar (Seattle, WA); Tal Neiman (Brooklyn, NY)
Assignee: AMAZON TECHNOLOGIES, INC.
G06F18/2148G06F18/251G06F40/30G06T9/002G06V30/262G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,141,236
App. No.
17/526,282
Granted
Nov 12, 2024
Kind
B1
Abstract

Systems and methods for improving training processes for image and text applications are described. A first set of embeddings may be generated based on a text input, and a second set of embeddings may be generated via a convolutional neural network (CNN), based on an input image. The first set of embeddings and the second set of embeddings may be utilized to generate a third set of embeddings including one or more placeholder values to be replaced. The placeholder values may be replaced based on predicted values, to reconstruct the input text and image.

Claims (58)

1. A computer-implemented method, comprising:

generating a first set of embeddings based on a text input;

generating a second set of embeddings corresponding to an input image;

associating the first set of embeddings with the second set of embeddings;

generating, based at least in part on the first set of embeddings and the second set of embeddings, a third set of embeddings including one or more placeholder values associated with one or more values removed from the first set of embeddings and the second set of embeddings;

predicting one or more values corresponding to known values associated with the first set of embeddings and the second set of embeddings; and

reconstructing at least one of the text input and the image input based, at least in part, on replacing the one or more placeholder values with the one or more predicted values.

2. The computer-implemented method of claim 1 , wherein the generating the third set of embeddings comprises:

removing at least a subset of the one or more values of the first set of embeddings and the second set of embeddings;

replacing the removed subset of the one or more values with the one or more placeholder values.

3. The computer-implemented method of claim 1 , wherein the predicting one or more values to be used to fill in the one or more placeholder values of the third set of embeddings further comprises:

extracting context information from at least one of the first set of embeddings and the second set of embeddings;

predicting, based at least in part on the context information, the one or more values to replace the one or more placeholder values of the third set of embeddings.

4. The computer-implemented method of claim 1 , wherein the filling in the one or more placeholder values of the third set of embeddings comprises:

determining, based at least in part upon a first loss function and a second loss function, one or more values corresponding to the one or more placeholder values to be filled in,

wherein the first loss function is utilized to determine one or more words to fill in the one or more placeholder values of the third set of embeddings, and

wherein the second loss function is utilized to determine pixel values for one or more image regions to fill in the one or more placeholder values of the third set of embeddings.

5. The computer-implemented method of claim 1 , wherein the second set of embeddings are generated based at least in part on converting the image input to a two-dimensional representation and assigning one or more sequential numbers corresponding to one or more positional values in the two-dimensional representation.

6. A computing system, comprising:

a computing device processor; and

a memory device including instructions that, when executed by the computing device processor, enable the computing system to:

generate a first text representation having one or more values based, at least in part, on a text input;

generate a first image representation having one or more values based, at least in part, on an image input;

provide a subset of the one or more values of the text representation and a subset of the one or more values of the image representation to a transformer;

generate, via the transformer, a second text representation including the subset of values of the text representation and one or more additional values determined based, at least in part, on the one or more values of the subset corresponding to the image representation; and

generate, via the transformer, a second image representation including the subset of values of the image representation and one or more additional values determined based, at least in part, on the one or more values of the subset corresponding to the text representation.

7. The computing system of claim 6 , wherein the instructions, when executed by the computing device processor, further enable the computing system to:

decode the second text representation and the second image representation.

8. The computing system of claim 7 , wherein the instructions, when executed by the computing device processor, further enable the computing system to:

apply a first loss function and a second loss function to the decoded second text representation and the decoded second image representation, respectively.

9. The computing system of claim 6 , wherein the instructions, when executed by the computing device processor, further enable the computing system to:

generate a new first text representation comprising the subset of the one or more values of the text representation; and

generate a new first image representation comprising the subset of the one or more values of the image representation.

10. The computing system of claim 9 , wherein the instructions, when executed by the computing device processor, further enable the computing system to:

extract context information from at least one of the new first text representation and the new first image representation.

11. The computing system of claim 6 , wherein the instructions, when executed by the computing device processor, further enable the computing system to:

apply one or more values associated with the first text representation to the first image representation.

12. The computing system of claim 6 , wherein the first image representation is generated based, at least in part, on converting the image input to a two-dimensional representation and assigning one or more sequential numbers corresponding to one or more positional values in the two-dimensional representation.

13. The computing system of claim 6 , wherein the instructions, when executed by the computing device processor, further enable the computing system to:

reconstruct at least one of the first text representation and the first image representation.

14. A computer-implemented method, comprising:

generate a first text representation having one or more values based, at least in part, on a text input;

generate a first image representation having one or more values based, at least in part, on an image input;

provide a subset of the one or more values of the text representation and a subset of the one or more values of the image representation to a transformer;

generate, via the transformer, a second text representation including the subset of values of the text representation and one or more additional values determined based, at least in part, on the one or more values of the subset corresponding to the image representation; and

generate, via the transformer, a second image representation including the subset of values of the image representation and one or more additional values determined based, at least in part, on the one or more values of the subset corresponding to the text representation.

15. The computer-implemented method of claim 14 , further comprising:

decoding the second text representation and the second image representation.

16. The computer-implemented method of claim 15 , further comprising:

applying a first loss function and a second loss function to the decoded second text representation and the decoded second image representation, respectively.

17. The computer-implemented method of claim 14 , further comprising:

generating a new first text representation comprising the subset of the one or more values of the text representation; and

generating a new first image representation comprising the subset of the one or more values of the image representation.

18. The computer-implemented method of claim 17 , further comprising:

extracting context information from at least one of the new first text representation and the new first image representation.

19. The computer-implemented method of claim 14 , further comprising:

apply one or more values associated with the first text representation to the first image representation.

20. The computer-implemented method of claim 14 , wherein the first image representation is generated based, at least in part, on converting the image input to a two-dimensional representation and assigning one or more sequential numbers corresponding to one or more positional values in the two-dimensional representation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2021
From: ARICI, TARIK; SEYFIOGLU, MEHMET SAYGIN; TUTAR, ISMAIL BAHA; NEIMAN, TAL
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 058112/0040 →
Cited By (2)
US 12,266,160 US 12,592,059