IP Library › Granted Patent US 12,131,406
Granted Patent B2
US 12,131,406 · App. 18/052,870 · Granted Oct 29, 2024

Generation of image corresponding to input text using multi-text guided image cropping

Inventors: Bingchen Liu (Los Angeles, CA); Yizhe Zhu (Los Angeles, CA); Xiao Yang (Los Angeles, CA)
Assignee: LEMON INC.
G06T11/00G06F40/40G06T5/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,131,406
App. No.
18/052,870
Granted
Oct 29, 2024
Kind
B2
Abstract

Systems and methods are provided that include a processor executing a program to receive an input from a user, where the input including a first input text and a second input text. The processor is further configured to provide an initial image and, for a predetermined number of iterations, define a first and second regions of the initial image associated with the first and second input texts, respectively, define a plurality of patches of the initial image, input the initial image into a diffusion process to generate a processed image, back-propagate the processed image through a text-image match gradient calculator by generating an image embedding based on the processed image, generating a text embedding based on the region and the input text that are associated with a patch, and calculating a differential between the image embedding and the text embedding.

Claims (61)

1. A computer system for generating an output image corresponding to an input text, the computing system comprising:

a processor and memory of a computing device, the processor being configured to execute a program using portions of the memory to:

receive an input from a user, the input including a first input text and a second input text;

provide an initial image; and

for a predetermined number of iterations:

define a first region of the initial image associated with the first input text;

define a second region of the initial image associated with the second input text;

define a plurality of patches of the initial image, each patch associated with at least one of the regions;

input the initial image into a diffusion process to generate a processed image;

back-propagate the processed image through a text-image match gradient calculator to calculate a gradient against the input from the user by:

for each of the plurality of patches:

 generating an image embedding based on the processed image;

 generating a text embedding based on the region and the input text that are associated with the patch; and

 calculating a differential between the image embedding and the text embedding; and

update the initial image with an image generated by applying the calculated gradient to the processed image.

2. The computer system of claim 1 , wherein the input further includes a third input text, and wherein the processor is further configured to define a third region of the initial image associated with the third input text.

3. The computer system of claim 1 , wherein each of the plurality of patches is associated with the region that has a largest intersection with the respective patch.

4. The computer system of claim 1 , wherein the generated text embedding for each of the plurality of patches is based on a weighted average of sub-text embeddings from the regions, where weights for the weighted average are proportional to an intersected area of the region and the respective patch.

5. The computer system of claim 1 , wherein the regions are defined based upon the input.

6. The computer system of claim 5 , wherein the first region is defined based upon the first input text.

7. The computer system of claim 1 , wherein the diffusion process is a denoising diffusion implicit model.

8. The computer system of claim 1 , wherein the diffusion process includes a gradient estimator model.

9. The computer system of claim 1 , wherein the diffusion process includes a diffusion model that has been trained using a curated dataset with safe content.

10. The computer system of claim 1 , wherein the predetermined number of iterations is between 70 and 100 iterations.

11. A method for generating an output image corresponding to an input text, the method comprising steps to:

receive an input from a user, the input including a first input text and a second input text;

provide an initial image; and

for a predetermined number of iterations:

define a first region of the initial image associated with the first input text;

define a second region of the initial image associated with the second input text;

define a plurality of patches of the initial image, each patch associated with at least one of the regions;

input the initial image into a diffusion process to generate a processed image;

back-propagate the processed image through a text-image match gradient calculator to calculate a gradient against the input from the user by:

for each of the plurality of patches:

generating an image embedding based on the processed image;

generating a text embedding based on the region and the input text that are associated with the patch; and

calculating a differential between the image embedding and the text embedding; and

update the initial image with an image generated by applying the calculated gradient to the processed image.

12. The method of claim 11 , wherein the input further includes a third input text, and wherein the method further includes steps to define a third region of the initial image associated with the third input text.

13. The method of claim 11 , wherein each of the plurality of patches is associated with the region that has a largest intersection with the respective patch.

14. The method of claim 11 , wherein the generated text embedding for each of the plurality of patches is based on a weighted average of sub-text embeddings from the regions, where weights for the weighted average are proportional to an intersected area of the region and the respective patch.

15. The method of claim 11 , wherein the regions are defined based upon the input.

16. The method of claim 15 , wherein the first region is defined based upon the first input text.

17. The method of claim 11 , wherein the diffusion process is a denoising diffusion implicit model.

18. The method of claim 11 , wherein the diffusion process includes a gradient estimator model.

19. The method of claim 11 , wherein the diffusion process includes a diffusion model that has been trained using a curated dataset with safe content.

20. A computer system for generating an output image corresponding to an input text, the computing system comprising:

a processor and memory of a computing device, the processor being configured to execute a program using portions of the memory to:

receive an input from a user, the input including a first input text and a second input text;

provide an initial image; and

for a predetermined number of iterations:

define a first region of the initial image associated with the first input text;

define a second region of the initial image associated with the second input text;

define a plurality of patches of the initial image, each patch associated with at least one of the regions;

input the initial image into a diffusion process to generate a processed image, wherein the diffusion process includes a diffusion model and a gradient estimator model, wherein the diffusion model has been trained using training data generated using curated safe phrase-image pairs;

back-propagate the processed image through a text-image match gradient calculator to calculate a gradient against the input from the user by:

for each of the plurality of patches:

 generating an image embedding based on the processed image;

 generating a text embedding based on the region and the input text that are associated with the patch; and

 calculating a differential between the image embedding and the text embedding; and

update the initial image with an image generated by applying the calculated gradient to the processed image.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2023
From: LIU, BINGCHEN; ZHU, YIZHE; YANG, XIAO
To: BYTEDANCE INC.
Reel/Frame 064300/0403 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2023
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 064300/0538 →
Continuity (1)
Related Publication 20240153153A1 · May 9, 2024
Cited By (2)
US 12,266,160 US 12,494,004