IP Library › Granted Patent US 12,354,244
Granted Patent B2
US 12,354,244 · App. 18/175,156 · Granted Jul 8, 2025

Systems and methods for reversible transformations using diffusion models

Inventors: Nikhil Naik (Mountain View, CA); Bram Wallace (Auburn, WA)
Assignee: Salesforce, Inc.
G06T5/70G06T5/30G06T5/50G06T2207/20081G06T2207/20084G06T2207/20216
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,244
App. No.
18/175,156
Granted
Jul 8, 2025
Kind
B2
Abstract

Embodiments described herein provide systems and methods for image editing, a first copy and a second copy of an input image are generated; noise is iteratively added to the first copy and the second copy by: updating the first copy based on a first inverted output of a denoising diffusion model (DDM) based on the second copy and a first caption and updating the second copy based on a second inverted output of the DDM based on the first copy and the first caption. A resultant noised image is iteratively denoised by a reverse process using the DDM conditioned on a second caption, thereby producing a final image.

Claims (74)

1. A method of image editing using a denoising diffusion model (DDM), the method comprising:

receiving, via a data interface, an input image, a first supplementary information, and a second supplementary information relating to the input image;

encoding, by an autoencoder, the input image into an image representation;

generating a first copy and a second copy based on the image representation;

iteratively adding noise to the first copy and the second copy by:

updating the first copy based on a first inverted output of the DDM based on the second copy and the first supplementary information, and

updating the second copy based on a second inverted output of the DDM based on the first copy and the first supplementary information;

obtaining a noised image representation based on at least one of the updated first copy or the updated second copy after a first pre-defined number of iterations of adding noise;

iteratively removing noise from the updated first copy and the updated second copy based on the second supplementary information, thereby producing a final image representation; and

generating, by an image decoder, a resulting image output based on the final image representation.

2. The method of claim 1 , wherein the iteratively removing noise comprises:

generating a first intermediate image representation based on the updated first copy and a first non-inverted output of the DDM based on the updated second copy and the second supplementary information; and

generating a second intermediate image representation based on the updated second copy and a second non-inverted output of the DDM based on the updated first copy and the second supplementary information,

wherein the final image representation is based on at least one of the first intermediate image representation and the second intermediate image representation after a second pre-defined number of iterations of denoising.

3. The method of claim 2 , wherein the iteratively removing noise further comprises:

updating the updated first copy based on a first dilation of the first copy and the second copy, and

updating the second copy based on a second dilation of the first copy and the second copy.

4. The method of claim 1 , wherein the iteratively adding noise further comprises:

updating the first copy based on a first average of the first copy and the second copy, and

updating the second copy based on a second average of the first copy and the second copy.

5. The method of claim 4 , wherein:

the first average and the second average are weighted averages; and

a weight applied to the first copy in the first average is the weight applied to the second copy in the second average.

6. The method of claim 1 , wherein the first supplementary information is a first text caption describing the input image, and the second supplementary information is a second text caption different from the first text caption.

7. The method of claim 6 , wherein the resulting image output includes a modification to the input image according to the second text caption.

8. The method of claim 1 , wherein the first supplementary information is a first reference image, and the second supplementary information is a second reference image different from the first reference image.

9. A system for image editing using a denoising diffusion model (DDM), the system comprising:

a memory that stores the DDM model and a plurality of processor executable instructions;

a communication interface that receives an input image, a first supplementary information, and a second supplementary information relating to the input image; and

one or more hardware processors that read and execute the plurality of processor-executable instructions from the memory to perform operations comprising:

encoding, by an autoencoder, the input image into an image representation;

generating a first copy and a second copy based on the image representation;

iteratively adding noise to the first copy and the second copy by:

updating the first copy based on a first inverted output of the DDM based on the second copy and the first supplementary information, and

updating the second copy based on a second inverted output of the DDM based on the first copy and the first supplementary information;

obtaining a noised image representation based on at least one of the updated first copy or the updated second copy after a first pre-defined number of iterations of adding noise;

iteratively removing noise from the updated first copy and the updated second copy based on the second supplementary information, thereby producing a final image representation; and

generating, by an image decoder, a resulting image output based on the final image representation.

10. The system of claim 9 , wherein the iteratively removing noise comprises:

generating a first intermediate image representation based on the updated first copy and a first non-inverted output of the DDM based on the updated second copy and the second supplementary information; and

generating a second intermediate image representation based on the updated second copy and a second non-inverted output of the DDM based on the updated first copy and the second supplementary information,

wherein the final image representation is based on at least one of the first intermediate image representation and the second intermediate image representation after a second pre-defined number of iterations of denoising.

11. The system of claim 10 , wherein the iteratively removing noise further comprises:

updating the updated first copy based on a first dilation of the first copy and the second copy, and

updating the second copy based on a second dilation of the first copy and the second copy.

12. The system of claim 9 , wherein the iteratively adding noise further comprises:

updating the first copy based on a first average of the first copy and the second copy, and

updating the second copy based on a second average of the first copy and the second copy.

13. The system of claim 12 , wherein:

the first average and the second average are weighted averages; and

a weight applied to the first copy in the first average is the weight applied to the second copy in the second average.

14. The system of claim 9 , wherein the first supplementary information is a first text caption describing the input image, and the second supplementary information is a second text caption different from the first text caption.

15. The system of claim 14 , wherein the resulting image output includes a modification to the input image according to the second text caption.

16. The system of claim 9 , wherein the first supplementary information is a first reference image, and the second supplementary information is a second reference image different from the first reference image.

17. A non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations comprising:

receiving, via a data interface, an input image, a first supplementary information, and a second supplementary information relating to the input image;

encoding, by an autoencoder, the input image into an image representation;

generating a first copy and a second copy based on the image representation;

iteratively adding noise to the first copy and the second copy by:

updating the first copy based on a first inverted output of a denoising diffusion model (DDM) based on the second copy and the first supplementary information, and

updating the second copy based on a second inverted output of the DDM based on the first copy and the first supplementary information;

obtaining a noised image representation based on at least one of the updated first copy or the updated second copy after a first pre-defined number of iterations of adding noise;

iteratively removing noise from the updated first copy and the updated second copy based on the second supplementary information, thereby producing a final image representation; and

generating, by an image decoder, a resulting image output based on the final image representation.

18. The non-transitory machine-readable medium of claim 17 , wherein the iteratively removing noise comprises:

generating a first intermediate image representation based on the updated first copy and a first non-inverted output of the DDM based on the updated second copy and the second supplementary information; and

generating a second intermediate image representation based on the updated second copy and a second non-inverted output of the DDM based on the updated first copy and the second supplementary information,

wherein the final image representation is based on at least one of the first intermediate image representation and the second intermediate image representation after a second pre-defined number of iterations of denoising.

19. The non-transitory machine-readable medium of claim 18 , wherein the iteratively removing noise further comprises:

updating the updated first copy based on a first dilation of the first copy and the second copy, and

updating the second copy based on a second dilation of the first copy and the second copy.

20. The non-transitory machine-readable medium of claim 17 , wherein the iteratively adding noise further comprises:

updating the first copy based on a first average of the first copy and the second copy, and

updating the second copy based on a second average of the first copy and the second copy.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2023
From: NAIK, NIKHIL; WALLACE, BRAM
To: SALESFORCE, INC.
Reel/Frame 062902/0555 →
Continuity (2)
Provisional Application 63383352 · Nov 11, 2022
Related Publication 20240161248A1 · May 16, 2024
References Cited (20)
US 11599972B1 · Xu · 2023 [cited by examiner]
US 11922550B1 · Ramesh · 2024 [cited by examiner]
US 20220122308A1 · Kalarot · 2022 [cited by examiner]
US 20220398697A1 · Vahdat · 2022 [cited by examiner]
US 20220405583A1 · Vahdat · 2022 [cited by examiner]
US 20230067841A1 · Saharia · 2023 [cited by examiner]
US 20230095092A1 · Xiao · 2023 [cited by examiner]
US 20230103638A1 · Saharia · 2023 [cited by examiner]
US 20230109379A1 · Kreis · 2023 [cited by examiner]
US 20230153949A1 · Huang · 2023 [cited by examiner]
US 20230153959A1 · Saharia · 2023 [cited by examiner]
US 20230368073A1 · Karras · 2023 [cited by examiner]
US 20230368337A1 · Karras · 2023 [cited by examiner]
US 20230377099A1 · Kreis · 2023 [cited by examiner]
US 20230419075A1 · Koike Akino · 2023 [cited by examiner]
US 20240087179A1 · Min · 2024 [cited by examiner]
US 20240087196A1 · Min · 2024 [cited by examiner]
US 20240104698A1 · Nie · 2024 [cited by examiner]
US 20240161250A1 · Balaji · 2024 [cited by examiner]
US 20240161327A1 · Chen · 2024 [cited by examiner]