IP Library Granted Patent US 12,586,364
Granted Patent B2
US 12,586,364 · App. 18/053,450 · Granted Mar 24, 2026

Single image concept encoder for personalization using a pretrained diffusion model

Inventors: Saeid Motiian (San Francisco, CA); Shabnam Ghadar (Menlo Park, CA)
Assignee: ADOBE INC.
G06V10/82G06V10/751G06V10/771
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,364
App. No.
18/053,450
Granted
Mar 24, 2026
Kind
B2
Abstract

Systems and methods for image processing are provided. One aspect of the systems and methods includes identifying a style image including a target style. A style encoder network generates a style vector representing the target style based on the style image. The style encoder can be trained based on a style loss that encourages the network to match a desired style. A a diffusion model generates a synthetic image that includes the target style based on the style vector. The diffusion model is trained independently of the style encoder network.

Claims (62)

1 . A method comprising:

identifying a style image and a text prompt, wherein the style image indicates a target style and the text prompt indicates image content;

generating a style vector representing the target style based on the style image using a style encoder network;

generating a guidance embedding by inserting the style vector corresponding to the style image into an embedding of the text prompt in place of a part of the text prompt, wherein the guidance embedding represents the target style from the style image in an embedding space for guidance of a diffusion model; and

generating, using the diffusion model, a synthetic image by denoising noisy image features based on the guidance embedding using the diffusion model, wherein the synthetic image depicts the image content from the text prompt with the target style from the style image.

2 . The method of claim 1 , wherein:

the style vector represents the target style in the embedding space for conditional guidance of the diffusion model.

3 . The method of claim 1 , wherein:

the style vector does not encode semantic content of the style image.

4 . The method of claim 1 , further comprising:

generating at least one guidance vector based on the text prompt, wherein the style vector is in a same latent space as the at least one guidance vector, and wherein the synthetic image is generated to include the content based on the at least one guidance vector.

5 . The method of claim 1 , further comprising:

identifying a target layout for the synthetic image; and

generating a layout vector representing the target layout using a layout encoder network, wherein the synthetic image is generated based on the layout vector and is arranged according to the target layout.

6 . The method of claim 1 , further comprising:

identifying audio data indicating content for the synthetic image; and

generating an audio vector representing the audio data using an audio encoder network, wherein the synthetic image is generated based on the audio vector and includes the content from the audio data.

7 . The method of claim 1 , further comprising:

identifying depth information for the synthetic image; and

generating a depth vector representing the depth information using a depth encoder network, wherein the synthetic image is generated based on the depth vector and is arranged according to the depth information.

8 . A method comprising:

identifying a plurality of training images including a target style;

computing an optimized style vector based on the plurality of training images;

identifying a training image that includes the target style;

encoding the training image using a style encoder network to obtain a style vector representing the target style in a latent space for guidance of a diffusion model, wherein the style vector represents the target style from the training image in an embedding space for guidance of the diffusion model;

computing a style loss based on the training image by comparing the style vector to the optimized style vector;

training the style encoder network by updating parameters of the style encoder network based on the style loss; and

inserting an output of the trained style encoder network into an embedding of a text prompt in place of a part of the text prompt to obtain a guidance embedding for the diffusion model.

9 . The method of claim 8 , further comprising:

identifying an original image from the plurality of images and a prompt describing the original image;

generating a synthetic image based on the original image, the prompt, and the optimized style vector;

comparing the synthetic image to the original image; and

updating the optimized style vector based on the comparison.

10 . The method of claim 8 , further comprising:

computing an artistic style score for a plurality of candidate training images using an artistic style classifier network; and

selecting the plurality of training images based on the artistic style score.

11 . The method of claim 8 , further comprising:

generating an original style representation of the training image;

generating a synthetic image based on the style vector using a diffusion network;

generating a predicted style representation of the synthetic image; and

comparing the original style representation and the predicted style representation to obtain the style loss.

12 . The method of claim 8 , further comprising:

training an audio encoder network to generate an audio vector representing audio data in the training image, wherein the synthetic image is generated based on the audio vector.

13 . The method of claim 8 , further comprising:

training a layout encoder network to generate a layout vector representing layout information of the training image, wherein the synthetic image is generated based on the layout vector and is arranged according to the layout information.

14 . The method of claim 8 , further comprising:

training a depth encoder network to generate a depth vector representing depth information of the training image, wherein the synthetic image is generated based on the depth vector and is arranged according to the depth information.

15 . An apparatus comprising:

a processor; and

a memory including instructions executable by the processor to:

identify a style image and a text prompt, wherein the style image indicates a target style and the text prompt indicates image content;

generate a style vector representing the target style based on the style image using a style encoder network;

generate a guidance embedding by inserting the style vector corresponding to the style image into an embedding of the text prompt in place of a part of the text prompt, wherein the guidance embedding represents the target style from the style image in an embedding space for guidance of a diffusion model; and

generate, using the diffusion model, a synthetic image by denoising noisy image features based on the guidance embedding using a diffusion model, wherein the synthetic image depicts the image content from the text prompt with the target style from the style image.

16 . The apparatus of claim 15 , the instructions further executable to:

generate a layout vector representing a target layout using a layout encoder network, wherein the synthetic image is generated based on the layout vector and is arranged according to the target layout.

17 . The method apparatus of claim 15 , the instructions further executable to:

generate an audio vector representing audio data using an audio encoder network, wherein the synthetic image is generated based on the audio vector and includes the content from the audio data.

18 . The apparatus of claim 15 , the instructions further executable to:

generate a depth vector representing depth information using a depth encoder network, wherein the synthetic image is generated based on the depth vector and is arranged according to the depth information.

19 . The apparatus of claim 15 , the instructions further executable to:

compute an artistic style score for candidate training images using an artistic style classifier network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2022
From: MOTIIAN, SAEID; GHADAR, SHABNAM
To: ADOBE INC.
Reel/Frame 061689/0454 →
Continuity (1)
Related Publication 20240153259A1 · May 9, 2024
References Cited (15)
US 11113578B1 · Brandt et al. · 2021 [cited by applicant]
US 20200342652A1 · Rowell · 2020 [cited by examiner]
US 20210110588A1 · Adamson, III · 2021 [cited by examiner]
US 20220108417A1 · Liu · 2022 [cited by examiner]
US 20220156987A1 · Chandran et al. · 2022 [cited by applicant]
US 20220245322A1 · Lundin · 2022 [cited by examiner]
US 20230289952A1 · Soborski · 2023 [cited by examiner]
CN 106952224B · 2019 [cited by applicant]
CN 114880441A · 2022 [cited by applicant]
CN 115222583A · 2022 [cited by applicant]
CN 115272121A · 2022 [cited by applicant]
Gu, Shuyang, et al. “Vector Quantized Diffusion Model for Text-to-Image Synthesis.” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2022.https://ieeexplore.ieee.org/abstract/document/98… [cited by examiner]
Hertz, Amir, et al. “Prompt-to-prompt image editing with cross attention control.” arXiv preprint arXiv:2208.01626 (2022).https://arxiv.org/abs/2208.01626 (Year: 2022). [cited by examiner]
Combined Search and Examination Report dated Feb. 22, 2024 in related Great Britain Patent Application No. GB2313666.6, 6 pages. [cited by applicant]
1Gal, et al, “An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion”, arXiv preprint arXiv:2208.01618v1 [cs.CV] Aug. 2, 2022, 26 pages. [cited by applicant]