IP Library › Granted Patent US 12,536,725
Granted Patent B2
US 12,536,725 · App. 18/498,768 · Granted Jan 27, 2026

Systems and methods for subject-driven image generation

Inventors: Junnan Li (Singapore, SG); Chu Hong Hoi (Singapore, SG); Dongxu Li (Singapore, SG)
Assignee: Salesforce, Inc.
G06T11/60G06T9/00G06V10/761G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,725
App. No.
18/498,768
Granted
Jan 27, 2026
Kind
B2
Abstract

Embodiments described herein provide systems and methods of subject-driven image generation. In at least one embodiment, a system receives, via a data interface, an image containing a subject, a text description of the subject in the image, and a text prompt relating to a different rendition of the subject. The system encodes, via an image encoder, the image into an image feature vector. The system encodes, via a text encoder, the text description int a text feature vector. The system generates, by a multimodal encoder, a vector representation of the subject based on the image feature vector and the text feature vector. The system generates, by a neural network based image generation model, an output image based on an input combining the text prompt and the vector representation.

Claims (64)

1 . A method of a subject-driven image generation framework, the method comprising:

receiving, via a data interface, a subject image containing a subject, a text description of the subject in the subject image, and a text prompt relating to a different rendition of the subject;

encoding, via an image encoder of the subject-driven image generation framework, the subject image into an image feature vector;

encoding, via a first text encoder of the subject-driven image generation framework, the text description into a text feature vector;

generating, by a multimodal encoder separated from but connected to both the image encoder and the first text encoder, a vector representation of the subject based on the image feature vector from the image encoder and the text feature vector from the text encoder;

generating, by a second text encoder of the subject-driven image generation framework, an image generation prompt based on the vector representation of the subject and the text prompt; and

generating, by a neural network based image generation model, an output image containing the subject rendered according to the text prompt based on image generation prompt combining the text prompt and the vector representation.

2 . The method of claim 1 , further comprising:

training jointly the multimodal encoder, the text encoder, and the neural network based image generation model based on a comparison of the output image and a modified image containing the subject on a different background than a background in the subject image.

3 . The method of claim 2 , wherein the generating the vector representation is further based on a plurality of query vectors, and

wherein the training includes updating the plurality of query vectors.

4 . The method of claim 1 , further comprising:

training the neural network based image generation model based on a comparison of the output image and the subject image.

5 . The method of claim 4 , further comprising:

keeping parameters of the text encoder frozen while training the neural network based image generation model.

6 . The method of claim 1 , further comprising:

generating, by the multimodal encoder, a plurality of vector representations of the subject based on a plurality of image feature vectors,

wherein the vector representation is an average of the plurality of vector representations.

7 . The method of claim 1 , further comprising:

receiving, via the data interface, a conditioning image,

wherein the generating the output image is further based on the conditioning image.

8 . A system for a subject-driven image generation framework, the system comprising:

a memory that stores a neural network based image generation model and a plurality of processor executable instructions;

a data interface that receives a subject image containing a subject, a text description of the subject in the subject image, and a text prompt relating to a different rendition of the subject; and

one or more hardware processors that read and execute the plurality of processor-executable instructions from the memory to perform operations comprising:

encoding, via an image encoder of the subject-driven image generation framework, the subject image into an image feature vector;

encoding, via a first text encoder of the subject-driven image generation framework the text description into a text feature vector;

generating, by a multimodal encoder separated from but connected to both the image encoder and the first text encoder, a vector representation of the subject based on the image feature vector from the image encoder and the text feature vector from the text encoder;

generating, by a second text encoder of the subject-driven image generation framework, an image generation prompt based on the vector representation of the subject and the text prompt; and

generating, by a neural network based image generation model, an output image containing the subject rendered according to the text prompt based on an input image generation prompt combining the text prompt and the vector representation.

9 . The system of claim 8 , the operations further comprising:

training jointly the multimodal encoder, the text encoder, and the neural network based image generation model based on a comparison of the output image and a modified image containing the subject on a different background than a background in the subject image.

10 . The system of claim 9 ,

wherein the generating the vector representation is further based on a plurality of query vectors, and

wherein the training includes updating the plurality of query vectors.

11 . The system of claim 8 , the operations further comprising:

training the neural network based image generation model based on a comparison of the output image and the subject image.

12 . The system of claim 11 , the operations further comprising:

keeping parameters of the text encoder frozen while training the neural network based image generation model.

13 . The system of claim 8 , the operations further comprising:

generating, by the multimodal encoder, a plurality of vector representations of the subject based on a plurality of image feature vectors,

wherein the vector representation is an average of the plurality of vector representations.

14 . The system of claim 8 , the operations further comprising:

receiving, via the data interface, a conditioning image,

wherein the generating the output image is further based on the conditioning image.

15 . A non-transitory machine-readable medium comprising a plurality of machine-executable instructions for a subject-driven image generation framework which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations comprising:

receiving, via a data interface, a subject image containing a subject, a text description of the subject in the subject image, and a text prompt relating to a different rendition of the subject;

encoding, via an image encoder of the subject-driven image generation framework, the subject image into an image feature vector;

encoding, via a first text encoder of the subject-driven image generation framework the text description into a text feature vector;

generating, by a multimodal encoder separated from but connected to both the image encoder and the first text encoder, a vector representation of the subject based on the image feature vector from the image encoder and the text feature vector from the text encoder;

generating, by a second text encoder of the subject-driven image generation framework, an image generation prompt based on the vector representation of the subject and the text prompt; and

generating, by a neural network based image generation model, an output image containing the subject rendered according to the text prompt based on image generation prompt combining the text prompt and the vector representation.

16 . The non-transitory machine-readable medium of claim 15 , the operations further comprising:

training jointly the multimodal encoder, the text encoder, and the neural network based image generation model based on a comparison of the output image and a modified image containing the subject on a different background than a background in the subject image.

17 . The non-transitory machine-readable medium of claim 16 ,

wherein the generating the vector representation is further based on a plurality of query vectors, and

wherein the training includes updating the plurality of query vectors.

18 . The non-transitory machine-readable medium of claim 15 , the operations further comprising:

training the neural network based image generation model based on a comparison of the output image and the subject image.

19 . The non-transitory machine-readable medium of claim 18 , the operations further comprising:

keeping parameters of the text encoder frozen while training the neural network based image generation model.

20 . The non-transitory machine-readable medium of claim 15 , the operations further comprising:

generating, by the multimodal encoder, a plurality of vector representations of the subject based on a plurality of image feature vectors,

wherein the vector representation is an average of the plurality of vector representations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 25, 2025
From: LI, JUNNAN; HOI, CHU HONG; LI, DONGXU
To: SALESFORCE, INC.
Reel/Frame 072108/0591 →
Continuity (3)
Provisional Application 63500767 · May 8, 2023
Provisional Application 63424413 · Nov 10, 2022
Related Publication 20240161369A1 · May 16, 2024
References Cited (15)
US 12198048B2 · Singh et al. · 2025 [cited by applicant]
US 20230281400A1 · Wang · 2023 [cited by examiner]
US 20230368510A1 · Chen · 2023 [cited by examiner]
US 20240087265A1 · Park · 2024 [cited by examiner]
US 20240282094A1 · Tsimpoukelli · 2024 [cited by examiner]
International Search Report and Written Opinion for PCT/US2024/027830, dated Jul. 17, 2024, 14 pages. [cited by applicant]
Chen et al., “Subject-driven Text-to-Image Generation via Apprenticeship Learning”, arxiv (Cornell University), Apr. 14, 2023, pp. 1-18, XP093186354, DOI: 10.48550/arxiv.2304.00186 Retrieved from the Internet: URL:https… [cited by applicant]
Jia et al., “Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 5, 2023, XP091… [cited by applicant]
Li et al., “BLIP-Diffusion: Pre-trained Subject Representation for Controllable Text-to-Image Generation and Editing”, arxiv.org, May 24, 2023, XP093180794, DOI: 10.48550/arxiv.2305.14720 Retrieved from the Internet: UR… [cited by applicant]
Ruiz et al., “DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation”, arxiv (Cornell University), Mar. 15, 2023, XP093179964, DOI: 10.48550/arxiv.2208.12242 Retrieved from the Internet: UR… [cited by applicant]
Deyao Zhu et al: “MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 20, 2023 (Apr.… [cited by applicant]
Jean-Baptiste Alayrac et al: “Flamingo: a Visual Language Model for Few-Shot Learning”, A Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 29, 2022 (Apr. 29, 2022), XP091… [cited by applicant]
Junnan Li et al: “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, May… [cited by applicant]
International Search Report and Written Opinion mailed Jul. 19, 2024, International Patent Application No. PCT/US2024/027695, 110 pages. [cited by applicant]
Alayrac et al., “Flamingo: a Visual Language Model for Few-Shot Learning”, DeepMind, Apr. 28, 2022., arXiv: 2204.14198v1, pp. 1-66. [cited by applicant]