IP Library › Granted Patent US 12,423,831
Granted Patent B2
US 12,423,831 · App. 18/500,263 · Granted Sep 23, 2025

Background separation for guided generative models

Inventors: Sudeep Katakol (San Francisco, CA); Siddharth Iyer (San Francisco, CA); Aliakbar Darabi (Newcastle, WA)
Assignee: ADOBE INC.
G06T7/194G06T7/11G06T11/60G06T2207/10024G06T2207/20081G06T2207/20084G06T2207/20212
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,831
App. No.
18/500,263
Filed
Nov 2, 2023
Granted
Sep 23, 2025
Kind
B2
Art Unit
2619
USPC
345/636
Abstract

Embodiments of the present disclosure include obtaining an input image and an approximate mask that approximately indicates a foreground region of the input image. Some embodiments generate an unconditional mask of the foreground region based on the input image. A conditional mask of the foreground region is generated based on the input image and the approximate mask. Then, an output image is generated based on the unconditional mask and the conditional mask. In some cases, the output image includes the foreground region of the input image.

Claims (50)

1. A method comprising:

obtaining an input image and an approximate mask that approximately indicates a foreground region of the input image;

generating, by an unconditional mask network, an unconditional mask of the foreground region based on the input image;

generating, by a conditional mask network, a conditional mask of the foreground region based on the input image and the approximate mask; and

generating an output image including the foreground region of the input image based on the unconditional mask and the conditional mask.

2. The method of claim 1 , wherein generating the output image comprises:

combining the unconditional mask and the conditional mask to obtain a combined mask, wherein the output image is generated based on the combined mask.

3. The method of claim 1 , wherein:

the input image is generated based on the approximate mask.

4. The method of claim 1 , wherein generating the output image comprises:

computing a distance transformation map based on the approximate mask, wherein the output image is generated based on the distance transformation map.

5. The method of claim 1 , wherein generating the output image comprises:

computing a color distance map based on the approximate mask, wherein the output image is generated based on the color distance map.

6. The method of claim 1 , wherein generating the output image comprises:

performing island removal on the input image, wherein the output image is generated based on the island removal.

7. The method of claim 1 , wherein:

the foreground region comprises text based on a font and modified with a text effect, and wherein the approximate mask is based on the text and the font without the text effect.

8. The method of claim 1 , wherein:

the unconditional mask comprises a probability mask and the conditional mask comprises a refined probability mask.

9. The method of claim 1 , wherein:

the unconditional mask network is trained to generate the unconditional mask of the foreground region based on the input image and the conditional mask network is trained to generate the conditional mask of the foreground region based on the input image and the approximate mask.

10. A method comprising:

initializing an unconditional mask network and a conditional mask network;

receiving training data including an input image, an approximate mask indicating a foreground region of the input image, and a ground-truth mask;

training, using the training data, the unconditional mask network to generate an unconditional mask of the foreground region based on the input image; and

training, using the training data, the conditional mask network to generate a conditional mask of the foreground region based on the input image and the approximate mask.

11. The method of claim 10 , further comprising:

combining the unconditional mask and the conditional mask to obtain a combined mask, wherein an output image is generated based on the combined mask.

12. The method of claim 10 , wherein:

the foreground region comprises text based on a font and modified with a text effect, and wherein the approximate mask is based on the text and the font without the text effect.

13. The method of claim 10 , wherein:

the conditional mask comprises a refined probability mask.

14. An apparatus comprising:

at least one processor;

at least one memory including instructions executable by the at least one processor;

a user interface comprising parameters stored in the at least one memory and configured to obtain an input image and an approximate mask that approximately indicates a foreground region of the input image;

an unconditional mask network comprising parameters stored in the at least one memory and configured to generate an unconditional mask of the foreground region based on the input image; and

a conditional mask network comprising parameters stored in the at least one memory and configured to generate a conditional mask of the foreground region based on the input image and the approximate mask.

15. The apparatus of claim 14 , further comprising:

an image generation model configured to generate an output image including the foreground region of the input image based on the unconditional mask and the conditional mask.

16. The apparatus of claim 15 , further comprising:

a distance transform component configured to compute a distance transformation map based on the approximate mask, wherein the output image is generated based on the distance transformation map.

17. The apparatus of claim 15 , further comprising:

a color distance component configured to compute a color distance map based on the approximate mask, wherein the output image is generated based on the color distance map.

18. The apparatus of claim 15 , further comprising:

an island removal component configured to perform island removal on the input image, wherein the output image is generated based on the island removal.

19. The apparatus of claim 15 , further comprising:

a mask combination component configured to combine the unconditional mask and the conditional mask to obtain a combined mask, wherein the output image is generated based on the combined mask.

20. The apparatus of claim 14 , further comprising:

a text-effect model configured to generate the input image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2023
From: KATAKOL, SUDEEP; IYER, SIDDHARTH; DARABI, ALIAKBAR
To: ADOBE INC.
Reel/Frame 065433/0103 →
Continuity (2)
Provisional Application 63495194 · Apr 10, 2023
Related Publication 20240338829A1 · Oct 10, 2024
References Cited (15)
US 12205207B2 · Smetanin et al. · 2025 [cited by applicant]
US 20200364913A1 · Bradski · 2020 [cited by applicant]
US 20220414949A1 · Guan · 2022 [cited by examiner]
US 20230135978A1 · Price et al. · 2023 [cited by applicant]
US 20240127510A1 · Darabi et al. · 2024 [cited by applicant]
Combined Search and Examination Report in related United Kingdom Patent Application No. GB2407653.1, 5 pages. [cited by applicant]
Saharia, et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, Advances in Neural Information Processing Systems, 35, 2022, pp. 36479-36494. [cited by applicant]
Ramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, arXiv preprint arXiv:2204.06125v1 [cs.CV] Apr. 13, 2022, 27 pages. [cited by applicant]
Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, (2022), (pp. 10684-10695). [cited by applicant]
Meng, et al., “Sdedit: Guided Image Synthesis and Editing With Stochastic Differential Equations”, arXiv preprint arXiv:2108.01073v2 [cs.CV] Jan. 5, 2022, 33 pages. [cited by applicant]
Balaji, et al., “eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers”, arXiv preprint arXiv:2211.01324v5 [cs.CV] Mar. 14, 2023, 24 pages. [cited by applicant]
Dehouche, et al., “What's in a Text-to-Image Prompt? The Potential of Stable Diffusion in Visual Arts Education”, 2023, arXiv, p. 1-11, https://doi.org/10.48550/arXiv.2301.01902 (Year: 2023). [cited by applicant]
Ma, et al., “Text Style Transfer With Decorative Elements”, 2021, 2021 IEEE 4th International Conference on Multimedia Information Processing and Retrieval (MIPR), p. 330-336, https://doi.org/10.1145/3544903.3544906 (Ye… [cited by applicant]
Xie, et al., “SmartBrush: Text and Shape Guided Object Inpainting with Diffusion Model”, 2022, arXiv, p. 1-10, https://doi.org/10.48550/arXiv.2212.05034 (Year: 2022). [cited by applicant]
Office Action dated Jun. 2, 2025 in related U.S. Appl. No. 18/479,379. [cited by applicant]