IP Library › Granted Patent US 12,354,247
Granted Patent B1
US 12,354,247 · App. 18/886,390 · Granted Jul 8, 2025

Seamless image integration and image personalization

Inventors: Arnab Ghosh (London, GB); Mandela Patrick (Brooklyn, NY); Oleksii Sadliak (Lviv, UA); Roman Vei (Lviv, UA)
Assignee: Optimatik Inc.
G06T5/77G06T3/40G06T5/50G06T5/60G06T7/11G06T7/30G06T7/70G06V40/162G06T2207/20081G06T2207/20084G06T2207/20132G06T2207/20221G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,247
App. No.
18/886,390
Granted
Jul 8, 2025
Kind
B1
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating an inpainted image of a replacement individual in the place of a reference individual in a reference image using a metadata comparison. In one aspect, a system comprises receiving a first image comprising one or more reference individuals and a second image comprising one or more replacement individuals; obtaining replacement metadata and reference metadata; determining, based on the obtained replacement metadata and reference metadata, at least one portion of the first image for replacing corresponding with the one or more reference individuals; and generating a third image that comprises a modification of the first image wherein the at least one portion of the first image is replaced with replacement content, the replacement content being generated based on (i) the one or more replacement individuals, (ii) the replacement metadata, and (iii) the reference metadata.

Claims (64)

1. A computing-device implemented method comprising:

receiving a first image comprising one or more reference individuals and a second image comprising one or more replacement individuals;

obtaining replacement metadata for each of the one or more replacement individuals;

obtaining reference metadata for each of the one or more reference individuals;

determining, based on the obtained replacement metadata and reference metadata, at least one portion of the first image for replacing, the at least one portion corresponding with at least one portion of the one or more reference individuals; and

generating a third image that comprises a modification of the first image, wherein the at least one portion of the first image is replaced with replacement content, the replacement content being generated based on (i) the one or more replacement individuals, (ii) the replacement metadata, and (iii) the reference metadata, and a portion of the replacement content being iteratively personalized by decreasing a discrepancy between an embedding of the portion of the replacement content and an embedding of a corresponding portion of reference content of the at least one portion of the first image.

2. The computing-device implemented method of claim 1 , wherein determining the at least one portion of the first image for replacing comprises:

processing the first image using an object detection machine learning model to identify a region of the first image corresponding to each of the one or more reference individuals;

extracting a plurality of segmentation masks from each identified region of the first image;

performing a comparison of the replacement and reference metadata that relates to each identified region; and

identifying one or more of the plurality of segmentation masks from each identified region as the at least one portion of the first image for replacing in accordance with the comparison.

3. The computing-device implemented method of claim 2 , wherein the plurality of segmentation masks comprise segmentation masks representing face, skin, hair, and one or more articles of clothing.

4. The computing-device implemented method of claim 2 , wherein identifying one or more of the plurality of segmentation masks comprises identifying one or more segmentation masks based on at least one discrepancy between the replacement and reference metadata as the at least one portion of the first image for replacing.

5. The computing-device implemented method of claim 1 , wherein generating the third image comprises:

obtaining a cropped image from the second image comprising a replacement individual; and

generating the third image from a model input comprising the first image, the cropped image, the at least one portion of the first image for replacing, and a text prompt comprising the replacement metadata using an inpainting machine learning model.

6. The computing-device implemented method of claim 5 , wherein the at least one portion of the first image for replacing comprises one or more segmentation masks.

7. The computing-device implemented method of claim 5 , wherein the inpainting machine learning model is a stable diffusion machine learning model.

8. The computing-device implemented method of claim 5 , wherein obtaining the cropped image from the second image further comprises:

identifying a region of the second image comprising the replacement individual from the second image using an object detection machine learning model; and

cropping the region of the second image comprising the replacement individual to generate the cropped image.

9. The computing-device implemented method of claim 8 , further comprising resizing the cropped image in accordance with the at least one portion of the first image for replacing.

10. The computing-device implemented method of claim 1 , wherein generating the third image further comprises:

generating a pose for each of the one or more reference individuals from the first image using a pose estimation machine learning model; and

conditioning the generation of the third image using the pose.

11. The computing-device implemented method of claim 1 , wherein obtaining the replacement metadata of the one or more replacement individuals comprises generating the replacement metadata by processing the second image using a metadata machine learning model, and wherein obtaining the reference metadata of the one or more reference individuals comprises generating the reference metadata by processing the first image using the metadata machine learning model.

12. The computing-device implemented method of claim 1 , wherein the second image is selected from a set of example individual images, wherein the replacement individuals are example individuals, wherein the replacement metadata is template metadata, and wherein the third image comprises a generated template image.

13. The computing-device implemented method of claim 12 , further comprising:

generating a corresponding set of template images for each example individual image in the set of example individual images; and

storing the set of template images and the template metadata in a template image database.

14. The computing-device implemented method of claim 13 , wherein generating the corresponding set of template images further comprises upsampling each template image in the set of template images using an upsampling model.

15. The computing-device implemented method of claim 2 , wherein identifying one or more of the plurality of segmentation masks comprises extracting a reference face segmentation mask for each reference individual face.

16. The computing-device implemented method of claim 15 , further comprising:

extracting a replacement face segmentation mask for each replacement individual face; and

replacing a reference face with a corresponding replacement face using a respective pair of reference face segmentation mask and corresponding replacement segmentation mask for each reference individual.

17. A system comprising:

a computing device comprising:

a memory configured to store instructions; and

a processor to execute the instructions to perform operations comprising:

receiving a first image comprising one or more reference individuals and a second image comprising one or more replacement individuals;

obtaining replacement metadata for each of the one or more replacement individuals;

obtaining reference metadata for each of the one or more reference individuals;

determining, based on the obtained replacement metadata and reference metadata, at least one portion of the first image for replacing, the at least one portion corresponding with at least one portion of the one or more reference individuals; and

generating a third image that comprises a modification of the first image, wherein the at least one portion of the first image is replaced with replacement content, the replacement content being generated based on (i) the one or more replacement individuals, (ii) the replacement metadata, and (iii) the reference metadata, and a portion of the replacement content being iteratively personalized by decreasing a discrepancy between an embedding of the portion of the replacement content and an embedding of a corresponding portion of reference content of the at least one portion of the first image.

18. The system of claim 17 , wherein determining the at least one portion of the first image for replacing comprises:

processing the first image using an object detection machine learning model to identify a region of the first image corresponding to each of the one or more reference individuals;

extracting a plurality of segmentation masks from each identified region of the first image;

performing a comparison of the replacement and reference metadata that relates to each identified region; and

identifying one or more of the plurality of segmentation masks from each identified region as the at least one portion of the first image for replacing in accordance with the comparison.

19. The system of claim 18 , wherein identifying one or more of the plurality of segmentation masks comprises extracting a reference face segmentation mask for each reference individual face.

20. The system of claim 19 , further comprising:

extracting a replacement face segmentation mask for each replacement individual face; and

replacing a reference face with a corresponding replacement face using a respective pair of reference face segmentation mask and corresponding replacement segmentation mask for each reference individual.

21. One or more non-transitory computer readable media storing instructions that are executable by a processing device, and upon such execution cause the processing device to perform operations comprising:

receiving a first image comprising one or more reference individuals and a second image comprising one or more replacement individuals;

obtaining replacement metadata for each of the one or more replacement individuals;

obtaining reference metadata for each of the one or more reference individuals;

determining, based on the obtained replacement metadata and reference metadata, at least one portion of the first image for replacing, the at least one portion corresponding with at least one portion of the one or more reference individuals; and

generating a third image that comprises a modification of the first image, wherein the at least one portion of the first image is replaced with replacement content, the replacement content being generated based on (i) the one or more replacement individuals, (ii) the replacement metadata, and (iii) the reference metadata, and a portion of the replacement content being iteratively personalized by decreasing a discrepancy between an embedding of the portion of the replacement content and an embedding of a corresponding portion of reference content of the at least one portion of the first image.

22. The non-transitory computer readable media of claim 21 , wherein determining the at least one portion of the first image for replacing comprises:

processing the first image using an object detection machine learning model to identify a region of the first image corresponding to each of the one or more reference individuals;

extracting a plurality of segmentation masks from each identified region of the first image;

performing a comparison of the replacement and reference metadata that relates to each identified region; and

identifying one or more of the plurality of segmentation masks from each identified region as the at least one portion of the first image for replacing in accordance with the comparison.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2024
From: GHOSH, ARNAB; PATRICK, MANDELA; SADLIAK, OLEKSII; VEI, ROMAN
To: OPTIMATIK INC.
Reel/Frame 068598/0890 →
Continuity (1)
Provisional Application 63694975 · Sep 16, 2024
References Cited (19)
US 10049477B1 · Kokemohr · 2018 [cited by examiner]
US 20150143209A1 · Sudai · 2015 [cited by examiner]
US 20170287136A1 · Dsouza · 2017 [cited by examiner]
US 20180047200A1 · O'Hara · 2018 [cited by examiner]
US 20230230198A1 · Zhang · 2023 [cited by examiner]
US 20230342893A1 · Hinz · 2023 [cited by examiner]
US 20240355010A1 · Ahafonov · 2024 [cited by examiner]
US 20240355022A1 · Shi · 2024 [cited by examiner]
Zhang, Zhixing, et al. “Sine: Single image editing with text-to-image diffusion models.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023. [cited by examiner]
Murphy-Chutorian, Erik, and Mohan Manubhai Trivedi. “Head pose estimation in computer vision: A survey.” IEEE transactions on pattern analysis and machine intelligence 31.4 (2008): 607-626. [cited by examiner]
Wu, Weijia, et al. “Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023. [cited by examiner]
Avrahami, Omri, Dani Lischinski, and Ohad Fried. “Blended diffusion for text-driven editing of natural images.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022. [cited by examiner]
Kirillov et al., “Segment anything,” CoRR, Submitted Apr. 5, 2023, arXiv:2304.02643v1, 30 pages. [cited by applicant]
Niu et al., “Painterly Image Harmonization by Learning from Painterly Objects,” CoRR, Submitted Dec. 15, 2023, arXiv:2312.10263v1, 14 pages. [cited by applicant]
Radford et al., “Learning transferable visual models from natural language supervision,” CoRR, Submitted Feb. 26, 2021, arXiv:2103.00020v1, 48 pages. [cited by applicant]
Wang et al., “Towards Real-World Blind Face Restoration with Generative Facial Prior,” CoRR, Submitted Jun. 11, 2021, arXiv:2101.04061v2, 11 pages. [cited by applicant]
Xie et al., “SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers,” Advances in neural information processing systems, Submitted Oct. 28, 2021, arXiv:2105.15203v3, 18 pages. [cited by applicant]
Ye et al., “IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models,”, CoRR, Submitted Aug. 13, 2023, arXiv:2308.06721v1, 16 pages. [cited by applicant]
Zhang et. al., “Adding Conditional Control to Text-to-Image Diffusion Models,” CoRR, Submitted Nov. 26, 2023, arXiv:2302.05543v3, 12 pages. [cited by applicant]
Cited By (1)
US 12,725,498