IP Library › Granted Patent US 12,299,896
Granted Patent B1
US 12,299,896 · App. 17/702,202 · Granted May 13, 2025

Synthetic image generation by manifold modification of a machine learning model

Inventors: Qianli Feng (Seattle, WA); Raghu Deep Gadde (Bothell, WA); Pietro Perona (Altadena, CA); Aleix Margarit Martinez (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G06T7/187G06N3/045G06V10/761G06V10/7747G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,896
App. No.
17/702,202
Granted
May 13, 2025
Kind
B1
Abstract

Described herein is a computer-implemented method for generating a synthetic image. An input image can be received by a computing device. A representation of the input image on an image approximation manifold can be identified by inputting the input image into a machine learning model. The image approximation manifold can be defined by the machine learning model. A local region of the image approximation manifold can be modified relative to the first representation to generate a modified image approximation manifold. The modified image approximation manifold can include a second representation of the input image. A synthetic image can be generated based on the modified image approximation manifold. A rendering of the synthetic image can be caused on a display.

Claims (41)

1. One or more non-transitory computer-readable media comprising computer-executable instructions that, when executed by one or more processors of a computer system, cause the computer system to perform operations comprising:

receiving a selection of an input image;

inputting the input image into a generative adversarial network (GAN) generator to identify a first representation of the input image on an image approximation manifold, the image approximation manifold being defined by the GAN generator;

modifying a local region of the image approximation manifold relative to the first representation to generate a modified image approximation manifold, the modified image approximation manifold comprising a second representation of the input image, wherein modifying the local region of the image approximation manifold comprises:

evaluating a first loss function associated with similarity between the input image and the second representation;

evaluating a second loss function associated with global cohesion of the image approximation manifold;

generating a synthetic image based on the modified image approximation manifold; and

causing rendering of the synthetic image on a display.

2. The one or more non-transitory computer-readable media of claim 1 , wherein the first representation comprises a first latent vector that represents the input image and constitutes a first mapping between a latent space and the first representation on the image approximation manifold.

3. The one or more non-transitory computer-readable media of claim 2 , wherein the second representation represents the synthetic image and constitutes a second mapping between the latent space and the second representation on the modified image approximation manifold.

4. The one or more non-transitory computer-readable media of claim 1 , wherein the input image comprises a set of input image attributes and the synthetic image comprises a set of synthetic image attributes that correspond to the set of input image attributes.

5. The one or more non-transitory computer-readable media of claim 4 , further comprising additional computer-executable instructions that, when executed by the one or more processors of the computer system, cause the computer system to perform additional operations comprising modifying the synthetic image by removing a synthetic image attribute of the set of synthetic image attributes, or by adding a different image attribute not previously included in either the set of input image attributes or the set of synthetic image attributes.

6. A computer-implemented method, comprising:

receiving a selection of an input image;

inputting the input image into a machine learning model to identify a first representation of the input image on an image approximation manifold, the image approximation manifold being defined by the machine learning model;

modifying a local region of the image approximation manifold relative to the first representation to generate a modified image approximation manifold, the modified image approximation manifold comprising a second representation of the input image;

generating a synthetic image based on the modified image approximation manifold; and

causing rendering of the synthetic image on a display.

7. The computer-implemented method of claim 6 , wherein modifying the local region of the image approximation manifold comprises evaluating a loss function that is dependent on the machine learning model.

8. The computer-implemented method of claim 7 , wherein the loss function comprises a first loss function associated with similarity between the input image and the second representation, and a second loss function associated with global cohesion of the image approximation manifold.

9. The computer-implemented method of claim 8 , wherein the first loss function is dependent on a reconstruction loss function, and an adversarial loss function.

10. The computer-implemented method of claim 9 , wherein the reconstruction loss function accounts for visual differences between the synthetic image and the input image, and the adversarial loss function accounts for the synthetic image being editable.

11. The computer-implemented method of claim 6 , wherein the input image and the synthetic image are visually identical.

12. The computer-implemented method of claim 6 , further comprising modifying the synthetic image to include an image attribute that was not in the input image.

13. The computer-implemented method of claim 12 , wherein modifying the synthetic image comprises adjusting a latent vector that corresponds to the second representation.

14. The computer-implemented method of claim 6 , wherein the local region is modified to represent a context of the input image.

15. The computer-implemented method of claim 6 , wherein the machine learning model comprises a generative adversarial network (GAN) generator.

16. The computer-implemented method of claim 6 , wherein the synthetic image is editable using generative adversarial network (GAN) editing algorithms.

17. The computer-implemented method of claim 6 , wherein the first representation corresponds to an intermediate image closest to the input image on the image approximation manifold, wherein the second representation is closer to the input image in latent space than the intermediate image, and wherein modifying the image approximation manifold comprises:

evaluating a first loss function associated with similarity between the input image and the second representation, wherein the first loss function is dependent on a reconstruction loss function and an adversarial loss function; and

evaluating a second loss function associated with global cohesion of the image approximation manifold.

18. A system, comprising:

one or more memories configured to store computer-executable instructions;

one or more processors configured to access the one or more memories and execute the computer-executable instructions to at least:

receive a selection of an input image;

input the input image into a machine learning model to identify a first representation of the input image on an image approximation manifold, the image approximation manifold being defined by the machine learning model;

modify a local region of the image approximation manifold relative to the first representation to generate a modified image approximation manifold, the modified image approximation manifold comprising a second representation of the input image;

generate a synthetic image based on the modified image approximation manifold; and

cause rendering of the synthetic image on a display.

19. The system of claim 18 , wherein modifying the local region of the image approximation manifold comprises evaluating a loss function that is dependent on the machine learning model.

20. The system of claim 19 , wherein the loss function comprises a first loss function associated with similarity between the input image and the second representation, and a second loss function associated with global cohesion of the image approximation manifold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2022
From: FENG, QIANLI; GADDE, RAGHU DEEP; PERONA, PIETRO; MARTINEZ, ALEIX MARGARIT
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 059377/0776 →
References Cited (44)
US 20210038198A1 · Jacob · 2021 [cited by examiner]
US 20210319302A1 · Li · 2021 [cited by examiner]
US 20220122305A1 · Smith · 2022 [cited by examiner]
US 20220414451A1 · Gurev · 2022 [cited by examiner]
US 20230289608A1 · Zhong · 2023 [cited by examiner]
Abal, R. et al., “Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?”, 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 4431-4440, 2019. [cited by applicant]
Abal, R. et al., “Image2StyleGAN++: How to Edit the Embedded Images?”, CVPR, pp. 8293-8302, 2020. [cited by applicant]
Adelson, E.H et al., “Pyramid methods in image processing”, RCA engineer, 29(6):33-41, 1984. [cited by applicant]
Alaluf, Y. et al., “ReStyle: A Residual-Based StyleGAN Encoder via Iterative Refinement”, In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021. [cited by applicant]
Balakrishnan, G. et al., “Towards causal benchmarking of bias in face analysis algorithms”, arXiv e-prints, pp. arXiv-2007, 2020. [cited by applicant]
Bau, D et al., “Seeing What a GAN Cannot Generate”, ICCV, pp. 4501-4510, 2019. [cited by applicant]
Beery, S. et al., “Recognition in Terra Incognita”, In Proceedings of the European conference on computer vision (ECCV), pp. 456-473, 2018. [cited by applicant]
Bojanowski, P. et al., “Optimizing the Latent Space of Generative Networks”, In International Conference on Machine Learning, pp. 600-609. PMLR, 2018. [cited by applicant]
Brock, A. et al., “Large Scale Gan Training for High Fidelity Natural Image Synthesis”, In International Conference on Learning Representations, 35 pages, 2018. [cited by applicant]
Chai, L. et al., “Using Latent Space Regression To Analyze and Leverage Compositionality in Gans”, ICLR, 30 pages, Jun. 3, 2021. [cited by applicant]
Chai, L. et al., “Ensembling with Deep Generative Views”, In CVPR, 24 pages, Apr. 29, 2021. [cited by applicant]
Choi, Y. et al., “StarGAN v2: Diverse Image Synthesis for Multiple Domains”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8188-8197, 2020. [cited by applicant]
Creswell, A. et al., “Inverting The Generator Of A Generative Adversarial Network”, IEEE Transactions on Neural Networks and Learning Systems, 30:1967-1974, 2019. [cited by applicant]
Raghudeep, G. et al., “Detail Me More: Improving GAN's photo-realism of complex scenes”, In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 13950-13959, Oct. 2021. [cited by applicant]
Seo Jo, E. et al., “Lessons from Archives: Strategies for Collecting Sociocultural Data in Machine Learning”, In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pp. 306-316, 2020. [cited by applicant]
Kang, K. et al., “GAN Inversion for Out-of-Range Images with Geometric Transformations”, ICCV, 9 pahges, Aug. 20, 2021. [cited by applicant]
Karras, T. et al., “Progressive Growing of GANs for Improved Quality, Stability, and Variation”, In International Conference on Learning Representations, 26 pages, 2018. [cited by applicant]
Karras, T. et al., “Alias-Free Generative Adversarial Networks”, In Proc. NeurIPS, 31 pages, Oct. 18, 2021. [cited by applicant]
Karras, T. et al., “A Style-Based Generator Architecture for Generative Adversarial Networks”, arXiv:1812.04948 [cs.NE], 12 pages, Mar. 29, 2019. [cited by applicant]
Karras, T. et al., “A Style-Based Generator Architecture for Generative Adversarial Networks”, Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4401-4410, 2019. [cited by applicant]
Karras, T. et al., “Analyzing and Improving the Image Quality of StyleGAN”, In Proc. CVPR, 21 pages, Mar. 23, 2020. [cited by applicant]
Karras, T. et al., “Analyzing and Improving the Image Quality of StyleGAN”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8110-8119, 2020. [cited by applicant]
Krause, J. et al., “3D Object Representations for Fine-Grained Categorization”, 2013 IEEE International Conference on Computer Vision Workshops, pp. 554-561, 2013. [cited by applicant]
Lipton, Z.C. et al., “Precise Recovery of Latent Vectors From Generative Adversarial Networks”, ArXiv, abs/1702.04782, 4 pages, 2017. [cited by applicant]
Menon, S. et al., “PULSE: Self-Supervised Photo Upsampling via Latent Space Exploration of Generative Models”, In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2437-2445, 2020. [cited by applicant]
Pidhorskyi, S. et al., “Adversarial Latent Autoencoders”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14104-14113, 2020. [cited by applicant]
Raji, D.I. et al., “ Actionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products”, In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, pp.… [cited by applicant]
Richardson, E. et al., “Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation”, CVPR, 21 pages, Apr. 21, 2021. [cited by applicant]
Shen, Y. et al., “InterFaceGAN: Interpreting the Disentangled Face Representation Learned by GANs”, IEEE transactions on pattern analysis and machine intelligence, 16 pages, 2020. [cited by applicant]
Shen, Y. et al., “Closed-Form Factorization of Latent Semantics in GANs”, In CVPR, 9 pages, Apr. 3, 2021. [cited by applicant]
Simonyan, K. et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, In International Conference on Learning Representations, 14 pages, Apr. 10, 2015. [cited by applicant]
Torralba, A. et al., “Unbiased Look at Dataset Bias”, In CVPR 2011, pp. 1521-1528. IEEE, 2011. [cited by applicant]
Tzelepis, C. et al., “WarpedGANSpace: Finding non-linear RBF paths in GAN latent space”, In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 6393-6402, Oct. 2021. [cited by applicant]
Wang, T. et al., “High-Fidelity GAN Inversion for Image Attribute Editing”, arxiv:2109.06590, 22 pages, Sep. 15, 2021. [cited by applicant]
Wu, Z. et al., “StyleSpace Analysis: Disentangled Controls for StyleGAN Image Generation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12863-12872, 2021. [cited by applicant]
Xia, W. et al., “GAN Inversion: A Survey”, arXiv preprint arXiv:2101.05278, 21 pages, Aug. 13, 2021. [cited by applicant]
Yu, F. et al., “LSUN: Construction of a Large-Scale Image Dataset using Deep Learning with Humans in the Loop”, arXiv preprint arXiv:1506.03365,, 9 pages, Jun. 4, 2016. [cited by applicant]
Zhang, R. et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586-595, 2018. [cited by applicant]
Zhu, Jun-Yan et al., “Generative Visual Manipulation on the Natural Image Manifold”, arXiv:1609.03552 [cs. CV], 16 pages, Dec. 16, 2018. [cited by applicant]
Cited By (2)
US 12,541,818 US 12,694,604