IP Library › Granted Patent US 12,314,862
Granted Patent B2
US 12,314,862 · App. 18/647,545 · Granted May 27, 2025

Optimizing generative networks via latent space regularizations

Inventor: Sheng Zhong (Santa Clara, CA)
Assignee: Agora Lab, Inc.
G06N3/084G06F18/217G06N20/00G06T3/4053G06T5/00G06V10/776G06V10/82G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,862
App. No.
18/647,545
Granted
May 27, 2025
Kind
B2
Abstract

A method for image generation based on a Generative AI Network. The Generative AI Network includes a generator and an encoder. The method includes determining, by the encoder, a first encoding E(Y) of a target image Y; generating, by the generator, a generated image G(Z) corresponding to the target image Y, wherein the generated image G(Z) is located in a close vicinity of a target neighborhood of the target image Y, and outputs of the generator are mapped, by the encoder, to a latent space adaptable to manipulate at least one characteristics of images generated by the Generative AI Network; and generating, by the encoder, a second encoding E(G(Z)) of the generated image G(Z) corresponding to the target image Y, wherein the first and second encodings E(Y) and E(G(Z)) map the target image Y and the generated image G(Z) to the latent space.

Claims (30)

1. A method for image generation based on a Generative AI Network, the Generative AI Network comprising a generator and an encoder, the method comprising:

determining, by the encoder, a first encoding E(Y) of a target image Y;

generating, by the generator, a generated image G(Z) corresponding to the target image Y, wherein the generated image G(Z) is located in a close vicinity of a target neighborhood of the target image Y, and outputs of the generator are mapped, by the encoder, to a latent space adaptable to manipulate at least one characteristics of images generated by the Generative AI Network; and

generating, by the encoder, a second encoding E(G(Z)) of the generated image G(Z) corresponding to the target image Y, wherein the first and second encodings E(Y) and E(G(Z)) map the target image Y and the generated image G(Z) to the latent space.

2. The method of claim 1 , wherein the encoder is trained to minimize the differences between the first and second encodings E(Y) and E(G(Z)), and the generator is trained by using the first and second encodings E(Y) and E(G(Z)) as part of a loss function of the generator.

3. The method of claim 1 , wherein a same set of weights are used for first encoding E(Y) of the target image Y and the second encoding E(G(Z)) of the generated image G(Z) corresponding to the target image Y, and the same set of weights are updated after one or more iterations of training the Generative AI Network.

4. The method of claim 1 , wherein the first encoding E(Y) comprises a compressed representation of the target image Y in the latent space, and the second encoding E(G(Z)) comprises a compressed representation of the generated image G(Z) corresponding to the target image Y in the latent space.

5. The method of claim 4 , wherein the compressed representation of the target image Y in the latent space comprises the most relevant feature extracted from the target image Y, and the compressed representation of the generated image G(Z) corresponding to the target image Y in the latent space comprises the most relevant feature extracted from the generated image G(Z).

6. The method of claim 1 , wherein the encoder implements a machine learning model.

7. The method of claim 1 , wherein the encoder implements a convolutional neural network.

8. The method of claim 1 , wherein the encoder implements a VGG network.

9. The method of claim 1 , wherein the generator implements an inverse convolutional network.

10. The method of claim 1 , wherein a first loss function between the first encoding E(Y) and the second encoding E(G(Z)) is minimized for the encoder; and

the generator is trained using the first encoding E(Y) and the second encoding E(G(Z)) as part of a second loss function for the generator.

11. An apparatus comprising:

at least one processor; and

at least one memory, wherein the at least one memory comprises instructions which, when executed by the at least one processor, cause the at least one processor to implement a Generative AI Network, the Generative AI Network comprising a generator and an encoder and to train the Generative AI Network by:

determining, by the encoder, a first encoding E(Y) of a target image Y;

generating, by the generator, a generated image G(Z) corresponding to the target image Y, wherein the generated image G(Z) is located in a close vicinity of a target neighborhood of the target image Y, and outputs of the generator are mapped, by the encoder, to a latent space adaptable to manipulate at least one characteristics of images generated by the Generative AI Network; and

generating, by the encoder, a second encoding E(G(Z)) of the generated image G(Z) corresponding to the target image Y, wherein the first and second encodings E(Y) and E(G(Z)) map the target image Y and the generated image G(Z) to the latent space.

12. The apparatus of claim 11 , the encoder is trained to minimize the differences between the first and second encodings E(Y) and E(G(Z)), and the generator is trained by using the first and second encodings E(Y) and E(G(Z)) as part of a loss function of the generator.

13. The apparatus of claim 11 , wherein a same set of weights are used for first encoding E(Y) of the target image Y and the second encoding E(G(Z)) of the generated image G(Z) corresponding to the target image Y, and the same set of weights are updated after one or more iterations of training the Generative AI Network.

14. The apparatus of claim 11 , wherein the first encoding E(Y) comprises a compressed representation of the target image Y in the latent space, and the second encoding E(G(Z)) comprises a compressed representation of the generated image G(Z) corresponding to the target image Y in the latent space.

15. The apparatus of claim 14 , wherein the compressed representation of the target image Y in the latent space comprises the most relevant feature extracted from the target image Y, and the compressed representation of the generated image G(Z) corresponding to the target image Y in the latent space comprises the most relevant feature extracted from the generated image G(Z).

16. The apparatus of claim 11 , wherein the encoder implements a machine learning model.

17. The apparatus of claim 11 , wherein the encoder implements a convolutional neural network.

18. The apparatus of claim 11 , wherein the encoder implements a VGG network.

19. The apparatus of claim 11 , wherein the generator implements an inverse convolutional network.

20. The apparatus of claim 11 , wherein a first loss function between the first encoding E(Y) and the second encoding E(G(Z)) is minimized for the encoder; and

the generator is trained using the first encoding E(Y) and the second encoding E(G(Z)) as part of a second loss function for the generator.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2024
From: ZHONG, SHENG
To: AGORA LAB, INC.
Reel/Frame 067246/0266 →
Continuity (5)
Continuation 18319109 · May 17, 2023
Continuation 17324831 · May 19, 2021
Continuation 16530692 · Aug 2, 2019
Provisional Application 62840635 · Apr 30, 2019
Related Publication 20240296332A1 · Sep 5, 2024
References Cited (12)
US 10713294B2 · Kim · 2020 [cited by examiner]
US 20180075581A1 · Shi et al. · 2018 [cited by applicant]
US 20190251721A1 · Hua et al. · 2019 [cited by applicant]
US 20190295302A1 · Fu et al. · 2019 [cited by applicant]
US 20200134415A1 · Haidar · 2020 [cited by examiner]
US 20200160153A1 · Elmoznino et al. · 2020 [cited by applicant]
US 20200342306A1 · Giovannini · 2020 [cited by examiner]
US 20200349447A1 · Zhong et al. · 2020 [cited by applicant]
US 20200356810A1 · Zhong et al. · 2020 [cited by applicant]
Alice Lucas et al.; Generative Adversarial Networks and Perceptual Losses for Video Super-Resolution; pp. 1-16, Feb. 19, 2021. [cited by applicant]
Ledig; Photo-Realistic Single Image Super-Resolution Using A Generative Adversarial Network, arxiv.org/pdf/1609.04802v1.pdf Sep. 15, 2016. [cited by applicant]
Simonyan; Very Deep Convolutional Networks For Large-Scale Image Recognition ICLR 2015; arxiv.org/pdf/1409.1556.pdf; Apr. 10, 2015. [cited by applicant]