IP Library › Granted Patent US 12,444,173
Granted Patent B2
US 12,444,173 · App. 18/177,084 · Granted Oct 14, 2025

Image component generation based on application of iterative learning on autoencoder model and transformer model

Inventors: Marzieh Edraki (San Jose, CA); Akira Nakamura (San Jose, CA)
Assignees: SONY GROUP CORPORATION; SONY CORPORATION OF AMERICA
G06V10/774G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,173
App. No.
18/177,084
Granted
Oct 14, 2025
Kind
B2
Abstract

An electronic device and method for image component generation based on application of iterative learning on autoencoder model and transformer model is provided. The electronic device fine-tunes, based on first training data including a first set of images, an autoencoder model and a transformer model. The autoencoder model includes an encoder model, a learned codebook, a generator model, and a discriminator model. The electronic device selects a subset of images from the first training data. The electronic device applies the encoder model on the selected subset of images. The electronic device generates second training data including a second set of images, based on the application of the encoder model. The generated second training data corresponds to a quantized latent representation of the selected subset of images. The electronic device pre-trains the autoencoder model to create a next generation of the autoencoder model, based on the generated second training data.

Claims (24)

1. An electronic device, comprising: circuitry configured to: fine-tune, based on first training data including a first set of images, an autoencoder model and a transformer model associated with the autoencoder model, wherein the autoencoder model includes an encoder model, a learned codebook associated with the transformer model, a generator model, and a discriminator model; select a subset of images from the first training data; apply the encoder model on the selected subset of images based on the learned codebook to determine encoded subset of images; generate second training data including a second set of images, based on the application of the encoder model, wherein the generated second training data corresponds to a quantized latent representation of the selected subset of images; and pre-train the autoencoder model to create a next generation of the autoencoder model, based on the generated second training data, wherein the next generation of the autoencoder model is used to generate specific training data for a next generation of the transformer model.

2. The electronic device according to claim 1 , wherein the circuitry is further configured to: apply the transformer model to predict a sequence of tokens for each of new synthetic images based on a start of the sequence of tokens; transform the predicted sequence of tokens to a quantized latent representation based on the learned codebook; apply the generator model on the quantized latent representation to generate a new synthetic image; generate third training data including a third set of images corresponding to the generated new synthetic image; and pre-train the transformer model to create the next generation of the transformer model, based on the generated third training data.

3. The electronic device according to claim 2 , wherein the predicted sequence of tokens corresponds to a sequence of indices from the learned codebook.

4. The electronic device according to claim 1 , wherein the fine-tuning of the autoencoder model and the transformer model, and the pre-training of the autoencoder model corresponds to an iterative learning model (ILM).

5. The electronic device according to claim 1 , wherein the autoencoder model corresponds to a convolutional neural network (CNN) model based on a vector quantized generative adversarial network (VQGAN).

6. The electronic device according to claim 1 , wherein the circuitry is further configured to map the selected subset of images from an image space to a signal space, based on an application of the encoder model on the selected subset of images.

7. The electronic device according to claim 6 , wherein the signal space corresponds to the learned codebook.

8. The electronic device according to claim 7 , wherein the quantized latent representation of the selected subset of images is determined based on a replacement of each vector, of a set of multi-dimensional code vectors associated with the selected subset of images, with a closest entry from the learned codebook.

9. The electronic device according to claim 1 , wherein the circuitry is further configured to:

determine a first loss function associated with the encoder model, the learned codebook, and the generator model;

determine a second loss function associated with the autoencoder model; and

determine a third loss function associated with the encoder model, wherein

the pre-training of the autoencoder model is further based on the determined first loss function, the determined second loss function, and the determined third loss function.

10. The electronic device according to claim 9 , wherein the determination of the third loss function is based on a second norm, associated with the encoder model of the next generation of the autoencoder model, with respect to the learned codebook.

11. An electronic device, comprising: circuitry configured to: fine-tune, based on first training data including a first set of images, an autoencoder model, wherein the autoencoder model includes an encoder model, a learned codebook associated with a transformer model, a generator model, and a discriminator model; apply the encoder model on the first set of images based on the learned codebook to determine encoded first set of images, wherein the encoded first set of images corresponds to a quantized latent representation of the first set of images; generate second training data including a second image dataset based on a subset of images from the first training data and the quantized latent representation of the subset of images; pre-train the autoencoder model to create a next generation of the autoencoder model, based on the generated second training data, wherein the next generation of the autoencoder model is used to generate specific training data for a next generation of the transformer model; and fine-tune the transformer model based on a last generation of the autoencoder model, wherein the last generation of the autoencoder model is previous to the next generation of the autoencoder model.

12. The electronic device according to claim 11 , wherein the fine-tuning of the autoencoder model, the pre-training of the autoencoder, and fine-tuning of the transformer model corresponds to an iterative learning model (ILM).

13. The electronic device according to claim 11 , wherein the autoencoder model corresponds to a convolutional neural network (CNN) model based on a vector quantized generative adversarial network (VQGAN).

14. The electronic device according to claim 11 , wherein the circuitry is further configured to map the first set of images from an image space to a signal space, based on an application of the encoder model on the first set of images.

15. The electronic device according to claim 14 , wherein the signal space corresponds to the learned codebook.

16. The electronic device according to claim 15 , wherein the quantized latent representation of the first set of images is determined based on a replacement of each vector, of a set of multi-dimensional code vectors associated with the first set of images, with a closest entry from the learned codebook.

17. An electronic device, comprising: circuitry configured to: fine-tune, based on first training data including a first set of images, an autoencoder model and a transformer model associated with the autoencoder model, wherein the autoencoder model includes an encoder model, a learned codebook associated with the transformer model, a generator model, and a discriminator model; apply the transformer model to predict a sequence of tokens for each of new synthetic images based on a start of the sequence of tokens; transform the predicted sequence of tokens to a quantized latent representation based on the learned codebook; apply the generator model on the quantized latent representation to generate a new synthetic image; generate third training data including a third set of images corresponding to the generated new synthetic image; and pre-train the transformer model to create a next generation of the transformer model, based on the generated third training data, wherein the next generation of the transformer model is used for the prediction in a subsequent training cycle.

18. The electronic device according to claim 17 , wherein the fine-tuning of the autoencoder model and the transformer model, and the pre-training of the transformer model corresponds to an iterative learning model (ILM).

19. A method, comprising: in an electronic device: fine-tuning, based on first training data including a first set of images, an autoencoder model and a transformer model associated with the autoencoder model, wherein the autoencoder model includes an encoder model, a learned codebook associated with the transformer model, a generator model, and a discriminator model; selecting a subset of images from the first training data; applying the encoder model on the selected subset of images based on the learned codebook to determine encoded subset of images; generating second training data including a second set of images, based on the application of the encoder model, wherein the generated second training data corresponds to a quantized latent representation of the selected subset of images; and pre-training the autoencoder model to create a next generation of the autoencoder model, based on the generated second training data, wherein the next generation of the autoencoder model is used to generate specific training data for a next generation of the transformer model.

20. The method according to claim 19 , further comprising: applying the transformer model to predict a sequence of tokens for each of new synthetic images based on a start of sequence of token; transforming the predicted sequence of tokens to a quantized latent representation based on the learned codebook; applying the generator model on the quantized latent representation to generate a new synthetic image; generating third training data including a third set of images corresponding to the generated new synthetic image; and pre-training the transformer model to create the next generation of the transformer model, based on the generated third training data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2023
From: EDRAKI, MARZIEH; NAKAMURA, AKIRA
To: SONY GROUP CORPORATION; SONY CORPORATION OF AMERICA
Reel/Frame 063194/0356 →
Continuity (2)
Provisional Application 63368264 · Jul 13, 2022
Related Publication 20240029411A1 · Jan 25, 2024
References Cited (15)
US 20220108183A1 · Arpit · 2022 [cited by applicant]
US 20230100413A1 · Zhu · 2023 [cited by examiner]
US 20230360294A1 · Aggarwal · 2023 [cited by examiner]
CN 113449135A · 2021 [cited by applicant]
JP 2020061023A · 2020 [cited by applicant]
JP 2022013136A · 2022 [cited by applicant]
WO WO2022125290A1 · 2022 [cited by applicant]
Mendez, et al., “How to Reuse and Compose Knowledge for a Lifetime of Tasks: A Survey on Continual Learning and Functional Composition”, ResearchGate.net, Jul. 15, 2022, 60 pages. [cited by applicant]
Rajeswar, et al., “Multi-label Iterated Learning for Image Classification with Label Ambiguity”, arxiv.org, Computer Vision and Pattern Recognition, Artificial Intelligence, Machine Learning, Nov. 23, 2021, 14 pages. [cited by applicant]
Ankit Vani, et al, “Iterated Learning for Emergent Systematicity in VQA”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, May 3, 2021 (May 3, 2021), pp. 1-21. [cited by applicant]
Chen, Mark, et al, “Generative Pretraining from Pixels”, Proceedings of the 37th International Conference On Machine Learning, [Online] Jul. 13, 2020 (Jul. 13, 2020), pp. 1-13. [cited by applicant]
Jiahui Yu, et al, “Vector-Quantized Image Modeling with Improved Vqgan”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jun. 5, 2022 (Jun. 5, 2022), pp. 1-17. [cited by applicant]
Patrick Esser, et al, “Taming Transformers for High-Resolution Image Synthesis”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Feb. 11, 2021 (Feb. 11, 2021), pp. 1-37. [cited by applicant]
Yuchen Lu, et al, “Supervised Seeded Iterated Learning for Interactive Language Learning”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 6, 2020 (Oct. 6, 2020), pp. 1-… [cited by applicant]
Cao, C. et al., “The Image Local Autoregressive Transformer”, Advances in Neural Information Processing Systems, 13 pages. [cited by applicant]