IP Library › Granted Patent US 12,670,627
Granted Patent B2
US 12,670,627 · App. 18/275,048 · Granted Jun 30, 2026

Generating images using sparse representations

Inventors: Charlie Thomas Curtis Nash (London, GB); Peter William Battaglia (London, GB)
Assignee: GDM Holding LLC
G06T9/002G06T7/11G06T7/73G06T2207/10024G06T2207/20021G06T2207/20052G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,627
App. No.
18/275,048
Filed
Jul 31, 2023
Granted
Jun 30, 2026
Kind
B2
Examiner
WU, MING HAN
Art Unit
2618
USPC
345/619
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating compressed representations of synthetic images. One of the methods is a method of generating a synthetic image using a generative neural network, and includes: generating, using the generative neural network, a plurality of coefficients that represent the synthetic image after the synthetic image has been encoded using a lossy compression algorithm; and decoding the synthetic image by applying the lossy compression algorithm to the plurality of coefficients.

Claims (58)

1 . A method of processing an input sequence using a generative neural network to generate an output sequence having a plurality of output elements, wherein:

the generative neural network comprises an encoder subnetwork and a sequence of a plurality of decoder subnetworks,

the encoder subnetwork comprises one or more self-attention layers,

each decoder subnetwork in the sequence of decoder subnetworks comprises i) one or more self-attention layers and ii) one or more encoder-decoder self-attention layers,

the method comprising:

processing an encoder subnetwork input comprising the input sequence using the encoder subnetwork to generate an embedding of the input sequence; and

at each of a plurality of time steps:

processing, using the first decoder subnetwork in the sequence of decoder subnetworks, a first decoder subnetwork input generated from i) the embedding of the input sequence and ii) a plurality of output tokens generated at previous time steps to generate a first decoder subnetwork output;

for each subsequent decoder subnetwork in the sequence of decoder subnetworks:

processing, using the subsequent decoder subnetwork, a subsequent decoder subnetwork input generated from i) the embedding of the input sequence and ii) the decoder subnetwork output generated by the previous decoder subnetwork in the sequence of decoder subnetworks to generate a subsequent decoder subnetwork output; and

generating a new output element in the output sequence using the decoder subnetwork outputs.

2 . The method of claim 1 , wherein the output elements in the output sequence represent an output image.

3 . The method of claim 2 , further comprising:

generating the output image from the output elements in the output sequence.

4 . The method of claim 3 , wherein the output elements in the output sequence specify a plurality of coefficients of a lossy compression algorithm; and

wherein generating the output image comprises decoding the output image by applying the lossy compression algorithm to the plurality of coefficients.

5 . The method of claim 3 , wherein each output element represents a respective coefficient, and:

the first decoder subnetwork output defines a channel of the respective coefficient, and the one or more subsequent decoder subnetwork outputs define a position and a value of the respective coefficient.

6 . The method of claim 1 , wherein, for the first decoder subnetwork, the embedding of the input sequence is processed using the one or more cross-attention layers of the first decoder subnetwork and an input generated from the plurality of output tokens generated at previous time steps is processed using the one or more self-attention layers of the first decoder subnetwork.

7 . The method of claim 1 , wherein, for each subsequent decoder subnetwork, the embedding of the input sequence is processed using the one or more cross-attention layers of the subsequent decoder subnetwork and an input generated from the decoder subnetwork output generated by the previous decoder subnetwork in the sequence is processed using the one or more self-attention layers of the subsequent decoder subnetwork.

8 . The method of claim 1 , wherein each output element has a respective first value defined by the first decoder subnetwork output and one or more additional values that are each defined by a respective subsequent decoder subnetwork output generated by one of the subsequent decoder subnetworks.

9 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations for processing an input sequence using a generative neural network to generate an output sequence having a plurality of output elements, wherein:

the generative neural network comprises an encoder subnetwork and a sequence of a plurality of decoder subnetworks,

the encoder subnetwork comprises one or more self-attention layers,

each decoder subnetwork in the sequence of decoder subnetworks comprises i) one or more self-attention layers and ii) one or more encoder-decoder self-attention layers,

the operations comprising:

processing an encoder subnetwork input comprising the input sequence using the encoder subnetwork to generate an embedding of the input sequence; and

at each of a plurality of time steps:

processing, using the first decoder subnetwork in the sequence of decoder subnetworks, a first decoder subnetwork input generated from i) the embedding of the input sequence and ii) a plurality of output tokens generated at previous time steps to generate a first decoder subnetwork output;

for each subsequent decoder subnetwork in the sequence of decoder subnetworks:

processing, using the subsequent decoder subnetwork, a subsequent decoder subnetwork input generated from i) the embedding of the input sequence and ii) the decoder subnetwork output generated by the previous decoder subnetwork in the sequence of decoder subnetworks to generate a subsequent decoder subnetwork output; and

generating a new output element in the output sequence using the decoder subnetwork outputs.

10 . The system of claim 9 , wherein the output elements in the output sequence represent an output image.

11 . The system of claim 10 , the operations further comprising:

generating the output image from the output elements in the output sequence.

12 . The system of claim 11 , wherein the output elements in the output sequence specify a plurality of coefficients of a lossy compression algorithm; and

wherein generating the output image comprises decoding the output image by applying the lossy compression algorithm to the plurality of coefficients.

13 . The system of claim 12 , wherein each output element represents a respective coefficient, and:

the first decoder subnetwork output defines a channel of the respective coefficient, and the one or more subsequent decoder subnetwork outputs define a position and a value of the respective coefficient.

14 . The system of claim 9 , wherein, for the first decoder subnetwork, the embedding of the input sequence is processed using the one or more cross-attention layers of the first decoder subnetwork and an input generated from the plurality of output tokens generated at previous time steps is processed using the one or more self-attention layers of the first decoder subnetwork.

15 . The system of claim 9 , wherein, for each subsequent decoder subnetwork, the embedding of the input sequence is processed using the one or more cross-attention layers of the subsequent decoder subnetwork and an input generated from the decoder subnetwork output generated by the previous decoder subnetwork in the sequence is processed using the one or more self-attention layers of the subsequent decoder subnetwork.

16 . The system of claim 9 , wherein each output element has a respective first value defined by the first decoder subnetwork output and one or more additional values that are each defined by a respective subsequent decoder subnetwork output generated by one of the subsequent decoder subnetworks.

17 . One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for processing an input sequence using a generative neural network to generate an output sequence having a plurality of output elements, wherein:

the generative neural network comprises an encoder subnetwork and a sequence of a plurality of decoder subnetworks,

the encoder subnetwork comprises one or more self-attention layers,

each decoder subnetwork in the sequence of decoder subnetworks comprises i) one or more self-attention layers and ii) one or more encoder-decoder self-attention layers,

the operations comprising:

processing an encoder subnetwork input comprising the input sequence using the encoder subnetwork to generate an embedding of the input sequence; and

at each of a plurality of time steps:

processing, using the first decoder subnetwork in the sequence of decoder subnetworks, a first decoder subnetwork input generated from i) the embedding of the input sequence and ii) a plurality of output tokens generated at previous time steps to generate a first decoder subnetwork output;

for each subsequent decoder subnetwork in the sequence of decoder subnetworks:

processing, using the subsequent decoder subnetwork, a subsequent decoder subnetwork input generated from i) the embedding of the input sequence and ii) the decoder subnetwork output generated by the previous decoder subnetwork in the sequence of decoder subnetworks to generate a subsequent decoder subnetwork output; and

generating a new output element in the output sequence using the decoder subnetwork outputs.

18 . The computer storage media of claim 17 , wherein the output elements in the output sequence represent an output image.

19 . The computer storage media of claim 18 , the operations further comprising:

generating the output image from the output elements in the output sequence.

20 . The computer storage media of claim 19 , wherein the output elements in the output sequence specify a plurality of coefficients of a lossy compression algorithm; and

wherein generating the output image comprises decoding the output image by applying the lossy compression algorithm to the plurality of coefficients.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2023
From: NASH, CHARLIE THOMAS CURTIS; BATTAGLIA, PETER WILLIAM
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 065030/0531 →
Continuity (2)
Provisional Application 63146474 · Feb 5, 2021
Related Publication 20240104785A1 · Mar 28, 2024
References Cited (62)
US 9813721B2 · Schulze · 2017 [cited by examiner]
US 10058290B1 · Proud · 2018 [cited by examiner]
US 11375242B1 · Said · 2022 [cited by examiner]
US 20180284872A1 · Schluessler · 2018 [cited by examiner]
US 20190130213A1 · Shazeer et al. · 2019 [cited by applicant]
CN 112218094 · 2021 [cited by applicant]
Ahmed et al., “Discrete Cosine Transform,” IEEE Transactions on Computers, Jan. 1974, pp. 90-93. [cited by applicant]
Bachlechner et al., “ReZero is All You Need: Fast Convergence at Large Depth,” CoRR, Submitted on Mar. 10, 2020, arXiv:2003.04887v1, 11 pages. [cited by applicant]
Brock et al., “Large Scale GAN Training for High Fidelity Natural Image Synthesis,” ICLR 2019, 2019, 35 pages. [cited by applicant]
Brown et al., “Language Models are Few-Shot Learners,” 34th Conference on Neural Information Processing Systems (NeurIPS 2020), 25 pages. [cited by applicant]
Chen et al., “Generative Pretraining From Pixels,” Proceedings of the 37th International Conference on Machine Learning, PMLR 119, 2020, 13 pages. [cited by applicant]
Child et al., “Generating Long Sequences with Sparse Transformers,” CoRR, Submitted on Apr. 23, 2019, arXiv:1904.10509v1, 10 pages. [cited by applicant]
Choromanski et al., “Rethinking Attention with Performers,” ICLR 2021, 2021, 38 pages. [cited by applicant]
cloud.google.com [online], “Accelerate AI development with Google Cloud TPUs,” available on or before May 17, 2017 via Internet Archive: Wayback Machine URL<https://web.archive.org/web/20170517174135/https://cloud.googl… [cited by applicant]
data.mendeley.com [online], “A Database of Leaf Images: Practice towards Plant Conservation with Plant Pathology,” Jun. 6, 2019, retrieved on Jul. 26, 2024, retrieved from URL< https://data.mendeley.com/datasets/hb74ynk… [cited by applicant]
De et al., “Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks,” Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS 2020), 2020, 12 pages. [cited by applicant]
Dhariwal et al., “Jukebox: A Generative Model for Music,” CoRR, Submitted on Apr. 30, 2020, arXiv:2005.00341v1, 20 pages. [cited by applicant]
Du et al., “Implicit Generation and Modeling with Energy Based Models,” Proceedings of the 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019, 11 pages. [cited by applicant]
GitHub.com [online], “jantic / DeOldify,” available on or before Nov. 2, 2018, via Internet Archive: Wayback Machine URL<http://web.archive.org/web/20200106043047/https://github.com/jantic/DeOldify>, retrieved on Jul. 2… [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets,” Proceedings of the Conference on Neural Information Processing Systems (NIPS 2014), 9 pages. [cited by applicant]
Han et al., “not-so-BigGAN: Generating High-Fidelity Images on Small Compute with Wavelet-based Super-Resolution,” CoRR, Submitted on Oct. 25, 2020, arXiv:2009.04433v1, 34 pages. [cited by applicant]
Henighan et al., “Scaling Laws for Autoregressive Generative Modeling,” CoRR, Submitted on Oct. 28, 2020, arXiv:2010.14701v1, 35 pages. [cited by applicant]
Heusel et al., “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium,” Proceedings of the 31st Conference on Neurla Information Processing Systems (NIPS 2017), 12 pages. [cited by applicant]
ijg.org [online], “Independent JPEG Group,” available on or before Feb. 3, 2020, via Internet Archive: Wayback Machine URL<https://web.archive.org/web/20200203051208/https://ijg.org/>, retrieved on Jul. 25, 2024, retrie… [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/EP2022/052895, Aug. 17, 2023, 10 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/EP2022/052895, Jun. 3, 2022, 17 pages. [cited by applicant]
Johnson et al., “CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2901-2910. [cited by applicant]
kaggle.com [online], “Diabetic Retinopathy Detection,” Jul. 2015, retrieved on Jul. 25, 2024, retrieved from URL< https://www.kaggle.com/c/diabetic-retinopathy-detection/data>, 2 pages. [cited by applicant]
Kaplan et al., “Scaling Laws for Neural Language Models,” CoRR, Submitted on Jan. 23, 2020, arXiv:2001.08361v1, 30 pages. [cited by applicant]
Karras et al., “A Style-Based Generator Architecture for Generative Adversarial Networks,” CoRR, Submitted on Dec. 12, 2018, arXiv: 1812.04948v1, 12 pages. [cited by applicant]
Karras et al., “Analyzing and Improving the Image Quality of StyleGAN,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 8110-8119. [cited by applicant]
Karras et al., “Progressive Growing of GANs for Improved Quality, Stability, and Variation,” ICLR 2018, 2018, 26 pages. [cited by applicant]
Kersten, “Predictability and redundancy of natural images,” J Opt Soc Am A, Dec. 1987, 4(12):2395-2400. [cited by applicant]
Kingma et al., “Adam: A Method for Stochastic Optimization,” CoRR, Submitted on Jul. 20, 2015, arXiv:1412.6980v7, 13 pages. [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes,” CoRR, Submitted on Apr. 10, 2014, arXiv:1312.6114v9, 14 pages. [cited by applicant]
Kingma et al., “Glow: Generative Flow with Invertible 1x1 Convolutions,” CoRR, Submitted on Jul. 9, 2018, arXiv:1807.03039v1, 15 pages. [cited by applicant]
Kitaev et al., “Reformer: The Efficient Transformer,” ICLR 2020, 12 pages. [cited by applicant]
Kuznetsova et al., “The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scal,” CoRR, Submitted on Nov. 2, 2018, arXiv:1811.00982v1, 20 pages. [cited by applicant]
Kynkäänniemi et al., “Improved Precision and Recall Metric for Assessing Generative Models,” 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019, 10 pages. [cited by applicant]
Luo et al., “Deep Wavelet Network with Domain Adaptation for Single Image Demoireing”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), IEEE, Jun. 14, 2020, pp. 1687-1694. [cited by applicant]
Mandava et al., “Pay Attention when Required,” CoRR, Submitted on Sep. 9, 2020, arXiv:2009.04534v1, 9 pages. [cited by applicant]
Menick et al., “Generating High Fidelity Images with Subscale Pixel Networks and Multidimensional Upscaling,” CoRR, Submitted on Dec. 4, 2018, arXiv:1812.01608v1, 15 pages. [cited by applicant]
Nash et al., “Generating Images with Sparse Representations,” CoRR, Submitted on Mar. 5, 2021, arXiv:2103.03841v1, 18 pages. [cited by applicant]
nvidia.com [online], “NVIDIA DLSS 2.0: A Big Leap In AI Rendering,” Mar. 23, 2020, retrieved on Jul. 26, 2024, retrieved from URL< https://www.nvidia.com/en-gb/geforce/news/nvidia-dlss-2-0-a-big-leap-in-ai-rendering/>, … [cited by applicant]
openai.com [online], “DALL-E: Creating images from text,” Jan. 5, 2021, retrieved on Jul. 26, 2024, retrieved from URL< https://openai.com/index/dall-e/>, 13 pages. [cited by applicant]
Parisotto et al., “Stabilizing Transformers for Reinforcement Learning,” Proceedings of the 37th International Conference on Machine Learning, PMLR 119, 2020, 12 pages. [cited by applicant]
Razavi et al., “Generating Diverse High-Fidelity Images with VQ-VAE-2,” 33rd Conference on Information Processing Systems (NeurIPS 2019), 2019, 11 pages. [cited by applicant]
Rezende et al., “Stochastic Backpropagation and Approximate Inference in Deep Generative Models,” Proceedings of the 31st International Conference on Machine Learning, 2014, JMLR:W&CP vol. 32, 9 pages. [cited by applicant]
Rezende et al., “Variational Inference with Normalizing Flows,” Proceedings of the 32nd International Conference on Machine Learning, 2015, JMLR: W&CP vol. 37, 9 pages. [cited by applicant]
Roy et al., “Efficient Content-Based Sparse Attention with Routing Transformers,” CORR, Submitted on Mar. 12, 2020, arXiv:2003.05997v1, 11 pages. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” CoRR, Submitted on Sep. 1, 2014, arXiv:1409.0575v1, 37 pages. [cited by applicant]
Shannon, “Prediction and Entropy of Printed English,” Bell System Technical Journal, Bell system technical journal, Jan. 1951, 30(1):50-64. [cited by applicant]
Szegedy et al., “Rethinking the Inception Architecture for Computer Vision,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 2818-2826. [cited by applicant]
Van den Oord et al., “Neural Discrete Representation Learning,” CoRR, Submitted on Nov. 2, 2017, arXiv:1711.00937v1, 10 pages. [cited by applicant]
Van den Oord et al., “Pixel Recurrent Neural Networks,” Proceedings of the 33rd International Conference on Machine Learning, 2016, JMLR:W&CVP vol. 48, 10 pages. [cited by applicant]
Vaswani et al., “Attention Is All You Need,” CoRR, Submitted on Jun. 12, 2017, arXiv:1706.03762v1, 15 pages. [cited by applicant]
Wallace, “The JPEG still picture compression standard,” IEEE Transactions on Consumer Electronics, 1992, 17 pages. [cited by applicant]
Wang et al., “Image quality assessment: from error visibility to structural similarity,” IEEE Transactions on Image Processing, Apr. 2004, vol. 13, No. 4, 14 pages. [cited by applicant]
Yu et al., “Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop,” CoRR, Submitted on Jun. 10, 2015, arXiv: 1506.03365v1, 9 pages. [cited by applicant]
Yu et al., “LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop,” CoRR, Submitted on Jun. 19, 2015, arXiv:1506.03365v2, 9 pages. [cited by applicant]
Zaheer et al., “Big bird: transformers for longer sequences,” Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Dec. 2020, 15 pages. [cited by applicant]
Office Action in European Appln. No. 22708041.3, mailed on Oct. 15, 2025, 9 pages. [cited by applicant]