IP Library Granted Patent US 12670627
Granted Patent B2
US 12670627 · App. 18/275,048 · Granted Jun 30, 2026

Generating images using sparse representations

Inventors: Charlie Thomas Curtis Nash (London, GB); Peter William Battaglia (London, GB)
Assignee: GDM Holding LLC
G06T9/002G06T7/11G06T7/73G06T2207/10024G06T2207/20021G06T2207/20052G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670627
App. No.
18/275,048
Granted
Jun 30, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating compressed representations of synthetic images. One of the methods is a method of generating a synthetic image using a generative neural network, and includes: generating, using the generative neural network, a plurality of coefficients that represent the synthetic image after the synthetic image has been encoded using a lossy compression algorithm; and decoding the synthetic image by applying the lossy compression algorithm to the plurality of coefficients.

Claims (58)

1 . A method of processing an input sequence using a generative neural network to generate an output sequence having a plurality of output elements, wherein:

the generative neural network comprises an encoder subnetwork and a sequence of a plurality of decoder subnetworks,

the encoder subnetwork comprises one or more self-attention layers,

each decoder subnetwork in the sequence of decoder subnetworks comprises i) one or more self-attention layers and ii) one or more encoder-decoder self-attention layers,

the method comprising:

processing an encoder subnetwork input comprising the input sequence using the encoder subnetwork to generate an embedding of the input sequence; and

at each of a plurality of time steps:

processing, using the first decoder subnetwork in the sequence of decoder subnetworks, a first decoder subnetwork input generated from i) the embedding of the input sequence and ii) a plurality of output tokens generated at previous time steps to generate a first decoder subnetwork output;

for each subsequent decoder subnetwork in the sequence of decoder subnetworks:

processing, using the subsequent decoder subnetwork, a subsequent decoder subnetwork input generated from i) the embedding of the input sequence and ii) the decoder subnetwork output generated by the previous decoder subnetwork in the sequence of decoder subnetworks to generate a subsequent decoder subnetwork output; and

generating a new output element in the output sequence using the decoder subnetwork outputs.

2 . The method of claim 1 , wherein the output elements in the output sequence represent an output image.

3 . The method of claim 2 , further comprising:

generating the output image from the output elements in the output sequence.

4 . The method of claim 3 , wherein the output elements in the output sequence specify a plurality of coefficients of a lossy compression algorithm; and

wherein generating the output image comprises decoding the output image by applying the lossy compression algorithm to the plurality of coefficients.

5 . The method of claim 3 , wherein each output element represents a respective coefficient, and:

the first decoder subnetwork output defines a channel of the respective coefficient, and the one or more subsequent decoder subnetwork outputs define a position and a value of the respective coefficient.

6 . The method of claim 1 , wherein, for the first decoder subnetwork, the embedding of the input sequence is processed using the one or more cross-attention layers of the first decoder subnetwork and an input generated from the plurality of output tokens generated at previous time steps is processed using the one or more self-attention layers of the first decoder subnetwork.

7 . The method of claim 1 , wherein, for each subsequent decoder subnetwork, the embedding of the input sequence is processed using the one or more cross-attention layers of the subsequent decoder subnetwork and an input generated from the decoder subnetwork output generated by the previous decoder subnetwork in the sequence is processed using the one or more self-attention layers of the subsequent decoder subnetwork.

8 . The method of claim 1 , wherein each output element has a respective first value defined by the first decoder subnetwork output and one or more additional values that are each defined by a respective subsequent decoder subnetwork output generated by one of the subsequent decoder subnetworks.

9 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations for processing an input sequence using a generative neural network to generate an output sequence having a plurality of output elements, wherein:

the generative neural network comprises an encoder subnetwork and a sequence of a plurality of decoder subnetworks,

the encoder subnetwork comprises one or more self-attention layers,

each decoder subnetwork in the sequence of decoder subnetworks comprises i) one or more self-attention layers and ii) one or more encoder-decoder self-attention layers,

the operations comprising:

processing an encoder subnetwork input comprising the input sequence using the encoder subnetwork to generate an embedding of the input sequence; and

at each of a plurality of time steps:

processing, using the first decoder subnetwork in the sequence of decoder subnetworks, a first decoder subnetwork input generated from i) the embedding of the input sequence and ii) a plurality of output tokens generated at previous time steps to generate a first decoder subnetwork output;

for each subsequent decoder subnetwork in the sequence of decoder subnetworks:

processing, using the subsequent decoder subnetwork, a subsequent decoder subnetwork input generated from i) the embedding of the input sequence and ii) the decoder subnetwork output generated by the previous decoder subnetwork in the sequence of decoder subnetworks to generate a subsequent decoder subnetwork output; and

generating a new output element in the output sequence using the decoder subnetwork outputs.

10 . The system of claim 9 , wherein the output elements in the output sequence represent an output image.

11 . The system of claim 10 , the operations further comprising:

generating the output image from the output elements in the output sequence.

12 . The system of claim 11 , wherein the output elements in the output sequence specify a plurality of coefficients of a lossy compression algorithm; and

wherein generating the output image comprises decoding the output image by applying the lossy compression algorithm to the plurality of coefficients.

13 . The system of claim 12 , wherein each output element represents a respective coefficient, and:

the first decoder subnetwork output defines a channel of the respective coefficient, and the one or more subsequent decoder subnetwork outputs define a position and a value of the respective coefficient.

14 . The system of claim 9 , wherein, for the first decoder subnetwork, the embedding of the input sequence is processed using the one or more cross-attention layers of the first decoder subnetwork and an input generated from the plurality of output tokens generated at previous time steps is processed using the one or more self-attention layers of the first decoder subnetwork.

15 . The system of claim 9 , wherein, for each subsequent decoder subnetwork, the embedding of the input sequence is processed using the one or more cross-attention layers of the subsequent decoder subnetwork and an input generated from the decoder subnetwork output generated by the previous decoder subnetwork in the sequence is processed using the one or more self-attention layers of the subsequent decoder subnetwork.

16 . The system of claim 9 , wherein each output element has a respective first value defined by the first decoder subnetwork output and one or more additional values that are each defined by a respective subsequent decoder subnetwork output generated by one of the subsequent decoder subnetworks.

17 . One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for processing an input sequence using a generative neural network to generate an output sequence having a plurality of output elements, wherein:

the generative neural network comprises an encoder subnetwork and a sequence of a plurality of decoder subnetworks,

the encoder subnetwork comprises one or more self-attention layers,

each decoder subnetwork in the sequence of decoder subnetworks comprises i) one or more self-attention layers and ii) one or more encoder-decoder self-attention layers,

the operations comprising:

processing an encoder subnetwork input comprising the input sequence using the encoder subnetwork to generate an embedding of the input sequence; and

at each of a plurality of time steps:

processing, using the first decoder subnetwork in the sequence of decoder subnetworks, a first decoder subnetwork input generated from i) the embedding of the input sequence and ii) a plurality of output tokens generated at previous time steps to generate a first decoder subnetwork output;

for each subsequent decoder subnetwork in the sequence of decoder subnetworks:

processing, using the subsequent decoder subnetwork, a subsequent decoder subnetwork input generated from i) the embedding of the input sequence and ii) the decoder subnetwork output generated by the previous decoder subnetwork in the sequence of decoder subnetworks to generate a subsequent decoder subnetwork output; and

generating a new output element in the output sequence using the decoder subnetwork outputs.

18 . The computer storage media of claim 17 , wherein the output elements in the output sequence represent an output image.

19 . The computer storage media of claim 18 , the operations further comprising:

generating the output image from the output elements in the output sequence.

20 . The computer storage media of claim 19 , wherein the output elements in the output sequence specify a plurality of coefficients of a lossy compression algorithm; and

wherein generating the output image comprises decoding the output image by applying the lossy compression algorithm to the plurality of coefficients.