IP Library › Granted Patent US 12,327,385
Granted Patent B2
US 12,327,385 · App. 17/969,551 · Granted Jun 10, 2025

End-to-end deep generative network for low bitrate image coding

Inventors: Yifei Pei (Santa Clara, CA); Ying Liu (Santa Clara, CA); Nam Ling (Santa Clara, CA); Yongxiong Ren (San Jose, CA); Lingzhi Liu (San Jose, CA)
Assignees: SANTA CLARA UNIVERSITY; KWAI INC.
G06T9/002G06N3/0455G06N3/0475G06N3/094H04N19/124
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,327,385
App. No.
17/969,551
Granted
Jun 10, 2025
Kind
B2
Abstract

A neural network system, a method and an apparatus for image compression are provided. The neural network may include a generator including an encoder, an entropy estimator, and a decoder, where the encoder receives an input image and generates an encoder output, a plurality of quantized feature entries are obtained based on the encoder output outputted at a last encoder block, the entropy estimator receives the plurality of quantized feature entries and calculates an entropy loss based on the plurality of quantized feature entries, and the decoder receives the plurality of quantized feature entries and generates a reconstructed image. Furthermore, the neural network may include a discriminator that determines whether the reconstructed image different from the input image based on a discriminator loss. Moreover, the generator may determine whether content of the reconstructed image matches content of the input image based on a generator loss including the entropy loss.

Claims (43)

1. A neural network system implemented by one or more computers for compressing an image, comprising:

a generator comprising an encoder, an entropy estimator, and a decoder,

wherein the encoder receives an input image and generates an encoder output, a plurality of quantized feature entries are obtained based on the encoder output outputted at a last encoder block in the encoder, the entropy estimator receives the plurality of quantized feature entries and calculates an entropy loss based on the plurality of quantized feature entries, and the decoder receives the plurality of quantized feature entries and generates a reconstructed image; and

a discriminator that determines whether the reconstructed image is different from the input image based on a discriminator loss,

wherein a generator loss comprises the entropy loss and a combined content loss, and the combined content loss comprises a mean absolute error (MAE) and a multi-scale structural similarity index measure (MS-SSIM) between the reconstructed image and the input image, and

wherein the generator is configured to determine whether content of the reconstructed image matches content of the input image based on the generator loss.

2. The neural network system of claim 1 , wherein the plurality of quantized feature entries are obtained after performing soft quantization on the encoder output.

3. The neural network system of claim 1 , wherein the plurality of quantized feature entries are encoded to binary representations using a standardized binary arithmetic coding and sent to the decoder.

4. The neural network system of claim 1 , wherein entropy of the plurality of quantized feature entries is estimated based on a standard normal distribution.

5. The neural network system of claim 1 , wherein the generator loss further comprises an adversarial loss.

6. The neural network system of claim 1 , wherein the discriminator loss comprises a hinge loss that penalizes incorrect classification results.

7. The neural network system of claim 1 , wherein the reconstructed image is at a bitrate lower than a predetermined threshold.

8. The neural network system of claim 1 , wherein the discriminator updates discriminator network weights based on the discriminator loss, and the generator updates generator network weights based on the generator loss.

9. A method for compressing an image, comprising:

receiving, by an encoder in a generator in a neural network system, an input image and generating, by the encoder, an encoder output;

obtaining, by the encoder, a plurality of quantized feature entries based on the encoder output outputted at a last encoder block in the encoder;

receiving, by an entropy estimator in the generator, the plurality of quantized feature entries and calculating, by the entropy estimator, an entropy loss based on the plurality of quantized feature entries;

receiving, by a decoder in the generator, the plurality of quantized feature entries and generating, by the decoder, a reconstructed image;

determining, by a discriminator in the neural network system, whether the reconstructed image is different from the input image based on a discriminator loss, and

determining, by the generator, whether content of the reconstructed image matches content of the input image based on a generator loss, wherein the generator loss comprises the entropy loss and a combined content loss comprising a mean absolute error (MAE) and a multi-scale structural similarity index measure (MS-SSIM) between the reconstructed image and the input image.

10. The method of claim 9 , further comprising:

obtaining the plurality of quantized feature entries after performing soft quantization on the encoder output.

11. The method of claim 9 , further comprising:

encoding the plurality of quantized feature entries using a standardized binary arithmetic coding to binary representations and sending the plurality of quantized feature entries that are encoded to the decoder.

12. The method of claim 9 , further comprising:

estimating entropy of the plurality of quantized feature entries based on a standard normal distribution.

13. The method of claim 9 , wherein the generator loss further comprises an adversarial loss.

14. The method of claim 9 , further comprising:

calculating, by the discriminator, the discriminator loss comprising a hinge loss that penalizes incorrect classification results.

15. The method of claim 9 , wherein the reconstructed image is at a bitrate lower than a predetermined threshold.

16. The method of claim 9 , further comprising:

updating discriminator network weights based on the discriminator loss, or

updating generator network weights based on the generator loss.

17. An apparatus for compressing an image, comprising:

one or more processors; and

a memory configured to store instructions executable by the one or more processors, wherein the one or more processors, upon execution of the instructions, are configured to perform acts comprising:

receiving, by an encoder in a generator in a neural network system, an input image and generating, by the encoder, an encoder output;

obtaining, by the encoder, a plurality of quantized feature entries based on the encoder output outputted at a last encoder block in the encoder;

receiving, by an entropy estimator in the generator, the plurality of quantized feature entries and calculating, by the entropy estimator, an entropy loss based on the plurality of quantized feature entries;

receiving, by a decoder in the generator, the plurality of quantized feature entries and generating, by the decoder, a reconstructed image;

determining, by a discriminator in the neural network system, whether the reconstructed image is different from the input image based on a discriminator loss, and determining, by the generator, whether content of the reconstructed image matches the content of the original image based on a generator loss, wherein the generator loss comprises the entropy loss and a combined content loss comprising a mean absolute error (MAE) and a multi-scale structural similarity index measure (MS-SSIM) between the reconstructed image and the input image.

18. The apparatus of claim 17 , wherein the one or more processors are configured to perform acts further comprising:

wherein the discriminator loss comprises a hinge loss that penalizes incorrect classification results.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2026
From: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
To: BEIJING TRANSTREAMS TECHNOLOGY CO., LTD.
Reel/Frame 074656/0721 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2025
From: KWAI INC.
To: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 073219/0688 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2022
From: PEI, YIFEI; LIU, YING; LING, NAM
To: SANTA CLARA UNIVERSITY
Reel/Frame 061504/0443 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2022
From: REN, YONGXIONG; LIU, LINGZHI
To: KWAI INC.
Reel/Frame 061504/0451 →
Continuity (1)
Related Publication 20240185473A1 · Jun 6, 2024
References Cited (30)
US 11153566B1 · Tao · 2021 [cited by examiner]
US 11415571B2 · Larsen · 2022 [cited by examiner]
US 11558620B2 · Besenbruch · 2023 [cited by examiner]
US 20090099986A1 · Varma · 2009 [cited by examiner]
US 20120294355A1 · Holcomb · 2012 [cited by examiner]
US 20240185075A1 · Du · 2024 [cited by examiner]
WO WO2022067656A1 · 2022 [cited by examiner]
Search machine translation of WO 2022/067656 A1 to Lyu, Image Processing Method and Apparatus, translated: Nov. 21, 2024, 33 pages. (Year: 2024). [cited by examiner]
Liu et al., Facial Image Inpainting Using Multi-Level Generative Network, Jul. 8-12, 2019 [retrieved Nov. 21, 2024], 2019 IEEE International Conference on Multimedia and Expo (ICME), pp. 1168-1173. DOI: 10.1109/ICME.201… [cited by examiner]
Yang, Jiayu et al., “Learned Low Bit-rate Image Compression with Adversarial Mechanism”, Peking University, China, CVPR 2020 workshop paper Open Access Version provided by Computer Vision Foundation and final published … [cited by applicant]
Mentzer, Fabian et al., “High-Fidelity Generative Image Compression”, 34th Conference on Neural Information Processing Systems (NeurIPS2020), Vancouver, Canada, (12p). [cited by applicant]
Setiawan, Antonius Darma et al., “Low-Bitrate Medical Image Compression”, 14-32 MVA2011 IAPR Conference on Machine Vision Applications, Jun. 13-15, 2011, Nara, Japan, (5p). [cited by applicant]
Johannes Balle et al., “End-To-End Optimized Image Compression”, arXiv:1611.01704v3, [cs.CV] Mar. 3, 2017, Published as a conference paper at ICLR 2017, (27p). [cited by applicant]
J. Ballé, et al., “Variational image compression with a scale hyperprior,” arXiv:1802.01436v2, [eess.IW] May 1, 2018, Published as a conference paper at ICLR 2018,(23p). [cited by applicant]
Z. Cheng, et al., “Learned image compression with discretized gaussian mixture likelihoods and attention modules,” in Proceedings of the IEEE Xplore, provided by the Computer Vision Foundation,(10p). [cited by applicant]
“Generative Adversarial Networks for Extreme Learned Image Compression,” Under review as a conference paper at CLR 2019, (31p). [cited by applicant]
N. Ling, et al, “The future of video coding,” APSIPA Transactions on Signal and Information Processing, 2022, 11, e16, (29p). [cited by applicant]
S. Iwai, et al., “Fidelity-controllable extreme image compression with generative adversarial networks,” arXiv:2008.10314v1 [eess.IV] Aug. 24, 2020,(8p). [cited by applicant]
F. Mentzer, et al., “High-fidelity generative image compression,” arXiv:2006.09965v3 [eess.IV] Oct. 23, 2020,(20p). [cited by applicant]
Goodfellow, J. et al., “Generative adversarial nets,” Advances in neural information processing systems, vol. 27, 2014, (9p). [cited by applicant]
M. Arjovsky, et al., “Wasserstein generative adversarial networks,” in International conference on machine earning. PMLR, 2017, (10p). [cited by applicant]
I. Gulrajani, et al., “Improved training of wasserstein gans,” Advances in neural information processing systems, vol. 30, 2017,(11p). [cited by applicant]
X. Mao, et al, “Least squares generative adversarial networks,” arXiv:1611.04076v3 [cs.CV] Apr. 5, 2017,(16p). [cited by applicant]
Ting-Chun Wang et al., “High-resolution image synthesis and semantic manipulation with conditional gans,” provided by the Computer Vision Foundation, (10p). [cited by applicant]
Haoyue Shi, et al., “Loss functions for person image generation.” in BMVC, 2020,(13p). [cited by applicant]
D. Ulyanov, et al, “Instance normalization: The missing ingredient for fast stylization,” arXiv:1607.08022v3 [cs:CV], Nov. 6, 2017, (6p). [cited by applicant]
E. Agustsson, et al., “Soft-to-hard vector quantization for end-to-end learning compressible representations,” arXiv:1704.00648v2 [cs.LG], Jan. 8, 2017, (16p). [cited by applicant]
Tsung-Yi Lin, et al., “Microsoft coco: Common objects in context,” in European conference on computer vision. Springer, 2014, (16p). [cited by applicant]
R. Franzen, “Kodak lossless true color image suite,” source: https://r0k.us/graphics/kodak, (3p). [cited by applicant]
M. Heusel, et al., “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, (12p). [cited by applicant]