IP Library › Granted Patent US 12,347,149
Granted Patent B2
US 12,347,149 · App. 17/950,569 · Granted Jul 1, 2025

System, method, and computer program for content adaptive online training for multiple blocks in neural image compression

Inventors: Ding Ding (Palo Alto, CA); Wei Wang (Palo Alto, CA); Shan Liu (Palo Alto, CA)
Assignee: TENCENT AMERICA LLC
G06T9/002G06N3/045G06T3/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,149
App. No.
17/950,569
Granted
Jul 1, 2025
Kind
B2
Abstract

Content-adaptive online training for end-to-end (E2E) neural image compression (NIC) using a neural network performed by at least one processor, is provided, including receiving an input image, to an E2E NIC framework, including one or more blocks, preprocessing a first neural network of the E2E NIC framework, based on the one or more blocks, computing updated parameters using the preprocessed first neural network, encoding the one or more blocks and the updated parameters, updating the first neural network based on the encoded updated parameters, and generating a compressed representation of the encoded one or more blocks using the updated first neural network.

Claims (50)

1. A method of content-adaptive online training for end-to-end (E2E) neural image compression (NIC) using a neural network performed by at least one processor, the method comprising:

receiving, by an E2E NIC framework, an input image including one or more blocks;

preprocessing a first neural network of the E2E NIC framework, based on the one or more blocks;

computing updated parameters using the preprocessed first neural network, wherein the updated parameters include a learning rate and a number of steps, and wherein the learning rate and the number of steps are selected based on characteristics of the input image;

encoding the one or more blocks and the updated parameters;

updating the first neural network based on the encoded updated parameters; and

generating a compressed representation of the encoded one or more blocks using the updated first neural network.

2. The method according to claim 1 , further comprising:

splitting the input image into the one or more blocks; and

compressing the one or more blocks individually.

3. The method according to claim 1 , further comprising;

decoding the compressed representation using arithmetic decoding; and

generating a reconstructed image based on the decoded compressed representation using a second neural network.

4. The method according to claim 1 , further comprising compressing the updated parameters.

5. The method according to claim 1 , wherein the characteristics of the input image are one of a RGB variance of the input image and an RD performance of the input image.

6. The method according to claim 1 , wherein when preprocessing the first neural network, the first neural network is fine-tuned using the one or more blocks.

7. An apparatus for content-adaptive online training for end-to-end (E2E) neural image compression (NIC) using a neural network, the apparatus comprising:

at least one memory configured to store computer program code; and

at least one processor configured to read the computer program code and operate as instructed by the computer program code, the computer program code including:

receiving code configured to cause the at least one processor to receive, by an E2E NIC framework, an input including one or more blocks;

preprocessing code configured to cause the at least one processor to preprocess a first neural network of the E2E NIC framework, based on the one or more blocks;

computing code configured to cause the at least one processor to compute updated parameters using the preprocessed first neural network, wherein the updated parameters include a learning rate and a number of steps, and wherein the learning rate and the number of steps are selected based on characteristics of the input image;

encoding code configured to cause the at least one processor to encode the one or more blocks and the updated parameters;

updating code configured to cause the at least one processor to update the first neural network based on the encoded updated parameters; and

first generating code configured to cause the at least one processor to generate a compressed representation of the encoded one or more blocks using the updated first neural network.

8. The apparatus of claim 7 , the computer program code further including:

splitting code configured to cause the at least one processor to split the input image into the one or more blocks; and

compressing code configured to cause the at least one processor to compress the one or more blocks individually.

9. The apparatus of claim 7 , the computer program code further including:

decoding code configured to cause the at least one processor to decode the compressed representation using arithmetic decoding; and

second generating code configured to cause the at least one processor to generate a reconstructed image based on the decoded compressed representation using a second neural network.

10. The apparatus of claim 7 , the computer program code further including compressing code configured to cause the at least one processor to compress the updated parameters.

11. The apparatus of claim 7 , wherein the characteristics of the input image are one of a RGB variance of the input image and an RD performance of the input image.

12. The apparatus of claim 7 , wherein when preprocessing the first neural network, the first neural network is fine-tuned using the one or more blocks.

13. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor of an apparatus for content-adaptive online training for end-to-end (E2E) neural image compression (NIC) using a neural network, cause the at least one processor to:

receive, by an E2E NIC framework, an input image including one or more blocks;

preprocess a first neural network of the E2E NIC framework, based on the one or more blocks;

compute updated parameters using the preprocessed first neural network;

encode the one or more blocks and the updated parameters, wherein the updated parameters include a learning rate and a number of steps, and wherein the learning rate and the number of steps are selected based on characteristics of the input image;

update the first neural network based on the encoded updated parameters; and

generate a compressed representation of the encoded one or more blocks using the updated first neural network.

14. The non-transitory computer-readable medium of claim 13 , wherein the instructions further cause the at least one processor to:

split the input image into the one or more blocks; and

compress the one or more blocks individually.

15. The non-transitory computer-readable medium of claim 13 , wherein the instructions further cause the at least one processor to:

decode the compressed representation using arithmetic decoding; and

generate a reconstructed image based on the decoded compressed representation using a second neural network.

16. The non-transitory computer-readable medium of claim 13 , wherein the instructions further cause the at least one processor to compress the updated parameters.

17. The non-transitory computer-readable medium of claim 13 , wherein the characteristics of the input image are one of a RGB variance of the input image and an RD performance of the input image.

18. The non-transitory computer-readable medium of claim 13 , wherein when preprocessing the first neural network, the first neural network is fine-tuned using the one or more blocks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2022
From: DING, DING; WANG, WEI; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 061185/0333 →
Continuity (2)
Provisional Application 63289033 · Dec 13, 2021
Related Publication 20230186525A1 · Jun 15, 2023
References Cited (19)
US 5729661A · Keeler et al. · 1998 [cited by applicant]
US 20170148431A1 · Catanzaro et al. · 2017 [cited by applicant]
US 20200029084A1 · Wierstra et al. · 2020 [cited by applicant]
US 20200366914A1 · Schroers et al. · 2020 [cited by applicant]
US 20220021870A1 · Jiang et al. · 2022 [cited by applicant]
US 20220051367A1 · Jiang et al. · 2022 [cited by applicant]
Pytorch, Github, 2018 (Year: 2018). [cited by examiner]
Toderici, Full Resolution Image Compression with Recurrent Neural Networks, CVPR (Year: 2017). [cited by examiner]
Kingma, “Adam: a Method for Stochastic Optimization”, ICLR (Year: 2015). [cited by examiner]
International Search Report dated Jan. 20, 2023 in International Application No. PCT/US2022/045007. [cited by applicant]
Written Opinion of the International Searching Authority dated Jan. 20, 2023 in International Application No. PCT/US2022/045007. [cited by applicant]
Wei Jiang et al., “Online Meta Adaptation for Variable-Rate Learned Image Compression”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 1-9 (9 pages total). [cited by applicant]
Guanbo Pan et al., “Content Adaptive Latents and Decoder for Neural Image Compression”, European Conference on Computer Vision, 2022, pp. 1-18 (34 pages total). [cited by applicant]
Hyunho Yeo et al., “Neural Adaptive Content-aware Internet Video Delivery”, 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI '18), 2018, pp. 645-661 (18 pages total). [cited by applicant]
Zhenghui Zhao, et al. “Learned Image Compression using Adaptive Block-wise Encoding and Reconstruction Network”, International Symposium on Circuits and Systems(ISCAS), 2021, IEEE (5 Pages). [cited by applicant]
Japanese Office Action dated Oct. 15, 2024 in Application No. 2023-560171. [cited by applicant]
Nannan Zou, et al.“L [cited by applicant]
Yat-Hong Lam, et al. “Efficient Adaptation of Neural Network Filter for Video Compression”, CHI Conference On Human Factors in Computing Systems, Oct. 12-26, 2020, pp. 358-366 (9 pages). [cited by applicant]
Extended European Search Report dated Mar. 3, 2025 in Application No. 22908169.0. [cited by applicant]