IP Library › Granted Patent US 12,518,432
Granted Patent B2
US 12,518,432 · App. 17/969,242 · Granted Jan 6, 2026

System, method, and computer program for content adaptive online training for multiple blocks based on certain patterns

Inventors: Ding Ding (Palo Alto, CA); Wei Wang (Palo Alto, CA); Shan Liu (Palo Alto, CA)
Assignee: TENCENT AMERICA LLC
G06T9/002G06T3/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,432
App. No.
17/969,242
Granted
Jan 6, 2026
Kind
B2
Abstract

Content-adaptive online training for end-to-end (E2E) neural image compression (NIC) using a neural network performed by at least one processor, is provided, including receiving an input image, to an E2E NIC framework, splitting the input image into a plurality of blocks, selecting a subset of blocks from the plurality of blocks, the subset of blocks sharing a same pattern, preprocessing a neural network of the E2E NIC framework, wherein the preprocessed neural network is applied to the selected subset of blocks, computing updated parameters using the preprocessed neural network, and generating an updated E2E NIC framework, based on the updated parameters.

Claims (47)

1 . A method of content-adaptive online training for end-to-end (E2E) neural image compression (NIC) using a neural network performed by at least one processor, the method comprising:

receiving, by an E2E NIC framework, an input image comprising a plurality of blocks;

preprocessing a neural network of the E2E NIC framework using a preprocessed neural network based on the plurality of blocks;

computing updated bias parameters using a last layer of the preprocessed neural network, wherein the updated bias parameters comprise first updated bias parameters that are associated with a first subset of blocks from the plurality of blocks that share a same pattern, the first subset of blocks having one of a same RGB variance, a same YUV variance, and a same dominant color; and

generating an updated E2E NIC framework, based on the updated bias parameters.

2 . The method according to claim 1 , further comprising:

encoding the plurality of blocks and the updated bias parameters to generate a compressed representation of the plurality of blocks and a compressed representation of the updated bias parameters;

decoding the compressed representation of the updated bias parameters to generate decoded updated bias parameters;

updating the E2E NIC framework based on the decoded updated bias parameters; and

decoding the compressed representation of the plurality of blocks based on the updated E2E NIC framework to generate a reconstructed image.

3 . The method according to claim 2 , further comprising determining a distortion loss of the reconstructed image based on a consumption of the compressed representation of the plurality of blocks and the updated bias parameters, trade-off hyperparameter, and a distortion between block residuals of the compressed representation of the plurality of blocks and a block residual of the decoded compressed representations of the plurality of blocks.

4 . The method according to claim 1 , wherein the same pattern is determined based on an RGB variance of the plurality of blocks or a YUV variance of the plurality of blocks.

5 . The method according to claim 1 , wherein the updated bias parameters include a learning rate and a number of steps, and the learning rate and the number of steps are selected based on characteristics of the input image.

6 . The method according to claim 5 , wherein the characteristics of the input image are one of a RGB variance of the input image and an RD performance of the input image.

7 . The method according to claim 1 , wherein when preprocessing the neural network, the neural network is fine-tuned using the plurality of blocks.

8 . An apparatus for content-adaptive online training for end-to-end (E2E) neural image compression (NIC) using a neural network, the apparatus comprising:

at least one memory configured to store computer program code; and

at least one processor configured to read the computer program code and operate as instructed by the computer program code, the computer program code including:

receiving code configured to cause the at least one processor to receive, by an E2E NIC framework, an input image comprising a plurality of blocks;

preprocessing code configured to cause the at least one processor to preprocess a neural network of the E2E NIC framework using a preprocessed neural network based on the plurality of blocks;

computing code configured to cause the at least one processor to compute updated bias parameters using a last layer of the preprocessed neural network, wherein the updated bias parameters comprise first updated bias parameters that are associated with a first subset of blocks from the plurality of blocks that share a same pattern, the first subset of blocks having one of a same RGB variance, a same YUV variance, and a same dominant color; and

generating code configured to cause the at least one processor to generate an updated E2E NIC framework, based on the updated bias parameters.

9 . The apparatus of claim 8 , the computer program code further including:

encoding code configured to cause the at least one processor to encode the plurality of blocks and the updated bias parameters to generate a compressed representation of the plurality of blocks and a compressed representation of the updated bias parameters;

first decoding code configured to cause the at least one processor to decode the compressed representation of the updated bias parameters to generate decoded updated bias parameters;

updating code configured to cause the at least one processor to update the E2E NIC framework based on the decoded updated bias parameters; and

second decoding code configured to cause the at least one processor to decode the compressed representation of the plurality of blocks based on the updated E2E NIC framework to generate a reconstructed image.

10 . The apparatus of claim 9 , the computer program code further including distortion loss determining code configured to cause the at least one processor to determine a distortion loss of the reconstructed image based on a consumption of the compressed representation of the plurality of blocks and the updated bias parameters, trade-off hyperparameter, and a distortion between block residuals of the compressed representation of the plurality of blocks and a block residual of the decoded compressed representations of the plurality of blocks.

11 . The apparatus of claim 8 , wherein the same pattern is determined based on an RGB variance of the plurality of blocks or a YUV variance of the plurality of blocks.

12 . The apparatus of claim 8 , wherein the updated bias parameters include a learning rate and a number of steps, and the learning rate and the number of steps are selected based on characteristics of the input image.

13 . The apparatus of claim 12 , wherein the characteristics of the input image are one of a RGB variance of the input image and an RD performance of the input image.

14 . The apparatus of claim 8 , wherein when preprocessing the neural network, the neural network is fine-tuned using the plurality of blocks.

15 . A non-transitory computer-readable medium storing instructions that, when executed by at least one processor of an apparatus for content-adaptive online training for end-to-end (E2E) neural image compression (NIC) using a neural network, cause the at least one processor to:

receive, by an E2E NIC framework, an input image comprising a plurality of blocks;

preprocess a neural network of the E2E NIC framework using a preprocessed neural network based on the plurality of blocks;

compute updated bias parameters using a last layer of the preprocessed neural network, wherein the updated bias parameters comprise first updated bias parameters that are associated with a first subset of blocks from the plurality of blocks that share a same pattern, the first subset of blocks having one of a same RGB variance, a same YUV variance, and a same dominant color; and

generate an updated E2E NIC framework, based on the updated bias parameters.

16 . The non-transitory computer-readable medium of claim 15 , wherein the instructions further cause the at least one processor to:

encode the plurality of blocks and the updated bias parameters to generate a compressed representation of the plurality of blocks and a compressed representation of the updated bias parameters;

decode the compressed representation of the updated bias parameters to generate decoded updated bias parameters;

update the E2E NIC framework based on the decoded updated bias parameters; and

decode the compressed representation of the plurality of blocks based on the updated E2E NIC framework to generate a reconstructed image.

17 . The non-transitory computer-readable medium of claim 16 ,

wherein the instructions further cause the at least one processor to determine a distortion loss of the reconstructed image based on a consumption of the compressed representation of the plurality of blocks and the updated bias parameters, trade-off hyperparameter, and a distortion between block residuals of the compressed representation of the plurality of blocks and a block residual of the decoded compressed representations of the plurality of blocks.

18 . The non-transitory computer-readable medium of claim 15 , wherein the same pattern is determined based on an RGB variance of the plurality of blocks or a YUV variance of the plurality of blocks.

19 . The non-transitory computer-readable medium of claim 15 , wherein the updated bias parameters include a learning rate and a number of steps, the learning rate and the number of steps are selected based on characteristics of the input image, and the characteristics of the input image are one of a RGB variance of the input image and an RD performance of the input image.

20 . The non-transitory computer-readable medium of claim 15 , wherein when preprocessing the neural network, the neural network is fine-tuned using the plurality of blocks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2022
From: DING, DING; WANG, WEI; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 061472/0136 →
Continuity (2)
Provisional Application 63289044 · Dec 13, 2021
Related Publication 20230186526A1 · Jun 15, 2023
References Cited (32)
US 10666962B2 · Wang · 2020 [cited by examiner]
US 20030227539A1 · Bonnery · 2003 [cited by examiner]
US 20090034622A1 · Huchet · 2009 [cited by examiner]
US 20110219033A1 · Oya · 2011 [cited by examiner]
US 20190362467A1 · Lee · 2019 [cited by examiner]
US 20210142524A1 · Djelouah et al. · 2021 [cited by applicant]
US 20210166434A1 · Miyauchi · 2021 [cited by applicant]
US 20210248748A1 · Turgutlu · 2021 [cited by examiner]
US 20220004810A1 · Sinha · 2022 [cited by examiner]
US 20220101492A1 · Ding · 2022 [cited by examiner]
US 20220103839A1 · Van Rozendaal · 2022 [cited by examiner]
US 20220215592A1 · Jiang · 2022 [cited by examiner]
US 20220343552A1 · Ding · 2022 [cited by examiner]
US 20220345717A1 · Lin · 2022 [cited by examiner]
US 20220385896A1 · Ding · 2022 [cited by examiner]
US 20220398455A1 · Dumas · 2022 [cited by examiner]
US 20220405979A1 · Ding · 2022 [cited by examiner]
US 20230186081A1 · Ding · 2023 [cited by examiner]
US 20230186525A1 · Ding · 2023 [cited by examiner]
US 20230186526A1 · Ding · 2023 [cited by examiner]
US 20240223775A1 · Gao · 2024 [cited by examiner]
US 20240378417A1 · Chilkuri · 2024 [cited by examiner]
JP 202190135A · 2021 [cited by applicant]
WO 2020008104A1 · 2020 [cited by applicant]
WO 2023113899A1 · 2023 [cited by applicant]
M. Li, W. Zuo, S. Gu, D. Zhao and D. Zhang, “Learning Convolutional Networks for Content-Weighted Image Compression,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018, p… [cited by examiner]
N. Zou et al., “L2C—Learning to Learn to Compress,” 2020 IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP), Tampere, Finland, 2020, pp. 1-6, doi: 10.1109/MMSP48831.2020.9287069. (Year: 2020). [cited by examiner]
International Search Report dated Mar. 2, 2023 in International Application No. PCT/US22/48656. [cited by applicant]
Written Opinion dated Mar. 2, 2023 in International Application No. PCT/US22/48656. [cited by applicant]
Communication issued Oct. 28, 2024 in Japanese Application No. 2023-561073. [cited by applicant]
Aytekin et al., “Block-optimized Variable Bit Rate Neural Image Compression”, arXiv: 1805.10887v1, May 28, 2018, pp. 1-4. [cited by applicant]
Extended European Search Report issued Feb. 28, 2025 in Application No. 22908184.9. [cited by applicant]