IP Library › Granted Patent US 12,283,075
Granted Patent B2
US 12,283,075 · App. 17/499,959 · Granted Apr 22, 2025

Method and apparatus for multi-learning rates of substitution in neural image compression

Inventors: Ding Ding (Palo Alto, CA); Wei Jiang (Sunnyvale, CA); Sheng Lin (San Jose, CA); Wei Wang (Palo Alto, CA); Xiaozhong Xu (State College, PA); Shan Liu (San Jose, CA)
Assignee: TENCENT AMERICA LLC
G06T9/002G06N3/08H04N19/184G06F2212/455G11B20/00007H04Q2213/034
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,283,075
App. No.
17/499,959
Granted
Apr 22, 2025
Kind
B2
Abstract

Neural network based substitutional end-to-end (E2E) image compression (NIC) being performed by at least one processor and includes receiving an input image to an E2E NIC framework, determining a substitute image based on a training model of the E2E NIC framework, encoding the substitute image to generate a bitstream, mapping the substitute image to the bitstream to generate a compressed representation of the input image. Further, the input may be partitioned into blocks for which a substitute representation is determined for each block and each block is encoded instead of the entire substitute image.

Claims (62)

1. A method of substitutional end-to-end (E2E) neural image compression (NIC) using a neural network performed by at least one processor, the method comprising:

receiving an input image to an E2E NIC framework;

splitting the input image into one or more blocks;

performing an encoding mapping, for each of the one or more blocks, by mapping the input image to a first bitstream having a first length;

performing a decoding mapping, for each of the one or more blocks, by mapping the first bitstream back to an original space with a first distortion loss;

determining a substitute image from the original space, based on a training model of the E2E NIC framework;

encoding the substitute image to generate a second bitstream; and

mapping the substitute image to the second bitstream to generate a compressed representation,

wherein the training model of the E2E NIC framework is trained based on a learning rate of the input image, a quantity of updates to the input image, and a second distortion loss,

wherein a plurality of substitute images are determined based on learning rates that are selected based on characteristics of the input image, and

wherein the substitute image is determined by performing an optimization process of the training model of the E2E NIC framework, comprising:

adjusting RGB variance of the split blocks to generate substitute block representations; and

selecting the RGB variance with a least distortion loss between the split blocks and the substitute block representations to use as the substitute block.

2. The method according to claim 1 , further comprising:

determining a substitute block for each of the one or more blocks, based on the training model of the E2E NIC framework;

encoding the substitute block to generate a block bitstream; and

mapping the substitute block to the block bitstream to generate a compressed block,

wherein the one or more blocks have a same size, and each block of the one or more blocks has a different learning rate.

3. The method according to claim 1 , wherein the training model of the E2E NIC framework is an artificial neural network based on pretrained image coding, and

wherein parameters of the artificial neural network are fixed and a gradient is used to update the input image.

4. An apparatus for substitutional end-to-end (E2E) neural image compression (NIC) using a neural network, the apparatus comprising:

at least one memory configured to store program code; and

at least one processor configured to read the program code and operate as instructed by the program code, the program code comprising:

receiving code configured to cause at least one processor to receive an input image to an E2E NIC framework;

splitting code configured to cause at least one processor to split the input image into one or more blocks;

first performing code configured to cause the at least one processor to perform an encoding mapping, for each of the one or more blocks, by mapping the input image to a first bitstream having a first length;

second performing code configured to cause the at least one processor to perform a decoding mapping, for each of the one or more blocks, by mapping the first bitstream back to an original space with a first distortion loss;

first determining code configured to cause at least one processor to determine a substitute image from the original space, based on a training model of the E2E NIC framework;

first encoding code configured to cause at least one processor to encode the substitute image to generate a second bitstream; and

first mapping code configured to cause at least one processor to map the substitute image to the second bitstream to generate a compressed representation,

wherein the training model of the E2E NIC framework is trained based on a learning rate of the input image, a quantity of updates to the input image, and a second distortion loss,

wherein a plurality of substitute images are determined based on learning rates that are selected based on characteristics of the input image, and

wherein the substitute image is determined by performing an optimization process of the training model of the E2E NIC framework, comprising:

adjusting code configured to cause at least one processor to adjust RGB variance of the split blocks to generate substitute block representations; and

selecting code configured to cause at least one processor to select the RGB variance with a least distortion loss between the split blocks and the substitute block representations to use as the substitute block.

5. The apparatus of claim 4 , further comprising:

second determining code configured to cause at least one processor to determine a substitute block for each of the one or more blocks, based on the training model of the E2E NIC framework;

second encoding code configured to cause at least one processor to encode the substitute block to generate a block bitstream; and

second mapping code configured to cause at least one processor to map the substitute block to the block bitstream to generate a compressed block,

wherein the one or more blocks have a same size, and each block of the one or more blocks has a different learning rate.

6. The apparatus according to claim 4 , wherein the training model of the E2E NIC framework is an artificial neural network based on pretrained image coding, and

wherein parameters of the artificial neural network are fixed and a gradient is used to update the input image.

7. A non-transitory computer readable medium storing instructions that, when executed by at least one processor for substitutional end-to-end (E2E) neural image compression (NIC), cause the at least one processor to:

receive an input image to an E2E NIC framework;

split the input image into one or more blocks;

perform an encoding mapping, for each of the one or more blocks, by mapping the input image to a first bitstream having a first length;

perform a decoding mapping, for each of the one or more blocks, by mapping the first bitstream back to an original space with a first distortion loss;

determine a substitute image from the original space, based on a training model of the E2E NIC framework;

encode the substitute image to generate a second bitstream; and

map the substitute image to the second bitstream to generate a compressed representation,

wherein the training model of the E2E NIC framework is trained based on a learning rate of the input image, a quantity of updates to the input image, and a second distortion loss,

wherein a plurality of substitute images are determined based on learning rates that are selected based on characteristics of the input image, and

wherein the instructions, when executed by at least one processor, further cause the at least one processor to performing an optimization process of the training model of the E2E NIC framework, comprising:

adjust RGB variance of the split blocks to generate substitute block representations; and

select the RGB variance with a least distortion loss between the split blocks and the substitute block representations to use as the substitute block.

8. The non-transitory computer readable medium of claim 7 , wherein the instructions, when executed by at least one processor, further cause the at least one processor to:

determine a substitute block for each of the one or more blocks, based on the training model of the E2E NIC framework;

encode the substitute block to generate a block bitstream; and

map the substitute block to the block bitstream to generate a compressed block,

wherein the one or more blocks have a same size, and each block of the one or more blocks has a different learning rate.

9. The non-transitory computer readable medium of claim 7 , wherein the training model of the E2E NIC framework is an artificial neural network based on pretrained image coding, and

wherein parameters of the artificial neural network are fixed and a gradient is used to update the input image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2021
From: DING, DING; JIANG, WEI; LIN, SHENG; WANG, WEI; XU, XIAOZHONG; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 057788/0525 →
Continuity (2)
Provisional Application 63176204 · Apr 16, 2021
Related Publication 20220343552A1 · Oct 27, 2022
References Cited (30)
US 20140177721A1 · Onno et al. · 2014 [cited by applicant]
US 20200051287A1 · Wihlidal · 2020 [cited by examiner]
US 20200137384A1 · Kwong et al. · 2020 [cited by applicant]
US 20200160565A1 · Ma · 2020 [cited by examiner]
US 20200351509A1 · Lee · 2020 [cited by examiner]
US 20200366914A1 · Schroers · 2020 [cited by examiner]
US 20220030246A1 · Xu · 2022 [cited by examiner]
US 20220230362A1 · Jiang · 2022 [cited by examiner]
ES 2310157A1 · 2008 [cited by examiner]
JP 7374340B2 · 2023 [cited by applicant]
WO 2020174216A1 · 2020 [cited by applicant]
Yang et al, Slimmable Compressive Autoencoders for Practical Neural Image Compression, 2021, arXiv:2102.15726v1, pp. 1-12. ( Year: 2021). [cited by examiner]
Ma et al, Improving Compression Artifact Reduction via End-to-End Learning of Side Information, 2020, IEEE international Conference on Visual Communications and Image Processing, pp. 1-4. (Year: 2020). [cited by examiner]
Rippel et al, Real-Time Adaptive Image Compression, 2017, arXiv:1705.05823v1, pp. 1-16. (Year: 2017). [cited by examiner]
Mentzer et al, Learning Better Lossless Compression Using Lossy Compression, 2020, IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1-12. (Year: 2020). [cited by examiner]
Lin et al, Spatial RNN Codec for End-to-End Image Compression, 2020, IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1-10. (Year: 2020). [cited by examiner]
Tellez et al, Neural Image Compression for Gigapixel Histopathology Image Analysis, 2020, arXiv: 1811.02840v2, pp. 1-15. (Year: 2020). [cited by examiner]
Ma et al, Improving Compression Artifact Reduction via End-to-End Learning of Side Information, 2020, IEEE International Conference on Visual Communications and Image Processing, pp. 1-5. (Year: 2020). [cited by examiner]
Galanopoulos et al, Measurement-driven Analysis of an Edge-Assisted Object Recognition System, 2020, arXiv: 2003.03584v1, pp. 1-7. (Year: 2020). [cited by examiner]
Yang et al, Slimmable Compressive Autoencoders for Practical Neural Image Compression, 2021, arXiv: 2103.15726v1, pp. 1-12. ( Year: 2021). [cited by examiner]
Chen et al, Neural Image Compression via Non-Local Attention Optimization and Improved Context Modeling, 2019, arXiv: 1910.06244v1, pp. 1-13. (Year: 2019). [cited by examiner]
Extended European Search Report issued Jul. 7, 2023 in European Application No. 21928340.5. [cited by applicant]
Wei Wang et al., “Substitutional Neural Image Compression” International Organisation for Standardisation Organisation Internationale de Normalisation ISO/IEC JTC1/SC29/WG11 Coding of Moving Pictures and Audio, ISO/IEC … [cited by applicant]
Caglar Aytekin et al., “Block-optimized Variable Bit Rate Neural Image Compression”, arXiv:1805.10887v1 [cs.LG], Cornell University Library, May 28, 2018, pp. 1-4 (4 pages total). [cited by applicant]
Xiao Wang et al., “Substitutional Neural Image Compression”, arXiv:2105.07512v1 [cs.CV], Cornell University Library, May 16, 2021 (8 pages total). [cited by applicant]
International Search Report dated Feb. 1, 2022 in International Application No. PCT/US21/55039. [cited by applicant]
Written Opinion of the International Searching Authority dated Feb. 1, 2022 in International Application No. PCT/US21/55039. [cited by applicant]
David Tellez et al., “Neural Image Compression for Gigapixel Histopathology Image Analysis”, IEEE transactions on pattern analysis and machine intelligence, 2020, arXiv:1811.02840v2, pp. 1-15 (15 pages). [cited by applicant]
Talebi, Hossein et al., “Better Compression with Deep Pre-Editing” arXiv:2002.00113v1, Feb. 1, 2020 13 pages total, Retrieved from the Internet: <URL: https://arxiv.org/abs/2002.00113v1>, https://doi.org/10.48550/arXiv.… [cited by applicant]
Communication dated Jan. 9, 2024, issued in Japanese Application No. 2022-564637. [cited by applicant]