IP Library › Granted Patent US 11,915,457
Granted Patent B2
US 11,915,457 · App. 17/365,371 · Granted Feb 27, 2024

Method and apparatus for adaptive neural image compression with rate control by meta-learning

Inventors: Wei Jiang (Sunnyvale, CA); Wei Wang (Palo Alto, CA); Shan Liu (San Jose, CA)
Assignee: TENCENT AMERICA LLC
G06T9/002G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,915,457
App. No.
17/365,371
Granted
Feb 27, 2024
Kind
B2
Abstract

A method of adaptive neural image compression with rate control by meta-learning includes receiving an input image and a hyperparameter; and encoding the received input image, based on the received hyperparameter, using an encoding neural network, to generate a compressed representation. The encoding includes performing a first shared encoding on the received input image, using a first shared encoding layer having first shared encoding parameters, performing a first adaptive encoding on the received input image, using a first adaptive encoding layer having first adaptive encoding parameters, combining the first shared encoded input image and the first adaptive encoded input image, to generate a first combined output, and performing a second shared encoding on the first combined output, using a second shared encoding layer having second shared encoding parameters.

Claims (113)

1. A method of adaptive neural image compression with rate control by meta-learning, the method being performed by at least one processor, and the method comprising:

receiving a recovered compressed representation and a hyperparameter of an image; and

decoding the received recovered compressed representation based on a received hyperparameter, using a decoded neural network, to reconstruct an output image, wherein the received recovered compressed representation is based on:

receiving an input image and the hyperparameter; and

encoding the received input image, based on the received hyperparameter, using an encoding neural network, to generate a compressed representation, wherein the encoding comprises:

performing a first shared encoding on the received input image, using a first shared encoding layer having first shared encoding parameters;

performing a first adaptive encoding on the received input image, using a first adaptive encoding layer having first adaptive encoding parameters;

combining the first shared encoded input image and the first adaptive encoded input image, to generate a first combined output; and

performing a second shared encoding on the first combined output, using a second shared encoding layer having second shared encoding parameters, and

wherein decoding the received recovered compressed representation comprises:

performing a first shared decoding on the received recovered compressed representation, using a first shared decoding layer having first shared decoding parameters;

performing a first adaptive decoding on the received recovered compressed representation, using a first adaptive decoding layer having first adaptive decoding parameters; and

combining the first shared decoded recovered compressed representation and the first adaptive decoded recovered compressed representation, to generate a second combined output.

2. The method of claim 1 , further comprising:

performing a second adaptive encoding on the first combined output, using a second adaptive encoding layer having second adaptive encoding parameters,

wherein the decoding further comprises:

performing a second shared decoding on the second combined output, using a second shared decoding layer having second shared decoding parameters; and

performing a second adaptive decoding on the second combined output, using a second adaptive decoding layer having second adaptive decoding parameters.

3. The method of claim 2 , wherein the encoding further comprises:

generating a shared feature, based the received input image and the first shared encoding parameters;

generating estimated adaptive encoding parameters, based on one or more of the received input image, the first adaptive encoding parameters, the generated shared feature, and the received hyperparameter, using a prediction neural network; and

generating the compressed representation, based on the estimated adaptive encoding parameters and the received hyperparameter.

4. The method of claim 3 , wherein the prediction neural network is trained by:

generating a first loss for training data corresponding to the received hyperparameter, and a second loss for validation data corresponding to the received hyperparameter, based on the received hyperparameter, the first shared encoding parameters, the first adaptive encoding parameters, the first shared decoding parameters, the first adaptive decoding parameters, and prediction parameters of the prediction neural network; and

updating the prediction parameters, based on gradients of the generated first loss and the generated second loss.

5. The method of claim 2 , wherein the decoding further comprises:

generating a shared feature, based the received input image and the first shared decoding parameters;

generating estimated adaptive decoding parameters, based on one or more of the received input image, the first adaptive decoding parameters, the generated shared feature, and the received hyperparameter, using a prediction neural network; and

reconstructing the output image, based on the estimated adaptive decoding parameters and the received hyperparameter.

6. The method of claim 5 , wherein the prediction neural network is trained by:

generating a first loss for training data corresponding to the received hyperparameter, and a second loss for validation data corresponding to the received hyperparameter, based on the received hyperparameter, the first shared encoding parameters, the first adaptive encoding parameters, the first shared decoding parameters, the first adaptive decoding parameters, and prediction parameters of the prediction neural network; and

updating the prediction parameters, based on gradients of the generated first loss and the generated second loss.

7. The method of claim 2 , wherein the encoding neural network and the decoding neural network are trained by:

generating an inner-loop loss for training data corresponding to the received hyperparameter, based on the received hyperparameter, the first shared encoding parameters, the first adaptive encoding parameters, the first shared decoding parameters, and the first adaptive decoding parameters;

first updating the first shared encoding parameters, the first adaptive encoding parameters, the first shared decoding parameters and the first adaptive decoding parameters, based on gradients of the generated inner-loop loss;

generating a meta loss for validation data corresponding to the received hyperparameter, based on the received hyperparameter, the first updated first shared encoding parameters, the first updated first adaptive encoding parameters, the first updated first shared decoding parameters, and the first updated first adaptive decoding parameters; and

second updating the first updated first shared encoding parameters, the first updated first adaptive encoding parameters, the first updated first shared decoding parameters, and the first updated first adaptive decoding parameters, based on gradients of the generated meta loss.

8. An apparatus for adaptive neural image compression with rate control by meta-learning, the apparatus comprising:

at least one memory configured to store program code; and

at least one processor configured to read the program code and operate as instructed by the program code, the program code comprising:

decoding code configured to cause the at least one processor to:

receive a recovered compressed representation and a hyperparameter of an image; and

decode the received recovered compressed representation based on a received hyperparameter, using a decoded neural network, to reconstruct an output image, wherein the received recovered compressed representation is based on:

first receiving code configured to cause the at least one processor to receiving an input image and the hyperparameter; and

encoding code configured to cause the at least one processor to encode the received input image, based on the received hyperparameter, using an encoding neural network, to generate a compressed representation,

wherein the encoding code is further configured to cause the at least one processor to:

perform a first shared encoding on the received input image, using a first shared encoding layer having first shared encoding parameters;

perform a first adaptive encoding on the received input image, using a first adaptive encoding layer having first adaptive encoding parameters;

combine the first shared encoded input image and the first adaptive encoded input image, to generate a first combined output; and

perform a second shared encoding on the first combined output, using a second shared encoding layer having second shared encoding parameters,

wherein decoding the received recovered compressed representation comprises:

performing a first shared decoding on the received recovered compressed representation, using a first shared decoding layer having first shared decoding parameters;

performing a first adaptive decoding on the received recovered compressed representation, using a first adaptive decoding layer having first adaptive decoding parameters; and

combining the first shared decoded recovered compressed representation and the first adaptive decoded recovered compressed representation, to generate a second combined output.

9. The apparatus of claim 8 , wherein the encoding code is further configured to cause the at least one processor to perform a second adaptive encoding on the first combined output, using a second adaptive encoding layer having second adaptive encoding parameters,

wherein the program code further comprises:

second receiving code configured to cause the at least one processor to receive the recovered compressed representation and the hyperparameter; and

decoding code configured to cause the at least one processor to decode the received recovered compressed representation, based on the received hyperparameter, using a decoding neural network, to reconstruct an output image, and

wherein the decoding code is further configured to cause the at least one processor to:

perform a second shared decoding on the second combined output, using a second shared decoding layer having second shared decoding parameters; and

perform a second adaptive decoding on the second combined output, using a second adaptive decoding layer having second adaptive decoding parameters.

10. The apparatus of claim 9 , wherein the encoding code is further configured to cause the at least one processor to:

generate a shared feature, based the received input image and the first shared encoding parameters;

generate estimated adaptive encoding parameters, based on one or more of the received input image, the first adaptive encoding parameters, the generated shared feature, and the received hyperparameter, using a prediction neural network; and

generate the compressed representation, based on the estimated adaptive encoding parameters and the received hyperparameter.

11. The apparatus of claim 10 , wherein the prediction neural network is trained by:

generating a first loss for training data corresponding to the received hyperparameter, and a second loss for validation data corresponding to the received hyperparameter, based on the received hyperparameter, the first shared encoding parameters, the first adaptive encoding parameters, the first shared decoding parameters, the first adaptive decoding parameters, and prediction parameters of the prediction neural network; and

updating the prediction parameters, based on gradients of the generated first loss and the generated second loss.

12. The apparatus of claim 9 , wherein the decoding code is further configured to cause the at least one processor to:

generate a shared feature, based the received input image and the first shared decoding parameters;

generate estimated adaptive decoding parameters, based on one or more of the received input image, the first adaptive decoding parameters, the generated shared feature, and the received hyperparameter, using a prediction neural network; and

reconstruct the output image, based on the estimated adaptive decoding parameters and the received hyperparameter.

13. The apparatus of claim 12 , wherein the prediction neural network is trained by:

generating a first loss for training data corresponding to the received hyperparameter, and a second loss for validation data corresponding to the received hyperparameter, based on the received hyperparameter, the first shared encoding parameters, the first adaptive encoding parameters, the first shared decoding parameters, the first adaptive decoding parameters, and prediction parameters of the prediction neural network; and

updating the prediction parameters, based on gradients of the generated first loss and the generated second loss.

14. The apparatus of claim 9 , wherein the encoding neural network and the decoding neural network are trained by:

generating an inner-loop loss for training data corresponding to the received hyperparameter, based on the received hyperparameter, the first shared encoding parameters, the first adaptive encoding parameters, the first shared decoding parameters, and the first adaptive decoding parameters;

first updating the first shared encoding parameters, the first adaptive encoding parameters, the first shared decoding parameters and the first adaptive decoding parameters, based on gradients of the generated inner-loop loss;

generating a meta loss for validation data corresponding to the received hyperparameter, based on the received hyperparameter, the first updated first shared encoding parameters, the first updated first adaptive encoding parameters, the first updated first shared decoding parameters, and the first updated first adaptive decoding parameters; and

second updating the first updated first shared encoding parameters, the first updated first adaptive encoding parameters, the first updated first shared decoding parameters, and the first updated first adaptive decoding parameters, based on gradients of the generated meta loss.

15. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor for adaptive neural image compression with rate control by meta-learning, cause the at least one processor to:

receive a recovered compressed representation and a hyperparameter of an image; and

decode the received recovered compressed representation based on a received hyperparameter, using a decoded neural network, to reconstruct an output image, wherein the received recovered compressed representation is based on:

receiving an input image and the hyperparameter; and

encoding the received input image, based on the received hyperparameter, using an encoding neural network, to generate a compressed representations based on:

performing a first shared encoding on the received input image, using a first shared encoding layer having first shared encoding parameters;

performing a first adaptive encoding on the received input image, using a first adaptive encoding layer having first adaptive encoding parameters;

combining the first shared encoded input image and the first adaptive encoded input image, to generate a first combined output; and

performing a second shared encoding on the first combined output, using a second shared encoding layer having second shared encoding parameters,

wherein decoding the received recovered compressed representation comprises:

performing a first shared decoding on the received recovered compressed representation, using a first shared decoding layer having first shared decoding parameters;

performing a first adaptive decoding on the received recovered compressed representation, using a first adaptive decoding layer having first adaptive decoding parameters; and

combining the first shared decoded recovered compressed representation and the first adaptive decoded recovered compressed representation, to generate a second combined output.

16. The non-transitory computer-readable medium of claim 15 , wherein the instructions, when executed by the at least one processor, further cause the at least one processor to:

perform a second adaptive encoding on the first combined output, using a second adaptive encoding layer having second adaptive encoding parameters;

receive the recovered compressed representation and the hyperparameter; and

wherein the instructions, when executed by the at least one processor, further cause the at least one processor to:

perform a second shared decoding on the second combined output, using a second shared decoding layer having second shared decoding parameters; and

perform a second adaptive decoding on the second combined output, using a second adaptive decoding layer having second adaptive decoding parameters.

17. The non-transitory computer-readable medium of claim 16 , wherein the instructions, when executed by the at least one processor, further cause the at least one processor to:

generate a shared feature, based the received input image and the first shared encoding parameters;

generate estimated adaptive encoding parameters, based on one or more of the received input image, the first adaptive encoding parameters, the generated shared feature, and the received hyperparameter, using a prediction neural network; and

generate the compressed representation, based on the estimated adaptive encoding parameters and the received hyperparameter.

18. The non-transitory computer-readable medium of claim 17 , wherein the prediction neural network is trained by:

generating a first loss for training data corresponding to the received hyperparameter, and a second loss for validation data corresponding to the received hyperparameter, based on the received hyperparameter, the first shared encoding parameters, the first adaptive encoding parameters, the first shared decoding parameters, the first adaptive decoding parameters, and prediction parameters of the prediction neural network; and

updating the prediction parameters, based on gradients of the generated first loss and the generated second loss.

19. The non-transitory computer-readable medium of claim 16 , wherein the instructions, when executed by the at least one processor, further cause the at least one processor to:

generate a shared feature, based the received input image and the first shared decoding parameters;

generate estimated adaptive decoding parameters, based on one or more of the received input image, the first adaptive decoding parameters, the generated shared feature, and the received hyperparameter, using a prediction neural network; and

reconstruct the output image, based on the estimated adaptive decoding parameters and the received hyperparameter.

20. The non-transitory computer-readable medium of claim 19 , wherein the prediction neural network is trained by:

generating a first loss for training data corresponding to the received hyperparameter, and a second loss for validation data corresponding to the received hyperparameter, based on the received hyperparameter, the first shared encoding parameters, the first adaptive encoding parameters, the first shared decoding parameters, the first adaptive decoding parameters, and prediction parameters of the prediction neural network; and

updating the prediction parameters, based on gradients of the generated first loss and the generated second loss.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2021
From: JIANG, WEI; WANG, WEI; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 056736/0776 →
Continuity (2)
Provisional Application 63139156 · Jan 19, 2021
Related Publication 20220230362A1 · Jul 21, 2022