IP Library › Granted Patent US 12,581,119
Granted Patent B2
US 12,581,119 · App. 18/679,679 · Granted Mar 17, 2026

Image encoding and decoding method and apparatus

Inventors: Tiansheng Guo (Beijing, CN); Jing Wang (Beijing, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
H04N19/61H04N19/136H04N19/147H04N19/184H04N19/124H04N19/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,581,119
App. No.
18/679,679
Granted
Mar 17, 2026
Kind
B2
Abstract

This application provides an image encoding and decoding method and apparatus. The image encoding method in this application includes: obtaining a to-be-processed first image feature; performing non-linear transformation processing on the to-be-processed first image feature to obtain a processed image feature, where the non-linear transformation processing sequentially includes a first non-linear operation, convolution processing, and an element-wise multiplication operation; and performing encoding based on the processed image feature to obtain a bitstream. This application can avoid limitation on a convolutional parameter, to implement efficient non-linear transformation processing in an encoding/decoding network, and further improve rate-distortion performance of an image/video compression algorithm.

Claims (38)

1 . An image encoding method, comprising:

obtaining a to-be-processed first image feature;

performing non-linear transformation processing on the to-be-processed first image feature to obtain a processed image feature, wherein the non-linear transformation processing sequentially comprises a first non-linear operation, convolution processing, and an element-wise multiplication operation; and

performing encoding based on the processed image feature to obtain a bitstream,

wherein the performing non-linear transformation processing on the to-be-processed first image feature to obtain a processed image feature comprises;

performing the first non-linear operation on each feature value in the to-be-processed first image feature to obtain a second image feature;

performing the convolution processing on the second image feature to obtain a third image feature, wherein a plurality of feature values in the third image feature correspond to a plurality of feature values in the to-be-processed first image feature; and

performing the element-wise multiplication operation on the plurality of corresponding feature values in the to-be-processed first image feature and the third image feature to obtain the processed image feature.

2 . The method according to claim 1 , wherein the first non-linear operation comprises an activation function of rectified linear unit series, Sigmoid, Tanh, or piecewise linear mapping.

3 . The method according to claim 1 , further comprising:

constructing a non-linear transformation unit in a training phase, wherein the non-linear transformation unit in the training phase comprises a first non-linear operation layer, a convolution processing layer, and an element-wise multiplication operation layer; and

performing training based on pre-obtained training data to obtain a trained non-linear transformation unit for implementing the non-linear transformation processing.

4 . An image encoding device, comprising:

one or more processors; and

a non-transitory computer-readable storage medium, coupled to the one or more processors and storing instructions, which, when executed by the one or more processors, cause the image encoding device to perform operations comprising:

obtaining a to-be-processed first image feature;

performing non-linear transformation processing on the to-be-processed first image feature to obtain a processed image feature, wherein the non-linear transformation processing sequentially comprises a first non-linear operation, convolution processing, and an element-wise multiplication operation; and

performing encoding based on the processed image feature to obtain a bitstream,

wherein the performing non-linear transformation processing on the to-be-processed first image feature to obtain processed image feature comprises:

performing the first non-linear operation on each feature value in the to-be-processed first image feature to obtain a second image feature;

performing the convolution processing on the second image feature to obtain a third image feature, wherein a plurality of features values in the third image feature correspond to a plurality of feature values in the to-be-processed first image feature, and

performing the element-wise multiplication operation on the plurality of corresponding feature values in the to-be-processed first image feature and the third image feature to obtain the processed image feature.

5 . The device according to claim 4 , wherein the first non-linear operation comprises an activation function of rectified linear unit series, Sigmoid, Tanh, or piecewise linear mapping.

6 . The device according to claim 4 , wherein the operations further comprise:

constructing a non-linear transformation unit in a training phase, wherein the non-linear transformation unit in the training phase comprises a first non-linear operation layer, a convolution processing layer, and an element-wise multiplication operation layer; and

performing training based on pre-obtained training data to obtain a trained non-linear transformation unit for implementing the non-linear transformation processing.

7 . A non-transitory computer-readable storage medium, comprising instructions, wherein when the instructions are run on a computer, cause the computer to perform operations comprising:

obtaining a to-be-processed first image feature;

performing non-linear transformation processing on the to-be-processed first image feature to obtain a processed image feature, wherein the non-linear transformation processing sequentially comprises a first non-linear operation, convolution processing, and an element-wise multiplication operation; and

performing encoding based on the processed image feature to obtain a bitstream,

wherein the performing non-linear transformation processing on the to-be-processed first image feature to obtain a processed image feature comprises:

performing the first non-linear operation on each feature value in the to-be processed first image feature to obtain a second image feature;

performing the convolution processing on the second image feature to obtain a third image feature, wherein a plurality of feature values in the third image feature correspond to a plurality of feature values in the to-be-processed first image feature; and

performing the element-wise multiplication operation on the plurality of corresponding feature values in the to-be-processed first image feature and the third image feature to obtain the processed image feature.

8 . The non-transitory computer-readable storage medium according to claim 7 , wherein the first non-linear operation comprises an activation function of rectified linear unit series, Sigmoid, Tanh, or piecewise linear mapping.

9 . The non-transitory computer-readable storage medium according to claim 7 , wherein the operations further comprise:

constructing a non-linear transformation unit in a training phase, wherein the non-linear transformation unit in the training phase comprises a first non-linear operation layer, a convolution processing layer, and an element-wise multiplication operation layer; and

performing training based on pre-obtained training data to obtain a trained non-linear transformation unit for implementing the non-linear transformation processing.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2024
From: GUO, TIANSHENG; WANG, JING
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 067943/0799 →
Priority Claims (1)
CN 202111470979.5 · Dec 3, 2021 · national
Continuity (2)
Continuation PCTCN2022135204 · Nov 30, 2022
Related Publication 20240323441A1 · Sep 26, 2024
References Cited (31)
US 20050276475A1 · Sawada · 2005 [cited by applicant]
US 20180075343A1 · van den Oord · 2018 [cited by examiner]
US 20200193296A1 · Dixit et al. · 2020 [cited by applicant]
US 20210012181A1 · Zhu et al. · 2021 [cited by applicant]
US 20210074036A1 · Fuchs et al. · 2021 [cited by applicant]
US 20210109966A1 · Ayush · 2021 [cited by examiner]
US 20210117687A1 · Ren et al. · 2021 [cited by applicant]
US 20230334829A1 · Du · 2023 [cited by examiner]
US 20240289590A1 · Cricrì · 2024 [cited by examiner]
CN 107736027A · 2018 [cited by applicant]
CN 110969626A · 2020 [cited by applicant]
CN 111243066A · 2020 [cited by applicant]
CN 111259904A · 2020 [cited by applicant]
CN 111754592A · 2020 [cited by applicant]
CN 113569790A · 2021 [cited by applicant]
CN 113709455A · 2021 [cited by applicant]
JP 2020028111A · 2020 [cited by applicant]
JP 2021520082A · 2021 [cited by applicant]
JP 2021522756A · 2021 [cited by applicant]
JP 2024532014A · 2024 [cited by applicant]
TW 202027033A · 2020 [cited by applicant]
WO 2021077620A1 · 2021 [cited by applicant]
Mohammad Akbari; Jie Liang; Jingning Han; Chengjie Tu:“ Learned Variable-Rate Image Compression with Residual Divisive Normalization”, 2020 IEEE International Conference on Multimedia and Expo (ICME), London, 2020, tota… [cited by applicant]
Johannes Ball , Valero Laparra and Eero P. Simoncelli: “Density Modeling of Images Using Ageneralized Normalization Transformation”, International Conference on Learning Representations (ICRL) , 2016, total 14 pages. [cited by applicant]
Johannes Ball : “End-to-end optimized image compression”, International Conference on Learning Representations (ICRL) , Toulon, 2017, total 43 pages. [cited by applicant]
Lei Zhou, Zhenhong Sun, Xiangji Wu, Junmin Wu: “End-to-end Optimized Image Compression with Attention Mechanism”, 2019, total 4 pages. [cited by applicant]
International Search Report and Written Opinion issued in PCT/CN2022/135204, dated Feb. 21, 2023, 8 pages. [cited by applicant]
Office Action issued in TW111146343, dated May 10, 2023, 7 pages. [cited by applicant]
Ge Gao et al., “Neural Image Compression via Attentional Multi-scale Back Projection and Frequency Decomposition”, IEEE International Conference on Computer Vision, Oct. 10, 2021, total 10 pages,XP034092432. [cited by applicant]
Ming Lu et al., “Decomposition, Compression, and Synthesis (DCS)-based Video Coding:A Neural Exploration via Resolution-Adaptive Learning”, arXiv:2012.00650v1 [cs.CV] Dec. 1, 2020, total 15 pages, XP081825519. [cited by applicant]
Tong Chen et al., “Neural Image Compression via Non-Local Attention Optimization and Improved Context Modeling”, arXiv:1910.06244v1 [eess.IV] Oct. 11, 2019, total 13 pages, XP081514975. [cited by applicant]