IP Library Granted Patent US 12,732,625
Granted Patent B2
US 12,732,625 · App. 18/338,143 · Granted Sep 8, 2026

Method and apparatus for encoding a picture and decoding a bitstream using a neural network

Inventors: Elena Alexandrovna Alshina (Munich, DE); Han Gao (Shenzhen, CN); Semih Esenlik (Munich, DE)
Assignee: Huawei Technologies Co., Ltd.
H04N19/42H04N19/132H04N19/172H04N19/187H04N19/463H04N19/80H04N19/85
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,732,625
App. No.
18/338,143
Granted
Sep 8, 2026
Kind
B2
Abstract

A method and an apparatus for encoding a picture and decoding a bitstream representing a matrix using a neural network is provided. The method includes obtaining a resizing method out of a plurality of resizing methods. The resizing method is applied to resize an input of size S to a size S . The resized input is processed by the neural network, wherein the neural network comprises one or more downsampling layers. Subsequently, the method provides an output of the neural network, the output having a size P that is smaller than S in the at least one dimension.

Claims (26)

1 . A method of decoding a bitstream representing a picture using a neural network (NN) to process an input representing a matrix having a size in at least one dimension, wherein the method comprises:

obtaining a resizing method out of a plurality of resizing methods, wherein obtaining the resizing method is based on one or more indications, wherein the one or more indications comprise a third indication, wherein the third indication has a value that indicates an interpolation filter that is to be used in the interpolation, wherein, the third indication is or comprises an index indicating an entry in a first lookup table (LUT) that has a plurality of entries and each entry in the first LUT specifies the interpolation filter, wherein obtaining the resizing method comprises comparing a larger size with another size and obtaining, based on the comparing, the resizing method, wherein the another size is obtained from a function comprising a combined upsampling parameter of the NN, the upsampling parameter being a product of upsampling ratios of all upsampling layers of the NN, wherein each upsampling layer has a different upsampling ratio;

processing the input with the size by the NN, wherein the NN comprises one or more upsampling layers, thereby obtaining an intermediate output having the larger size that is larger than the size in the at least one dimension; and

resizing the intermediate output from the larger size to the another size by applying the obtained resizing method, thereby obtaining a decoded picture.

2 . The method according to claim 1 , wherein the step of obtaining the resizing method comprises determining the resizing method from the plurality of resizing methods based on information relating to at least one of the input, the one or more layers of the NN, an output to be provided by the NN, or the decoded picture.

3 . The method according to claim 1 , wherein the plurality of resizing methods comprises padding, padding with zeros, reflection padding, repetition padding, cropping, interpolation to increase the larger size of the intermediate output to the another size, interpolation to decrease the larger size of the intermediate output to the another size.

4 . The method according to claim 1 , wherein the another size is further obtained from:

(ii) a product of the size and the unsampling ratios of all unsampling layers of the neural network; or

(iii) the bitstream or an additional bitstream or an index in the bitstream or an index in the additional bitstream and the index indicating an entry in a table, wherein the table is a lookup table (LUT), comprising a plurality of entries and each entry indicates the another size, wherein the method further comprises obtaining the another size using the index.

5 . The method according to claim 1 , wherein:

based on determining that the larger size is not equal to the another size, the resizing method is applied; or

based on determining that the larger size is smaller than the another size, the resizing method is applied that increases the larger size; or

based on determining that the larger size is larger than the another size, the resizing method is applied that decreases the larger size.

6 . The method according claim 2 , wherein the one or more indications comprise at least one of:

a first indication being or comprising a first flag having a size of 1 bit, wherein a first value of the first indication indicates that padding or cropping is to be applied as the resizing method and a second value of the first indication indicates that interpolation is to be applied as the resizing method;

a second indication, wherein the second indication has a first value that indicates that the larger size to be increased and a second value that indicates that the larger size is to be decreased;

a fourth indication, being or comprising a second flag having a size of 1 bit, the fourth indication having a first value that indicates that padding is to be applied as the resizing method and a second value that indicates that cropping is to be applied as the resizing method;

a fifth indication having a first value that indicates whether padding with zeros, reflection padding or repetition padding is to be applied as the resizing method; and

a sixth indication that is or comprises an index indicating an entry in a second look-up table (LUT) wherein the second LUT comprises a plurality of entries, wherein each entry specifies a resizing method.

7 . A decoder for decoding a bitstream representing a picture, wherein the decoder comprises one or more processors for implementing a neural network (NN) to process an input representing a matrix having a size in at least one dimension, wherein the one or more processors are adapted to perform a method comprising:

obtaining a resizing method out of a plurality of resizing methods, wherein obtaining the resizing method is based on one or more indications, wherein the one or more indications comprise a third indication, wherein the third indication has a value that indicates an interpolation filter that is to be used in the interpolation, wherein, the third indication is or comprises an index indicating an entry in a first lookup table (LUT) that has a plurality of entries and each entry in the first LUT specifies the interpolation filter, wherein obtaining the resizing method comprises comparing a larger size with another size and obtaining, based on the comparing, the resizing method, wherein the another size is obtained from a function comprising a combined upsampling parameter of the NN, the upsampling parameter being a product of upsampling ratios of all upsampling layers of the NN, wherein each upsampling layer has a different upsampling ratio;

processing the input with the size by the NN, wherein the NN comprises one or more upsampling layers, thereby obtaining an intermediate output having the larger size that is larger than the size in the at least one dimension; and

resizing the intermediate output from the larger size to the another size by applying the obtained resizing method, thereby obtaining a decoded picture.

8 . The method according to claim 1 , wherein the NN further comprises one or more downsampling layers configured to apply a downsampling by a certain factor.

9 . The method according to claim 8 , wherein the NN further comprises an inverse generalized divisive normalization (IGDN).

10 . The method according to claim 2 , wherein the upsampling parameter includes information regarding upsampling ratios of a decoder of the NN.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2024
From: ALSHINA, ELENA ALEXANDROVNA; GAO, HAN; ESENLIK, SEMIH
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 068223/0451 →
Continuity (2)
Continuation PCTEP2020087332 · Dec 18, 2020
Related Publication 20230353766A1 · Nov 2, 2023
References Cited (24)
US 11218696B2 · Dorovic · 2022 [cited by examiner]
US 11677948B2 · Besenbruch · 2023 [cited by examiner]
US 12001935B2 · Sjögren · 2024 [cited by examiner]
US 20030067973A1 · Lee · 2003 [cited by examiner]
US 20180091825A1 · Zhao · 2018 [cited by examiner]
US 20190096049A1 · Kim · 2019 [cited by examiner]
US 20190251707A1 · Gupta · 2019 [cited by examiner]
US 20190297326A1 · Reda · 2019 [cited by examiner]
US 20200145661A1 · Jeon · 2020 [cited by examiner]
US 20200162751A1 · Kim · 2020 [cited by examiner]
US 20200175290A1 · Raja · 2020 [cited by examiner]
US 20210044804A1 · Lu · 2021 [cited by examiner]
US 20210127101A1 · Roh · 2021 [cited by examiner]
US 20210241429A1 · Pan · 2021 [cited by examiner]
US 20210319420A1 · Yu · 2021 [cited by examiner]
US 20210350163A1 · Ojard · 2021 [cited by examiner]
US 20220094977A1 · Kim · 2022 [cited by examiner]
Sze et al., “High Efficiency Video Coding (HEVC): Algorithms and Architectures,” Integrated Circuits and Systems, Springer, Total 384 pages (Sep. 2014). [cited by applicant]
Hashemi, “Enlarging smaller images before inputting into convolutional neural network: zero-padding vs. interpolation,” Journal of Big Data, vol. 6, Springer Open, XP055725533, Total 14 pages (Nov. 14, 2019). [cited by applicant]
Hassaballah et al., “Recent Advances in Computer Vision: Theories and Applications,” Theories and Applications, Springer, Total 430 pages (Jan. 2019). [cited by applicant]
Theis et al., “Lossy Image Compression with Compressive Autoencoders,” Published as a conference paper at ICLR 2017, arxiv.org, Cornell University Library, 201 Online Library Cornell University Ithaca, NY 14853, XP08075… [cited by applicant]
Balle et al., “Variational Image Compression With a Scale Hyperprior,” arXiv:1802.01436v2 [eess.IV], Total 23 pages (May 1, 2018). [cited by applicant]
Ozer et al., “Noise robust sound event classification with convolutional neural network,” Neurocomputing, vol. 272, Elsevier, XP085275968, Accepted Jul. 2017, Total 8 pages (Jan. 2018). [cited by applicant]
Guo et al., “3-D Context Entropy Model for Improved Practical Image Compression,” IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Total 4 pages (Apr. 2020). [cited by applicant]