IP Library Granted Patent US 12,586,255
Granted Patent B2
US 12,586,255 · App. 18/479,507 · Granted Mar 24, 2026

Configurable positions for auxiliary information input into a picture data processing neural network

Inventors: Timofey Mikhailovich Solovyev (Munich, DE); Biao Wang (Shenzhen, CN); Elena Alexandrovna Alshina (Munich, DE); Han Gao (Shenzhen, CN); Panqi Jia (Munich, DE); Esin Koyuncu (Munich, DE); Alexander Alexandrovich Karabutov (Munich, DE); Mikhail Vyacheslavovich Sosulnikov (Munich, DE); Semih Esenlik (Munich, DE); Sergey Yurievich Ikonin (Moscow, RU)
Assignee: Huawei Technologies Co., Ltd.
G06T9/002G06T3/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,255
App. No.
18/479,507
Granted
Mar 24, 2026
Kind
B2
Abstract

Methods and apparatuses for processing of picture data or picture feature data using a neural network with two or more layers are provided. The present disclosure may be applied in the field of artificial intelligence (AI)-based video or picture compression technologies, and in particular, to the field of neural network-based video compression technologies. According to the present disclosure, position within the neural network, at which auxiliary information may be entered for processing is selectable based on a gathering condition. The gathering condition may assess whether some prerequisite is fulfilled. Advantages may include better performance in terms of rate and/or disclosure due to the effect of increased flexibility in neural network configurability.

Claims (78)

1 . A method for processing picture feature data from a bitstream using a neural network comprising a plurality of neural network layers, the method comprising:

obtaining the picture feature data from the bitstream;

processing the picture feature data using the neural network, wherein for each of one or more preconfigured positions within the neural network, the processing the picture feature data comprises:

determining, based on a gathering condition, whether to gather auxiliary data for processing by one of the plurality of neural network layers at the preconfigured position,

in response to determining to gather the auxiliary data, processing the picture feature data with the neural network layer at the preconfigured position based on the auxiliary data.

2 . The method according to claim 1 , wherein after applying the gathering condition in determining to gather the auxiliary data, the auxiliary data is to be gathered for a single one of the one or more preconfigured positions.

3 . The method according to claim 1 , wherein after applying the gathering condition in determining to gather the auxiliary data, the auxiliary data is to be gathered for more than one of the preconfigured positions.

4 . The method according to claim 1 , wherein

there are more than one of the preconfigured positions;

the auxiliary data is scalable in size to match dimensions of an input channel processed by the layer at two or more of the preconfigured positions; and

after applying the gathering condition in determining to gather the auxiliary data, the auxiliary data is

i) gathered or

ii) gathered and scaled

for a single one of the preconfigured positions.

5 . The method according to claim 1 , wherein the gathering condition is based on a picture characteristic or a picture feature data characteristic obtained from the bitstream.

6 . The method according to claim 5 , wherein

the picture characteristic or the picture feature data characteristic includes resolution; and

the gathering condition includes a comparison of the resolution with a preconfigured resolution threshold.

7 . The method according to claim 5 , wherein

the picture is a video picture and the picture characteristic includes a picture type; and

the gathering condition includes determining whether the picture type is a temporally predicted picture type or spatially predicted picture type.

8 . The method according to claim 1 , further comprising:

obtaining from the bitstream an indication specifying for the one or more preconfigured positions whether or not to gather the auxiliary data,

the gathering condition for each of the one or more preconfigured positions is as follows:

based on the indication specifies for the preconfigured position that the auxiliary data being to be gathered, the determination is affirmative; and

based on the indication specifies for the preconfigured position that the auxiliary data being not to be gathered, the determination is negative.

9 . The method according to claim 1 , wherein the auxiliary data provides information about the picture feature data processed by the neural network to generate an output.

10 . The method according to claim 1 , wherein the auxiliary data includes prediction data which is a prediction of the picture or a prediction of picture feature data after processing by one or more of the layers of the neural network;

wherein the auxiliary data are a coupled pair of the prediction data and supplementary data to be combined with the prediction data; and

wherein the prediction data and the supplementary data have dimensions of data processed by layers at mutually different positions in the neural network.

11 . The method according to claim 1 , wherein

the neural network includes a sub-network for lossless decoding with at least one layer; and

the auxiliary data is input into the sub-network for lossless decoding.

12 . The method according to claim 1 , wherein the neural network is trained to perform at least one of still picture decoding, video picture decoding, still picture filtering, video picture filtering, and machine vision processing including object detection, object recognition or object classification.

13 . The method according to claim 1 , wherein the method is performed for each of a plurality of auxiliary data, including first auxiliary data and second auxiliary data,

wherein the first auxiliary data is associated with a first set of one or more preconfigured positions and the second auxiliary data is associated with a second set of one or more preconfigured positions; and

wherein the first set of one or more preconfigured positions and the second set of one or more preconfigured positions share at least one preconfigured position.

14 . The method according to claim 1 ,

wherein the neural network is trained to perform the processing of video pictures; and

wherein the determining whether to gather the auxiliary data for processing by the neural network layer at the preconfigured position is performed on every predetermined number of video pictures, wherein the predetermined number of video pictures is one or more.

15 . A method for processing a picture with a neural network comprising a plurality of neural network layers to generate a bitstream, the method comprising:

processing the picture with the neural network, wherein for each of one or more preconfigured positions within the neural network, the processing the picture comprises:

determining, based on a gathering condition, whether to gather auxiliary data for processing by a layer at the preconfigured position,

in response to determining to gather the auxiliary data, processing the picture with the layer at the preconfigured position based on the auxiliary data; and

inserting into the bitstream data obtained through processing the picture by the neural network.

16 . The method according to claim 15 , wherein after applying the gathering condition in determining to gather the auxiliary data, the auxiliary data is to be gathered for a single one of the one or more preconfigured positions.

17 . The method according to claim 15 , wherein after applying the gathering condition in determining to gather the auxiliary data, the auxiliary data is to be gathered for more than one of the preconfigured positions.

18 . The method according to claim 15 , wherein

there are more than one of the preconfigured positions;

the auxiliary data is scalable in size to match dimensions of an input channel processed by the layer at two or more of the preconfigured positions; and

after applying the gathering condition in determining to gather the auxiliary data, the auxiliary data is

i) gathered or

ii) gathered and scaled

for a single one of the preconfigured positions.

19 . The method according to claim 15 , wherein the gathering condition is based on a picture characteristic or a picture feature data characteristic which is included into the bitstream.

20 . The method according to claim 19 , wherein

the picture characteristic or the picture feature data characteristic includes resolution; and

the gathering condition includes a comparison of the resolution with a preconfigured resolution threshold.

21 . The method according to claim 19 , wherein

the picture is a video picture and the picture characteristic includes a picture type; and

the gathering condition includes determining whether the picture type is a temporally predicted picture type or spatially predicted picture type.

22 . The method according to claim 15 , further comprising:

generating the indication specifying for the one or more preconfigured positions whether to gather the auxiliary data;

including the indication into the bitstream; and

selecting for the one or more preconfigured positions whether to gather the auxiliary data based on an optimization of a cost function including at least one of rate, distortion, accuracy, or complexity.

23 . A non-transitory computer-readable medium including code instructions which upon execution by one or more processors cause the one or more processors to perform the method according to claim 1 .

24 . An apparatus for processing picture feature data from a bitstream using a neural network comprising a plurality of neural network layers, the apparatus comprising:

a processing circuitry configured to:

obtain the picture feature data from the bitstream;

process the picture feature data using the neural network, wherein for each of one or more preconfigured positions within the neural network, the processing the picture feature data comprises:

determine, based on a gathering condition, whether to gather auxiliary data for processing by one of the plurality of neural network layers at the preconfigured position,

in response to determining to gather the auxiliary data, process the picture feature data with the neural network layer at the preconfigured position based on the auxiliary data.

25 . An apparatus for processing a picture with a neural network comprising a plurality of neural network layers to generate a bitstream, the apparatus comprising:

a processing circuitry configured to:

process the picture with the neural network, wherein for each of one or more preconfigured positions within the neural network, the processing the picture comprises:

determine, based on a gathering condition, whether to gather auxiliary data for processing by a layer at the preconfigured position,

in response to determining to gather the auxiliary data, process the picture with the layer at the preconfigured position based on the auxiliary data; and

insert into the bitstream data obtained through processing the picture by the neural network.

Continuity (2)
Continuation PCTRU2021000136 · Apr 1, 2021
Related Publication 20240037802A1 · Feb 1, 2024
References Cited (20)
US 20210014531A1 · Pfaff et al. · 2021 [cited by applicant]
US 20210120255A1 · Fang · 2021 [cited by examiner]
US 20210409783A1 · Wan et al. · 2021 [cited by applicant]
US 20240187640A1 · Racape · 2024 [cited by examiner]
JP 2022528604A · 2022 [cited by applicant]
JP 2024510433A · 2024 [cited by applicant]
WO 2020177133A1 · 2020 [cited by applicant]
WO 2022211658A1 · 2022 [cited by applicant]
Zhao et al., Enhanced CTU-Level Inter Prediction With Deep Frame Rate Up-Conversion for High Efficiency Video Coding, 2018 25th IEEE International Conference on Image Processing (ICIP) (Year: 2018). [cited by examiner]
Lu et al., “Tests on Decomposition, Compression, Synthesis (DCS)-based Technology,” JVET-U0096, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 21th Meeting by teleconference, Total 9 pages … [cited by applicant]
Zhao et al., “Enhanced Ctu-Level Inter Prediction with Deep Frame Rate Up-Conversion for High Efficiency Video Coding,” 2018 25th IEEE International Conference on Image Processing (ICIP), XP033455036, Total 5 pages (Oct… [cited by applicant]
“Series H:Audiovisual and Multimedia Systems, Infrastructure of audiovisual services-Coding of moving video, High efficiency video coding,” Telecommunication Standardization Sector of ITU, ITU-T, H.265, Total 664 pages … [cited by applicant]
Ma et al., “Image and Video Compression With Neural Networks: A Review,” IEEE Transactions on Circuits and Systems for Video Technology, XP055765818, Total 16 pages (Apr. 2019). [cited by applicant]
Boyce et al., “JVET common test conditions and software reference configurations,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG11, 10th Meeting, San Diego, JVET-J1010-v1, Total 6 pages … [cited by applicant]
Shen et al., “SHVC CU Processing Aided by a Feedforward Neural Network,” IEEE Transactions on Industrial Informatics, vol. 15, No. 11, Total 13 pages, XP011754874 (Nov. 2019). [cited by applicant]
Cui et al., “G-VAE: A Continuously Variable Rate Deep Image Compression Framework,” arXiv:2003.02012v2, total 8 pages (Apr. 2020). [cited by applicant]
Egmont-Petersen et al., “Image processing with neural networks-a review,” The Journal of the Pattern Recognition Society, Pattern Recognition 35, XP004366785, Total 23 pages (Oct. 2002). [cited by applicant]
Cooley et al., “An algorithm for the machine calculation of complex Fourier series,” American Mathematical Society, Mathematics of Computation, vol. 19, No. 90, Total 6 pages (Apr. 1965). [cited by applicant]
Schiopu et el., “CNN-Based Intra-Prediction for Lossless Hevc,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, No. 7, XP011796774, total 13 pages (Jul. 2020). [cited by applicant]
Alshina et al., “Description of Exploration Experiments on NN-based video coding,” Document: JVET-T2023_r2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, XP030293658, Total 12 pages (Oct. … [cited by applicant]