IP Library › Granted Patent US 12,278,957
Granted Patent B2
US 12,278,957 · App. 18/127,616 · Granted Apr 15, 2025

Video encoding and decoding methods, encoder, decoder, and storage medium

Inventors: Yanzhuo Ma (Dongguan, CN); Ruipeng Qiu (Dongguan, CN); Junyan Huo (Dongguan, CN); Shuai Wan (Dongguan, CN); Fuzheng Yang (Dongguan, CN)
Assignee: GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP., LTD.
H04N19/117G06V10/82H04N19/17H04N19/184
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,278,957
App. No.
18/127,616
Granted
Apr 15, 2025
Kind
B2
Abstract

A video encoding method is applicable to an encoder, and comprises: determining side information having a correlation with a current coding unit; filtering first relevant information to be coded of the current coding unit by using a preset network model and the side information, to obtain second relevant information to be coded; and inputting the second relevant information to be coded into a subsequent coding module for encoding, to obtain a bitstream.

Claims (52)

1. A method for encoding a video, applicable to an encoder, comprises:

determining side information having a correlation with a current coding unit;

filtering first relevant information to be coded of the current coding unit by using a preset network model and the side information, to obtain second relevant information to be coded; and

inputting the second relevant information to be coded into a subsequent coding module for encoding, to obtain a bitstream;

wherein the preset network model comprises a first network model, and the first network model comprises at least a first neural network model and a first adder; and

the first neural network model comprises: a first convolution layer, at least one second convolution layer, first residual layers, second residual layers, average pooling layers, and sampling rate conversion modules; wherein

the first convolution layer is followed by at least two of the first residual layers connected in series, and an average pooling layer is connected between adjacent first residual layers;

at least one of the first residual layers is followed by at least two of the second residual layers connected in series, and a sampling rate conversion module is connected between adjacent second residual layers; and

at least one of the second residual layers is connected in series with the at least one second convolution layer;

wherein filtering the first relevant information to be coded of the current coding unit by using the preset network model and the side information, to obtain the second relevant information to be coded comprises:

obtaining the first decoding information corresponding to first information to be coded of the current coding unit; and

inputting the first decoding information and the side information into the first network model to output the second decoding information, wherein the first network model is used for performing enhancement on the first decoding information by using the side information;

wherein part of input ends of the first network model is shorted to an output end of the first network model, and an input end for the first decoding information in the first network model is shorted to the output end.

2. The method of claim 1 , wherein the coding unit is a picture or an area in the picture.

3. The method of claim 1 , wherein the first relevant information to be coded comprises first decoding information, and the second relevant information to be coded comprises second decoding information.

4. The method of claim 3 , wherein

inputting the second relevant information to be coded into the subsequent coding module for encoding, to obtain the bitstream comprises:

reconstructing the current coding unit by using the second decoding information, to obtain a reconstructed unit of the current coding unit; and

performing subsequent coding according to the reconstructed unit of the current coding unit, to obtain the bitstream.

5. The method of claim 1 , wherein inputting the first decoding information and the side information into the first network model to output the second decoding information, comprises:

inputting the first decoding information and the side information into the first neural network model to output a first intermediate value; and

adding the first intermediate value to the first decoding information by the first adder, to obtain the second decoding information.

6. A method for decoding a video, applicable to a decoder, comprises:

decoding a bitstream to obtain information to be decoded;

inputting the information to be decoded into a first decoding module to output first decoding information of a current coding unit;

determining side information having a correlation with the first decoding information;

filtering the first decoding information by using a preset network model and the side information, to obtain second decoding information; and

reconstructing the current coding unit by using the second decoding information, to obtain a reconstructed unit of the current coding unit;

wherein the preset network model comprises a first network model, and the first network model comprises at least a first neural network model and a first adder; and

the first neural network model comprises: a first convolution layer, at least one second convolution layer, first residual layers, second residual layers, average pooling layers, and sampling rate conversion modules; wherein

the first convolution layer is followed by at least two of the first residual layers connected in series, and an average pooling layer is connected between adjacent first residual layers;

at least one of the first residual layers is followed by at least two of the second residual layers connected in series, and a sampling rate conversion module is connected between adjacent second residual layers; and

at least one of the second residual layers is connected in series with the at least one second convolution layer;

wherein filtering the first decoding information by using the preset network model and the side information, to obtain the second decoding information comprises:

inputting the first decoding information and the side information into the first network model to output the second decoding information;

wherein the first network model is used for performing filtering on the first decoding information according to the correlation between the first decoding information and the side information, part of input ends of the first network model is shorted to an output end of the first network model, and an input end for the first decoding information in the first network model is shorted to the output end.

7. The method of claim 6 , wherein the coding unit is a picture or an area in the picture.

8. The method of claim 6 , wherein inputting the first decoding information and the side information into the first network model to output the second decoding information comprises:

inputting the first decoding information and the side information into the first neural network model to output a first intermediate value; and

adding the first intermediate value to the first decoding information by the first adder, to obtain the second decoding information.

9. A decoder, comprises: a processor and a memory for storing a computer program executable by the processor, wherein the processor is configured to execute the computer program to:

decode a bitstream to obtain information to be decoded, and input the information to be decoded into a first decoding module to output first decoding information of a current coding unit;

determine side information having a correlation with the first decoding information;

filter the first decoding information by using a preset network model and the side information, to obtain second decoding information; and

reconstruct the current coding unit by using the second decoding information, to obtain a reconstructed unit of the current coding unit;

wherein the preset network model comprises a first network model, and the first network model comprises at least a first neural network model and a first adder; and

the first neural network model comprises: a first convolution layer, at least one second convolution layer, first residual layers, second residual layers, average pooling layers, and sampling rate conversion modules; wherein

the first convolution layer is followed by at least two of the first residual layers connected in series, and an average pooling layer is connected between adjacent first residual layers;

at least one of the first residual layers is followed by at least two of the second residual layers connected in series, and a sampling rate conversion module is connected between adjacent second residual layers; and

at least one of the second residual layers is connected in series with the at least one second convolution layer;

wherein the processor is further configured to execute the computer program to input the first decoding information and the side information into the first network model to output the second decoding information;

wherein the first network model is used for performing filtering on the first decoding information according to the correlation between the first decoding information and the side information, part of input ends of the first network model is shorted to an output end of the first network model, and an input end for the first decoding information in the first network model is shorted to the output end.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: MA, YANZHUO; QIU, RUIPENG; HUO, JUNYAN; WAN, SHUAI; YANG, FUZHENG
To: GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP., LTD.
Reel/Frame 063137/0909 →
Continuity (2)
Continuation PCTCN2020119732 · Sep 30, 2020
Related Publication 20230239470A1 · Jul 27, 2023
References Cited (32)
US 10735735B2 · Andersson · 2020 [cited by examiner]
US 10803583B2 · Wang · 2020 [cited by examiner]
US 10869036B2 · Coelho · 2020 [cited by examiner]
US 11196992B2 · Huang · 2021 [cited by examiner]
US 11494658B2 · Chen · 2022 [cited by examiner]
US 11689726B2 · Mukherjee · 2023 [cited by examiner]
US 20200162789A1 · Ma · 2020 [cited by examiner]
US 20200213587A1 · Galpin et al. · 2020 [cited by applicant]
US 20200252654A1 · Su et al. · 2020 [cited by applicant]
US 20200280717A1 · Li · 2020 [cited by examiner]
US 20200404340A1 · Xu · 2020 [cited by applicant]
US 20210021823A1 · Na et al. · 2021 [cited by applicant]
US 20210092413A1 · Tsukuba · 2021 [cited by examiner]
US 20230300354A1 · Li · 2023 [cited by examiner]
CN 110971915A · 2020 [cited by applicant]
CN 111133756A · 2020 [cited by applicant]
CN 111194555A · 2020 [cited by applicant]
CN 111464815A · 2020 [cited by applicant]
WO 2019154152A1 · 2019 [cited by applicant]
WO 2019194425A1 · 2019 [cited by applicant]
WO 2020062074A1 · 2020 [cited by applicant]
Guo Lu, et al. “DVC: An End-to-end Deep Video Compression Framework”, Apr. 7, 2019. [cited by applicant]
J. Balle, V. Laparra, and E. P. Simoncelli, “End-to-end optimized image compression” , arXiv preprint, arXiv:1611.01704, Nov. 5, 2016. 1, 2, 4, 5. [cited by applicant]
J. Balle, D. Minnen, S. Singh, S. J. Hwang, and N. Johnston, “Variational image compression with a scale hyperprior”, arXiv preprint arXiv:1802.01436, 2018. 1, 2, 5, 8. [cited by applicant]
International Search Report in the international application No. PCT/CN2020/119732, mailed on Jul. 7, 2021. [cited by applicant]
Written Opinion of the International Searching Authority in the international application No. PCT/CN2020/119732, mailed on Jul. 7, 2021. [cited by applicant]
Lu Guo et al: “An End-to-End Learning Framework for Video Compression”, IEEE Transactions on Pattern Analysis and Machine Intelligence, IEEE Computer Society, USA, vol. 43, No. 10, Apr. 20, 2020 (Apr. 20, 2020), pp. 329… [cited by applicant]
Huo Shuai et al: “Convolutional Neural Network-Based Motion Compensation Refinement for Video Coding”, 2018 IEEE International Symposium on Circuits and Systems (ISCAS), IEEE, May 27, 2018 (May 27, 2018), pp. 1-4, XP033… [cited by applicant]
Siwei Ma et al: “Image and Video Compression With Neural Networks: A Review”, Apr. 10, 2019 (Apr. 10, 2019), pp. 1-16, XP055765818, the whole document. 16 pages. [cited by applicant]
Supplementary European Search Report in the European application No. 20955826.1, mailed on Oct. 9, 2023. 10 pages. [cited by applicant]
Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia, Margret Keuper, Alexey Dosovitskiy, and Thomas Brox. “Flownet 2.0: Evolution of optical flow estimation with deep networks”. Proceedings of the IEEE Conference on Computer Vision… [cited by applicant]
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. “U-net: Convolutional networks for biomedical image segmentation”. Medical Image Computing and Computer Assisted Intervention-MICCAI, pp. 234-241. Springer, 2015. 8 pa… [cited by applicant]