IP Library Granted Patent US 11,166,014
Granted Patent B2
US 11,166,014 · App. 16/772,443 · Granted Nov 2, 2021

Image encoding and decoding method and device using prediction network

Inventors: Seung-Hyun Cho (Daejeon, KR); Joo-Young Lee (Daejeon, KR); Youn-Hee Kim (Daejeon, KR); Jin-Wuk Seok (Daejeon, KR); Woong Lim (Daejeon, KR); Jong-Ho Kim (Daejeon, KR); Dae-Yeol Lee (Daejeon, KR); Se-Yoon Jeong (Daejeon, KR); Hui-Yong Kim (Daejeon, KR); Jin-Soo Choi (Daejeon, KR)
Assignee: Electronics and Telecommunications Research Institute
H04N19/105H04N19/159H04N19/176H04N19/85
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,166,014
App. No.
16/772,443
Granted
Nov 2, 2021
Kind
B2
Abstract

Disclosed herein are a method and apparatus for video decoding and a method and apparatus for video encoding. A prediction block for a target block is generated by predicting the target block using a prediction network, and a reconstructed block for the target block is generated based on the prediction block and a reconstructed residual block. The prediction network includes an intra-prediction network and an inter-prediction network and uses a spatial reference block and/or a temporal reference block when it performs prediction. For learning in the prediction network, a loss function is defined, and learning in the prediction network is performed based on the loss function.

Claims (49)

1. A decoding method, comprising:

performing prediction for a target block using a prediction network, thereby generating a prediction block for the target block; and

generating a reconstructed block for the target block based on the prediction block, wherein

a batch normalization layer is inserted in an input end of a hidden layer of the prediction network.

2. The decoding method of claim 1 , wherein the prediction network is an intra-prediction network that performs prediction for the target block using a spatial reference block.

3. The decoding method of claim 1 , wherein the prediction network is an inter-prediction network that performs prediction for the target block using a temporal reference block.

4. The decoding method of claim 3 , wherein the inter-prediction network performs prediction for the target block using a spatial reference block.

5. The decoding method of claim 1 , wherein:

the prediction network comprises multiple prediction networks, and

the multiple prediction networks are used respectively for multiple color channels, multiple block sizes, or multiple quantization parameters.

6. The decoding method of claim 1 , wherein:

the prediction network generates the prediction block by performing prediction for the target block using a reference block;

preprocessing is performed for a sample of the reference block before the sample of the reference block is input to the prediction network, and

postprocessing, which is a reverse process of the preprocessing, is performed for a sample of the prediction block when the sample of the prediction block is output from the prediction network.

7. The decoding method of claim 6 , wherein the preprocessing is mean subtraction, normalization, principal component analysis (PCA), or whitening.

8. The decoding method of claim 1 , wherein learning in the prediction network is performed for a single image.

9. The decoding method of claim 1 , wherein learning in the prediction network is performed for a batch.

10. A decoding method, comprising:

performing prediction for a target block using a prediction network, thereby generating a prediction block for the target block; and

generating a reconstructed block for the target block based on the prediction block, wherein

a loss function for learning in the prediction network is defined based on a prediction image and an original image, and

the loss function is defined based on a square of a difference between the prediction image and the original image.

11. A decoding method, comprising:

performing prediction for a target block using a prediction network, thereby generating a prediction block for the target block; and

generating a reconstructed block for the target block based on the prediction block, wherein

a loss function for learning in the prediction network is defined based on a prediction image and an original image, and

the loss function is defined based on an absolute value of a difference between the prediction image and the original image.

12. The decoding method of claim 11 , wherein a loss function for online update of the prediction network is determined based on a residual block.

13. The decoding method of claim 11 , wherein a loss function for online update of the prediction network is determined based on a reconstructed block.

14. A decoding method, comprising:

performing prediction for a target block using a prediction network, thereby generating a prediction block for the target block; and

generating a reconstructed block for the target block based on the prediction block,

wherein online update of a network parameter of the prediction network is continuously performed during decoding of video.

15. A decoding method, comprising:

performing prediction for a target block using a prediction network, thereby generating a prediction block for the target block; and

generating a reconstructed block for the target block based on the prediction block, wherein

a network parameter of the prediction network is initialized for a specified target, and

wherein the specified target is a slice, a picture, or a picture having an identifier that is different from a temporal identifier of a previous picture.

16. An encoding method, comprising:

performing prediction for a target block using a prediction network, thereby generating a prediction block for the target block; and

generating a reconstructed block for the target block based on the prediction block, wherein

online update of a network parameter of the prediction network is continuously performed during decoding of video.

17. A non-transitory computer-readable recording medium in which a bitstream for decoding of an image is stored, the bitstream comprising:

information about a target block,

wherein:

a prediction block for the target block is generated by performing prediction for the target block using a prediction network,

a reconstructed block for the target block is generated based on the information about the target block and the prediction block,

a loss function for learning in the prediction network is defined based on a prediction image and an original image, and

the loss function is defined based on an absolute value of a difference between the prediction image and the original image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2020
From: CHO, SEUNG-HYUN; LEE, JOO-YOUNG; KIM, YOUN-HEE; SEOK, JIN-WUK; LIM, WOONG; KIM, JONG-HO; LEE, DAE-YEOL; JEONG, SE-YOON; KIM, HUI-YONG; CHOI, JIN-SOO
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 052927/0852 →
Priority Claims (2)
KR 10-2017-0172537 · Dec 14, 2017 · national
KR 10-2018-0160775 · Dec 13, 2018 · national
Continuity (1)
Related Publication 20210084290A1 · Mar 18, 2021
Cited By (3)
US 12,225,233 US 12,354,312 US 12,689,770