Modeling-based image decoding method and device in image coding system
According to the present invention, an image decoding method performed by a decoding device comprises the steps of: deriving a prediction mode of a current block; generating a prediction block of the current block on the basis of the prediction mode of the current block; deriving a neighboring block of the current block; deriving correlation information on the basis of a reconstruction block of the neighboring block and a prediction block of the neighboring block; deriving a modified prediction block of the current block on the basis of the prediction block of the current block and the correlation information; and deriving a reconstruction block of the current block on the basis of the modified prediction block. According to the present invention, a current block can be predicted in consideration of the correlation between a prediction sample and a reconstruction sample of a neighboring block, and a complicated image can be more efficiently reconstructed through the prediction.
1. A video decoding method performed by a video decoder, the method comprising:
deriving a prediction mode of a current block;
generating a prediction block of the current block based on the prediction mode of the current block;
deriving a spatial neighboring block of the current block;
deriving correlation information based on a reconstructed block of the spatial neighboring block and a prediction block of the spatial neighboring block;
deriving a modified prediction block of the current block based on the prediction block of the current block and the correlation information; and
deriving a reconstructed block of the current block based on the modified prediction block,
wherein the correlation information comprises a weight per block unit and an offset per block unit,
wherein the weight per block unit and the offset per block unit are derived through Least Square Method (LSM) with respect to values of reconstructed samples included in the reconstructed block of the spatial neighboring block and values of prediction samples included in the prediction block of the spatial neighboring block,
wherein the modified prediction block is derived based on an equation as below:
P ′( x,y )=α P ( x,y )+β
where P′(x,y) denotes a modified prediction sample value at coordinates (x,y) included in the modified prediction block, P(x,y) denotes a prediction sample value at coordinates (x,y) included in the prediction block of the current block, α denotes the weight per block unit based on the LSM, and β denotes the offset per block unit based on the LSM,
wherein, when the prediction block of the current block comprises I number of samples, the α denoting the weight per block unit is derived based on an equation as below:
α
=
I
·
∑
i
=
1
I
Rec
A
(
i
)
·
Pred
A
(
i
)
-
∑
i
=
1
I
Rec
A
(
i
)
·
∑
i
=
1
I
Pred
A
(
i
)
I
·
∑
i
=
1
I
Pred
A
(
i
)
·
Pred
A
(
i
)
-
(
∑
i
=
1
I
Pred
A
(
i
)
)
2
where RecA(i) denotes an ith reconstructed sample value of the reconstructed block of the spatial neighboring block, and PredA(i) denotes an ith prediction sample value of the prediction block of the spatial neighboring block, and
wherein the β denoting the offset per block unit is derived based on an equation as below:
β
=
∑
i
=
1
I
Rec
A
(
i
)
-
α
·
∑
i
=
1
I
Pred
A
(
i
)
I
.
2. The video decoding method of claim 1 , wherein the spatial neighboring block is a spatial neighboring block having a smallest difference value between samples, according to a phase, relative to the current block among spatial neighboring blocks of the current block.
3. The method of claim 1 , wherein, when the prediction mode of the current block is an intra prediction mode, the modified prediction block of the current block is derived.
4. The video decoding method of claim 1 , wherein:
the correlation information comprises a weight per sample unit and an offset per sample unit; and
the weight per sample unit and the offset per sample unit are derived through Least Square Method (LSM) with respect to values of reconstructed samples in a first region of the reconstructed block of the spatial neighboring block, the first region located to correspond to a prediction sample of the prediction block of the current block, and values of prediction samples in a second region of the prediction block of the spatial neighboring block, the second region corresponding to the first region.
5. The video decoding method of claim 4 , wherein the first region of the reconstructed block of the spatial neighboring block and the second region of the prediction block of the spatial neighboring block have an identical size.
6. The video decoding method of claim 1 , wherein:
the correlation information comprises information on a difference value that corresponds to each prediction sample included in the prediction block of the current block; and
the difference value indicates a difference value between a reconstructed sample value of the reconstructed block of the spatial neighboring block, which corresponds to a prediction sample of the prediction block of the current block, and a prediction sample value of the prediction block of the spatial neighboring block, which corresponds to a prediction sample to the prediction block of the current block.
7. The video decoding method of claim 6 , wherein:
the difference value corresponding to each prediction sample included in the prediction block of the current block is derived based on an equation as below:
Δ
1
=
Rec
A
(
0
,
0
)
-
Pred
A
(
0
,
0
)
Δ
2
=
Rec
A
(
0
,
1
)
-
Pred
A
(
0
,
1
)
⋮
Δ
N
=
Rec
A
(
N
,
N
)
-
Pred
A
(
N
,
N
)
where RecA(x,y) denotes a reconstructed sample value at coordinates (x,y) of the reconstructed block of the spatial neighboring block, PredA(x,y) denotes a prediction sample value at coordinates (x,y) of the prediction block of the spatial neighboring block, ΔN denotes a Nth difference value, and a final prediction block comprises N number of samples.
8. A video encoding method performed by a video encoder, the method comprising:
deriving a prediction mode of the current block;
generating a prediction block of the current block based on the prediction mode of the current block;
determining a spatial neighboring block of the current block;
generating correlation information based on a reconstructed block of the spatial neighboring block and a prediction block of the spatial neighboring block;
deriving a modified prediction block of the current block based on the prediction block of the current block and the correlation information;
generating residual information based on an original block of the current block and the modified prediction block; and
encoding information on the prediction mode of the current block and the residual information, and outputting the encoded information,
wherein the correlation information comprises a weight per block unit and an offset per block unit,
wherein the weight per block unit and the offset per block unit are derived through Least Square Method (LSM) with respect to values of reconstructed samples included in the reconstructed block of the spatial neighboring block and values of prediction samples included in the prediction block of the spatial neighboring block,
wherein the modified prediction block is derived based on an equation as below:
P ′( x,y )=α P ( x,y )+β
where P′(x,y) denotes a modified prediction sample value at coordinates (x,y) included in the modified prediction block, P(x,y) denotes a prediction sample value at coordinates (x,y) included in the prediction block of the current block, α denotes the weight per block unit based on the LSM, and β denotes the offset per block unit based on the LSM,
wherein, when the prediction block of the current block comprises I number of samples, the α denoting the weight per block unit is derived based on an equation as below:
α
=
I
·
∑
i
=
1
I
Rec
A
(
i
)
·
Pred
A
(
i
)
-
∑
i
=
1
I
Rec
A
(
i
)
·
∑
i
=
1
I
Pred
A
(
i
)
I
·
∑
i
=
1
I
Pred
A
(
i
)
·
Pred
A
(
i
)
-
(
∑
i
=
1
I
Pred
A
(
i
)
)
2
where RecA(i) denotes an ith reconstructed sample value of the reconstructed block of the spatial neighboring block, and PredA(i) denotes an ith prediction sample value of the prediction block of the spatial neighboring block, and
wherein the 3 denoting the offset per block unit is derived based on an equation as below:
β
=
∑
i
=
1
I
Rec
A
(
i
)
-
α
·
∑
i
=
1
I
Pred
A
(
i
)
I
.
9. The video encoding method of claim 8 , wherein the spatial neighboring block is determined to be a spatial neighboring block that has a smallest difference value between samples, according to a phase, relative to the current block among spatial neighboring blocks of the current block.