Process and arrangement for encoding video pictures
Today's video codecs require the intelligent choice between many coding options. This choice can efficiently be done using Lagrangian coder control. But Lagrangian coder control only provides results given a particular Lagrange parameter, which correspond to some unknown transmission rate. On the other hand, rate control algorithms provide coding results at a given bitrate but without the optimization performance of Lagrangian coder control. The combination of rate control and Lagrangian optimization for hybrid video coding is investigated. A new approach is suggested to incorporate these two known methods into the video coder control using macroblock mode decision and quantizer adaptation. The rate-distortion performance of the proposed approach is validated and analyzed via experimental results. It is shown that for most bit-rates the combined rate control and Lagrangian optimization producing a constant number of bits per picture achieves similar rate distortion performance as the constant slope case only using Lagrangian optimization.
1. Process for encoding video pictures in a video coder,
wherein
a pre-analysis of pictures is carried out, wherein for at least a part of macroblocks at least one control-parameter which assists the encoding process is determined based on at least one estimated parameter,
in a second step the picture is encoded on macroblock level with encoding-parameters calculated based on the control-parameters determined in the pre-analysis step, wherein the macroblock encoding process comprises the following steps given a target quantization parameter QP i *:
calculating Lagrangian multipliers used for motion estimation and mode decision of a macroblock i according to:
(λ motion,i ) 2 =λ mode,i =0.85· QP i * 2 for H.263,MPEG-4, or
(λ motion,i ) 2 =λ mode,i =0.85·2^(( QP* i −12)/3)for H.264/AVC,
for all motion-compensated macroblock/block modes determining associated motion vectors m i and a reference index r i by minimizing a Lagrangian functional
[
m
i
,
r
i
]
=
arg
min
m
∈
M
,
r
∈
R
{
D
DFD
(
i
,
m
,
r
)
+
λ
motion
,
i
·
R
MV
(
i
,
m
,
r
)
}
,
with a distortion term being given as
D
DFD
(
i
,
m
,
r
)
=
∑
(
x
,
y
)
∈
B
i
s
(
x
,
y
,
t
)
-
s
′
(
x
-
m
x
,
y
-
m
y
,
t
r
)
,
where s( . . . , t) and s′( . . . , t r ) represent an array of luminance samples of an original picture and a decoded reference picture given by a reference index r, respectively,
R denotes a set of reference pictures stored in the decoded picture buffer,
M specifies a motion vectors search range inside a reference picture,
t r is a sampling time of a reference picture referred by the reference index r,
B is the area of a corresponding block or macroblock, and
R MV (i,m,r) specifies a number of bits needed to transmit all components of a motion vector m=[m x ,m y ] T as well as the reference index r;
determining macroblock/block encoding modes p i of a macroblock i by minimizing a Lagrangian cost function
p
i
=
arg
min
p
∈
S
mod
e
{
D
REC
(
i
,
p
❘
QP
i
*
)
+
λ
mod
e
·
R
all
(
i
,
p
❘
QP
i
*
)
}
where D REC (i,p|QP i *) is defined as
D
REC
(
i
,
p
❘
QP
i
*
)
=
∑
(
x
,
y
)
∈
B
(
s
(
x
,
y
)
-
s
′
(
x
,
y
❘
p
,
QP
i
*
)
)
2
s( . . . ) and s′( . . . ) represent an array of original macroblock samples and their reconstruction, respectively,
B specifies a set of corresponding macroblock/block samples,
R all (i,p|QP i *) is a number of bits associated with choosing a mode p and quantization parameter QP i *, including bits for the macroblock/block modes, motion vectors and reference indices as well as quantized transform coefficients of all luminance and chrominance blocks, and
S mode is a given set of possible macroblock/block modes.
2. Process according to claim 1 ,
wherein
an energy measure of a residual signal of a macroblock representing the difference between an original macroblocks samples and their prediction is used as control parameter, which is determined based on at least one estimated parameter in the pre-analysis step.
3. Process according to claim 2 ,
wherein
the energy measure of the residual signal is calculated as the average of variances of the residual signals of luminance and chrominance blocks inside a macroblocks i that are used for transform coding according to:
σ
i
2
=
1
N
B
·
N
P
∑
j
=
1
N
B
∑
k
=
1
N
P
(
d
i
,
j
(
k
)
-
d
i
,
j
_
)
2
where
N B and N p are the number of blocks (luminance and chrominance) used for transform coding inside a macroblock and the number of samples inside such a block, respectively,
d i,j is the residual signal of the block j inside the macroblock i, and
d i,j represents the average of the d i,j .
4. Process according to claim 1 ,
wherein
for predictive coded pictures, the prediction signal of macroblocks used for determining the control parameters is estimated by motion compensated prediction using one or more displacement vectors and reference indices that are estimated in the pre-analysis step.
5. Process according to claim 1 ,
wherein the pre-analysis step includes an estimation of displacement vectors m and reference index r by minimizing a Lagrangian cost function
[
m
^
,
r
]
=
arg
min
m
,
r
{
D
DFD
(
m
,
r
)
+
λ
motion
·
R
MV
(
m
,
r
)
}
,
where
D
DFD
(
m
,
r
)
=
∑
(
x
,
y
)
∈
B
s
(
x
,
y
,
t
)
-
s
′
(
x
-
m
x
,
y
-
m
y
,
t
r
)
determines a distortion term,
s( . . . , t) and s′( . . . , t r ) represent an array of luminance samples of an original picture and a decoded reference picture given by the reference index r, respectively,
R MV (m,r) specifies a number of bits needed to transmit all components of a displacement vector [m x ,m y ] T and the reference index r,
B is the area of a macroblock, macroblock partition, or sub-macroblock partition for which the displacement vector and the reference index are estimated, and
λ motion ≧0 is the Lagrangian multiplier.
6. Process according to claim 5 ,
wherein
the Lagrangian multiplier λ motion used for displacement vector estimation in the pre-analysis step is set in accordance with
λ motion =√{square root over (0.85· QP 2 )}for H.263,MPEG-4 or
λ motion =√{square root over (0.85·2^(( QP −12)/3))}for H.264/AVC
where QP represents an average quantization parameter of a last encoded picture of a same picture type.
7. Process according to claim 1 ,
wherein
a displacement vector estimation in the pre-analysis step is done for the entire macroblock covering an area of 16×16 luminance samples, and the reference index r is not estimated but determined in a way that it refers to the temporally closest reference picture that is stored in the decoded picture buffer.
8. Process according to claim 1 ,
wherein
during the encoding process for each macroblock i a target quantization parameter QP i * is determined in dependence of the control parameters estimated in the pre-analysis step.
9. Arrangement with at least one chip and/or processor that is (are) installed in such a manner, that a process for encoding video pictures can be executed in a manner so that a pre-analysis of pictures is carried out, wherein for at least a part of macroblocks at least one control-parameter which assists the encoding process is determined based on at least one estimated parameter,
in a second step the picture is encoded with encoding-parameters calculated based on the control-parameters determined in the pre-analysis step wherein the macroblock encoding process comprises the following steps given a target quantization parameter QP i *:
calculating Lagrangian multipliers used for motion estimation and mode decision of a macroblock i according to:
(λ motion,i ) 2 =λ mode,i =0.85· QP i * 2 for H.263,MPEG-4, or
(λ motion,i ) 2 =λ mode,i =0.85·2^(( QP* i −12)/3)for H.264/AVC,
for all motion-compensated macroblock/block modes determining associated motion vectors m i and a reference index r i by minimizing a Lagrangian functional
[
m
i
,
r
i
]
=
arg
min
m
∈
M
,
r
∈
R
{
D
DFD
(
i
,
m
,
r
)
+
λ
motion
,
i
·
R
MV
(
i
,
m
,
r
)
}
,
with a distortion term being given as
D
DFD
(
i
,
m
,
r
)
=
∑
(
x
,
y
)
∈
B
i
s
(
x
,
y
,
t
)
-
s
′
(
x
-
m
x
,
y
-
m
y
,
t
r
)
,
where s( . . . , t) and s′( . . . , t r ) an array of luminance samples of an original picture and a decoded reference picture given by a reference index r, respectively,
R denotes a set of reference pictures stored in the decoded picture buffer,
M specifies a motion vectors search range inside a reference picture,
t r is a sampling time of a reference picture referred by the reference index r,
B is the area of a corresponding block or macroblock, and
R MV (i,m,r) specifies a number of bits needed to transmit all components of a motion vector m=[m x ,m y ] T as well as the reference index r;
determining macroblock/block encoding modes p i of a macroblock i by minimizing a Lagrangian cost function
p
i
=
arg
min
p
∈
S
mod
e
{
D
REC
(
i
,
p
|
QP
i
*
)
+
λ
mod
e
·
R
all
(
i
,
p
|
QP
i
*
)
}
where D REC (i,p|QP i *) is defined as
D
REC
(
i
,
p
|
QP
i
*
)
=
∑
(
x
,
y
)
∈
B
(
s
(
x
,
y
)
-
s
′
(
x
,
y
|
p
,
QP
i
*
)
)
2
s( . . . ) and s′( . . . ) represent an array of original macroblock samples and their reconstruction, respectively,
B specifies a set of corresponding macroblock/block samples,
R all (i,p|QP i *) is a number of bits associated with choosing a mode p and quantization parameter QP i *, including bits for the macroblock/block modes, motion vectors and reference indices as well as quantized transform coefficients of all luminance and chrominance blocks, and
S mode is a given set of possible macroblock/block modes.
10. Computer program embodied on a non-transitory computer readable storage medium within a computer that enables the computer to run a process for encoding video pictures, where a pre-analysis of pictures is carried out, wherein for at least a part of macroblocks at least one control-parameter which assists the encoding process is determined based on at least one estimated parameter,
in a second step the picture is encoded with encoding-parameters calculated based on the control-parameters determined in the pre-analysis step wherein the macroblock encoding process comprises the following steps given a target quantization parameter QP i *:
calculating Lagrangian multipliers used for motion estimation and mode decision of a macroblock i according to:
(λ motion,i ) 2 =λ mode,i =0.85· QP i * 2 for H.263,MPEG-4, or
(λ motion,i ) 2 =λ mode,i =0.85·2^(( QP* i −12)/3)for H.264/AVC,
for all motion-compensated macroblock/block modes determining associated motion vectors m i and a reference index r i by minimizing a Lagrangian functional
[
m
i
,
r
i
]
=
arg
min
m
∈
M
,
r
∈
R
{
D
DFD
(
i
,
m
,
r
)
+
λ
motion
,
i
·
R
MV
(
i
,
m
,
r
)
}
,
with a distortion term being given as
D
DFD
(
i
,
m
,
r
)
=
∑
(
x
,
y
)
∈
B
i
s
(
x
,
y
,
t
)
-
s
′
(
x
-
m
x
,
y
-
m
y
,
t
r
)
,
where s( . . . , t) and s′( . . . , t r )represent an array of luminance samples of an original picture and a decoded reference picture given by a reference index r, respectively,
R denotes a set of reference pictures stored in the decoded picture buffer,
M specifies a motion vectors search range inside a reference picture,
t r is a sampling time of a reference picture referred by the reference index r,
B is the area of a corresponding block or macroblock, and
R MV (i,m,r) specifies a number of bits needed to transmit all components of a motion vector m=[m x ,m y ] T as well as the reference index r;
determining macroblock/block encoding modes p i of a macroblock i by minimizing a Lagrangian cost function
p
i
=
arg
min
p
∈
S
mod
e
{
D
REC
(
i
,
p
❘
QP
i
*
)
+
λ
mod
e
·
R
all
(
i
,
p
❘
QP
i
*
)
}
where D REC (i,p|QP i *) is defined as
D
REC
(
i
,
p
❘
QP
i
*
)
=
∑
(
x
,
y
)
∈
B
(
s
(
x
,
y
)
-
s
′
(
x
,
y
❘
p
,
QP
i
*
)
)
2
s( . . . ) and s′( . . . ) represent an array of original macroblock samples and their reconstruction, respectively,
B specifies a set of corresponding macroblock/block samples,
R all (i,p|QP i *) is a number of bits associated with choosing a mode p and quantization parameter QP i *, including bits for the macroblock/block modes, motion vectors and reference indices as well as quantized transform coefficients of all luminance and chrominance blocks, and
S mode is a given set of possible macroblock/block modes.
11. A non-transitory computer-readable storage medium, on which an executable program is stored, that enables a computer to execute a process for encoding video pictures according to claim 1 .
12. Process in that a computer program as described in claim 10 is downloaded from a network for data transfer to a data processing unit, that is connected to said network.