Method and system for decoder-side intra mode derivation for block-based video coding
View Patent ↗Systems and methods are disclosed for encoding and for decoding of video data, including predicting a video block of the video data using decoder-side intra mode derivation (DIMD). The disclosed predicting techniques include selecting candidate intra prediction modes based on intra prediction modes used by neighboring video blocks. Techniques disclosed further include determining respective costs of using the selected candidate intra prediction modes to predict samples in a template region adjacent to the video block, deriving an intra prediction mode from the selected candidate intra prediction modes having the lowest cost, and predicting samples in the video block using the derived intra prediction mode.
1 . A method of encoding video data into a bitstream, comprising:
predicting a video block from a frame of the video data using decoder-side intra mode derivation (DIMD), the predicting comprising:
selecting candidate intra prediction modes from among intra prediction modes that are used by neighboring video blocks,
determining respective costs of using the selected candidate intra prediction modes to predict samples in a template region adjacent to the video block,
deriving an intra prediction mode from among the selected candidate intra prediction modes having the lowest cost, and
predicting samples in the video block using the derived intra prediction mode.
2 . The method according to claim 1 , wherein the neighboring video blocks include video blocks in a spatial neighborhood of the video block, in a temporal neighborhood of the video block, or in both.
3 . The method according to claim 1 , wherein the selecting of the candidate intra prediction modes comprises:
selecting candidates, of the intra prediction modes used by the neighboring video blocks, that are most frequently used by the neighboring video blocks.
4 . The method according to claim 1 , further comprising:
adding, before the determining, to the selected candidate intra prediction modes at least one of a DC mode and a planar mode, if not already included.
5 . The method according to claim 1 , further comprising:
refining the derived intra prediction mode by:
obtaining candidate intra prediction modes within a range centered at the derived intra prediction mode,
determining respective costs of using the obtained candidate intra prediction modes to predict samples in the template region adjacent to the video block,
deriving a refined intra prediction mode from the obtained candidate intra prediction modes having the lowest cost,
wherein the predicting of the samples is performed by using the refined intra prediction mode.
6 . The method according to claims 1 , wherein each of the determined costs is a measure of distortion between the template region and a prediction of the template region using a respective candidate intra prediction mode.
7 . The method according to claim 1 , further comprising:
determining at least one of a size of the template region and a layout of the template region.
8 . An apparatus for encoding video data into a bitstream, comprising:
at least one processor; and
memory storing instructions that, when executed by the at least one processor, cause the apparatus to predict a video block from a frame of the video data using DIMD, the predicting comprising:
selecting candidate intra prediction modes from among intra prediction modes that are used by neighboring video blocks,
determining respective costs of using the selected candidate intra prediction modes to predict samples in a template region adjacent to the video block,
deriving an intra prediction mode from among the selected candidate intra prediction modes having the lowest cost, and
predicting samples in the video block using the derived intra prediction mode.
9 . The apparatus according to claim 8 , wherein the instructions further cause the apparatus to:
code, into the bitstream, a flag indicating that the DIMD is used for the video block.
10 . The apparatus according to claim 8 , wherein the neighboring video blocks include video blocks in a spatial neighborhood of the video block, in a temporal neighborhood of the video block, or in both.
11 . The apparatus according to claim 8 , wherein the selecting of the candidate intra prediction modes comprises:
selecting candidates, of the intra prediction modes used by the neighboring video blocks, that are most frequently used by the neighboring video blocks; and
adding to the selected candidate intra prediction modes at least one of a DC mode and a planar mode, if not already included.
12 . A method of decoding video data from a bitstream, comprising:
predicting a video block from a frame of the video data using DIMD, the predicting comprising:
selecting candidate intra prediction modes from among intra prediction modes that are used by neighboring video blocks,
determining respective costs of using the selected candidate intra prediction modes to predict samples in a template region adjacent to the video block,
deriving an intra prediction mode from among the selected candidate intra prediction modes having the lowest cost, and
predicting samples in the video block using the derived intra prediction mode.
13 . The method according to claim 12 , further comprising:
decoding, from the bitstream, a flag indicating that the DIMD is used for the video block.
14 . The method according to claim 12 , wherein the neighboring video blocks include blocks in a spatial neighborhood of the current block, in a temporal neighborhood of the current block, or in both.
15 . The method according to claim 12 , wherein the selecting of the candidate intra prediction modes comprises:
selecting candidates, of the intra prediction modes used by the neighboring video blocks, that are most frequently used by the neighboring video blocks.
16 . The method according to claim 12 , further comprising:
adding, before the determining, to the selected candidate intra prediction modes at least one of a DC mode and a planar mode, if not already included.
17 . An apparatus for decoding video data from a bitstream, comprising:
at least one processor; and
memory storing instructions that, when executed by the at least one processor, cause the apparatus to predict a video block from a frame of the video data using DIMD, the predicting comprising:
selecting candidate intra prediction modes from among intra prediction modes that are used by neighboring video blocks,
determining respective costs of using the selected candidate intra prediction modes to predict samples in a template region adjacent to the video block,
deriving an intra prediction mode from among the selected candidate intra prediction modes having the lowest cost, and
predicting samples in the video block using the derived intra prediction mode.
18 . The apparatus according to claim 17 , wherein the instructions further cause the apparatus to:
decode, from the bitstream, a flag indicating that the DIMD is used for the video block.
19 . The apparatus according to claim 17 , wherein the neighboring video blocks include blocks in a spatial neighborhood of the current block.
20 . The apparatus according to claim 17 , wherein the neighboring video blocks include blocks in a temporal neighborhood of the current block.