IP Library › Granted Patent US 12,149,703
Granted Patent B2
US 12,149,703 · App. 18/206,918 · Granted Nov 19, 2024

Method and system for decoder-side intra mode derivation for block-based video coding

Inventors: Xiaoyu Xiu (San Diego, CA); Yuwen He (San Diego, CA); Yan Ye (San Diego, CA)
Assignee: InterDigital Madison Patent Holdings, SAS
H04N19/154H04N19/11H04N19/147H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,149,703
App. No.
18/206,918
Granted
Nov 19, 2024
Kind
B2
Abstract

Systems and methods are disclosed for video encoding and video decoding using decoder-side intra mode derivation (DIMD). Techniques are provided to code and to decode a video block of a video frame into a bitstream, including determining costs of using respective candidate intra prediction modes to predict samples in a template region adjacent to the video block, deriving an intra prediction mode based on candidates of the candidate intra prediction modes and their respective costs, and predicting samples in the current video block using the derived intra prediction mode. The provided techniques further include determining the costs in multiple stages. In an initial stage, the costs of using respective intra prediction modes from an initial set of candidate modes are determined. Then, in at least one subsequent stage, the costs of using respective intra prediction modes from a subsequent set of candidate modes are determined, where the subsequent set of candidate modes is selected based on the candidate mode in the previous stage having the lowest cost.

Claims (59)

1. A method for encoding video data, comprising:

coding a current video block of a video frame of the video data into a bitstream, the coding comprises:

determining costs of using respective candidate intra prediction modes to predict samples in a template region adjacent to the current video block, wherein determining the costs comprises:

in an initial stage, determining costs of using respective intra prediction modes from an initial set of candidate modes, and

in at least one subsequent stage, determining costs of using respective intra prediction modes from a subsequent set of candidate modes, the subsequent set of candidate modes is selected based on the candidate mode in the previous stage having the lowest cost,

deriving an intra prediction mode based on candidates of the candidate intra prediction modes and their respective costs, and

predicting samples in the current video block using the derived intra prediction mode.

2. The method of claim 1 , wherein each of the determined costs is a measure of distortion between the template region and a prediction of the template region using a respective candidate intra prediction mode.

3. The method of claim 2 , wherein the template region comprises reconstructed samples and the prediction of the template region is based on a set of reconstructed reference samples.

4. The method of claim 1 , further comprising:

coding, into the bitstream, a flag indicating that decoder-side intra mode derivation is used for the current video block.

5. The method of claim 1 , wherein the initial set of candidate modes includes a planar mode and a DC mode, and wherein the at least one subsequent stage is performed only in response to a determination that neither the planar nor the DC mode is the candidate mode having the lowest cost.

6. The method of claim 1 , wherein the initial set of candidate modes is adaptively determined based on at least one intra prediction mode of a video block in the neighborhood of the current video block.

7. The method of claim 1 , wherein the candidate intra prediction modes include at least one intra prediction mode of a video block in the spatial or temporal neighborhood of the current video block.

8. The method of claim 1 , further comprising:

in a current stage, of the initial stage or of the at least one subsequent stage, computing a cost variation measure of the determined costs; and

if the cost variation measure is below a predetermined threshold, setting the current stage as the last stage.

9. The method of claim 1 , wherein at least some video blocks of the video data are predicted using a predetermined set of explicitly-signaled intra modes, and wherein the candidate intra prediction modes have a finer granularity than the predetermined set of explicitly signaled intra modes.

10. The method of claim 1 , wherein:

in the initial stage, the intra prediction modes from the initial set are separated by an initial interval; and

in the at least one subsequent stage, the intra prediction modes from the subsequent set are separated by a subsequent interval smaller than the interval used in the previous stage.

11. A method for decoding video data from a bitstream, comprising:

decoding, from the bitstream, a current video block of a video frame of the video data, the decoding comprises:

determining costs of using respective candidate intra prediction modes to predict samples in a template region adjacent to the current video block, wherein determining the costs comprises:

in an initial stage, determining costs of using respective intra prediction modes from an initial set of candidate modes, and

in at least one subsequent stage, determining costs of using respective intra prediction modes from a subsequent set of candidate modes, the subsequent set of candidate modes is selected based on the candidate mode in the previous stage having the lowest cost,

deriving an intra prediction mode based on candidates of the candidate intra prediction modes and their respective costs, and

predicting samples in the current video block using the derived intra prediction mode.

12. The method of claim 11 , further comprising:

decoding, from the bitstream, a flag indicating that decoder-side intra mode derivation is to be used for the current video block.

13. The method of claim 11 , wherein the initial set of candidate modes is adaptively determined based on at least one intra prediction mode of a video block in the neighborhood of the current video block.

14. The method of claim 11 , wherein the candidate intra prediction modes include at least one intra prediction mode of a video block in the spatial or temporal neighborhood of the current video block.

15. A system for encoding video data, comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the system to code a current video block of a video frame of the video data into a bitstream, the coding comprises:

determining costs of using respective candidate intra prediction modes to predict samples in a template region adjacent to the current video block, wherein determining the costs comprises:

in an initial stage, determining costs of using respective intra prediction modes from an initial set of candidate modes, and

in at least one subsequent stage, determining costs of using respective intra prediction modes from a subsequent set of candidate modes, the subsequent set of candidate modes is selected based on the candidate mode in the previous stage having the lowest cost,

deriving an intra prediction mode based on candidates of the candidate intra prediction modes and their respective costs, and

predicting samples in the current video block using the derived intra prediction mode.

16. The system of claim 15 , wherein the coding further comprising:

coding, into the bitstream, a flag indicating that decoder-side intra mode derivation is used for the current video block.

17. The system of claim 15 , wherein the derived intra prediction mode is included in a list of most probable modes and wherein an index is coded into the bitstream identifying the derived intra prediction mode from the list of most probable modes.

18. The system of claim 15 , wherein prediction residuals for the samples in the current video block are coded in a bitstream using a transform coefficient scanning order, and wherein the selection of the transform coefficient scanning order is independent of the derived intra prediction mode.

19. The system of claim 18 , wherein the transform coefficient scanning order is a predetermined scanning order.

20. The system of claim 18 , wherein the transform coefficient scanning order is based on intra prediction modes of spatial neighbors of the current video block.

21. A system for decoding video data from a bitstream, comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the system to decode, from the bitstream, a current video block of a video frame of the video data, the decoding comprises:

determining costs of using respective candidate intra prediction modes to predict samples in a template region adjacent to the current video block, wherein determining the costs comprises:

in an initial stage, determining costs of using respective intra prediction modes from an initial set of candidate modes, and

in at least one subsequent stage, determining costs of using respective intra prediction modes from a subsequent set of candidate modes, the subsequent set of candidate modes is selected based on the candidate mode in the previous stage having the lowest cost,

deriving an intra prediction mode based on candidates of the candidate intra prediction modes and their respective costs, and

predicting samples in the current video block using the derived intra prediction mode.

22. The system of claim 21 , further comprising:

decoding, from the bitstream, a flag indicating that decoder-side intra mode derivation is to be used for the current video block.

23. The system of claim 21 , wherein the initial set of candidate modes is adaptively determined based on at least one intra prediction mode of a video block in the neighborhood of the current video block.

24. The system of claim 21 , wherein the candidate intra prediction modes include at least one intra prediction mode of a video block in the spatial or temporal neighborhood of the current video block.

25. The system of claim 21 , wherein the derived intra prediction mode is included in a list of most probable modes and wherein an index is decoded from the bitstream identifying the derived intra prediction mode from the list of most probable modes.

Continuity (5)
Continuation 16096236
Provisional Application 62367414 · Jul 27, 2016
Provisional Application 62335512 · May 12, 2016
Provisional Application 62332871 · May 6, 2016
Related Publication 20230319289A1 · Oct 5, 2023