IP Library Granted Patent US 11,350,100
Granted Patent B2
US 11,350,100 · App. 17/343,086 · Granted May 31, 2022

Transform coding based on matrix-based intra prediction

Inventors: Zhipin Deng (Beijing, CN); Kai Zhang (San Diego, CA); Li Zhang (San Diego, CA); Hongbin Liu (Beijing, CN); Jizheng Xu (San Diego, CA)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/132H04N19/105H04N19/117H04N19/12H04N19/159H04N19/176H04N19/186H04N19/60H04N19/70H04N19/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,350,100
App. No.
17/343,086
Granted
May 31, 2022
Kind
B2
Abstract

Devices, systems and methods for digital video coding, which includes matrix-based intra prediction methods for video coding, are described. In a representative aspect, a method for video processing includes performing a conversion between a current video block of a video and a bitstream representation of the current video block according to a rule, where the rule specifies a relationship between applicability of a matrix based intra prediction (MIP) mode or a transform mode during the conversion, where the MIP mode includes determining a prediction block of the current video block by performing, on previously coded samples of the video, a boundary downsampling operation, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation, and where the transform mode specifies use of a transform operation for the determining the prediction block for the current video block.

Claims (40)

1. A video processing method, comprising:

determining, for a conversion between a video block of a video and the bitstream of the video, that a first intra mode is applied on the video block of the video,

performing a boundary downsampling operation on reference samples of the video block based on a size of the video block, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation to generate prediction samples for the video block; and

performing the conversion based on the prediction samples and residual samples of the video block;

wherein a secondary transform tool is applied on the video block, and a secondary transform matrix is selected to generate the residual samples for the video block, and

wherein the secondary transform tool includes applying, during encoding, a forward secondary transform to an output of a forward primary transform before applying a quantization process, or applying, during decoding, an inverse secondary transform to an output of a dequantization process before applying an inverse primary transform.

2. The method of claim 1 , wherein the secondary transform tool includes a non-separable secondary transform.

3. The method of claim 1 , wherein the secondary transform tool includes a Reduced Secondary Transform (RST) or a rotation transform.

4. The method of claim 1 , wherein based on the first intra mode being applied on the video block, a second intra mode is converted from the first intra mode, and wherein a secondary transform classification is derived based on the second intra mode.

5. The method of claim 4 , wherein the selected secondary transform matrix is determined based on the secondary transform classification.

6. The method of claim 1 , wherein a down-sampling factor derived in the boundary down-sampling operation is larger than or equal to 1, and wherein a one-dimensional vector array is further derived based on concatenating down-sampled reference samples derived from the boundary down-sampling operation and the one-dimensional vector array is used as input of the matrix vector multiplication operation.

7. The method of claim 4 , wherein in the second intra mode, distance-based weighted calculations are applied on reference values in vertical direction and horizontal direction to derive prediction values.

8. The method of claim 4 , wherein the second intra mode includes a planar mode.

9. The method of claim 1 , wherein whether to apply the first intra mode is specified by a first syntax element included in a sequence parameter set and a second syntax element included in a coding unit level set.

10. The method of claim 9 , wherein in response to a width-height ratio of the video block being greater than 2, a context with an index of 3 is used for a first bin of the second syntax element.

11. The method of claim 9 , wherein in response to a width-height ratio of the video block being smaller than or equal to 2, a single context selected from contexts with indices of 0, 1 or 2 is used for a first bin of the second syntax element.

12. The method of claim 1 , wherein the first intra mode includes multiple types, and a type index for the video block is derived excluding referring to type indices of previous video blocks.

13. The method of claim 12 , wherein the type index for the video block is explicitly included in the bitstream.

14. The method of claim 1 , wherein whether to apply the secondary transform tool is based on a height (H) or a width (W) of the video block.

15. The method of claim 1 , wherein the conversion includes encoding the video block into the bitstream.

16. The method of claim 1 , wherein the conversion includes decoding the video block from the bitstream.

17. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

determine, for a conversion between a video block of a video and the bitstream of the video, that a first intra mode is applied on the video block of the video,

perform a boundary downsampling operation on reference samples of the video block based on a size of the video block, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation to generate prediction samples for the video block; and

perform the conversion based on the prediction samples and residual samples of the video block;

wherein a secondary transform tool is applied on the video block, and a secondary transform matrix is selected to generate the residual samples for the video block, and

wherein the secondary transform tool includes applying, during encoding, a forward secondary transform to an output of a forward primary transform before applying a quantization process, or applying, during decoding, an inverse secondary transform to an output of a dequantization process before applying an inverse primary transform.

18. The apparatus of claim 17 , wherein based on the first intra mode being applied on the video block, a second intra mode is converted from the first intra mode, and wherein a secondary transform classification is derived based on the second intra mode.

19. A non-transitory computer-readable storage medium storing instructions that cause a processor to:

determine, for a conversion between a video block of a video and the bitstream of the video, that a first intra mode is applied on the video block of the video,

perform a boundary downsampling operation on reference samples of the video block based on a size of the video block, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation to generate prediction samples for the video block; and

perform the conversion based on the prediction samples and residual samples of the video block;

wherein a secondary transform tool is applied on the video block, and a secondary transform matrix is selected to generate the residual samples for the video block, and

wherein the secondary transform tool includes applying, during encoding, a forward secondary transform to an output of a forward primary transform before applying a quantization process, or applying, during decoding, an inverse secondary transform to an output of a dequantization process before applying an inverse primary transform.

20. A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by a video processing apparatus, wherein the method comprises:

determining that a first intra mode is applied on a video block of the video,

performing a boundary downsampling operation on reference samples of the video block based on a size of the video block, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation to generate prediction samples for the video block; and

generating the bitstream based on the prediction samples and residual samples of the video block;

wherein a secondary transform tool is applied on the video block, and a secondary transform matrix is selected to generate the residual samples for the video block, and

wherein the secondary transform tool includes applying, during encoding, a forward secondary transform to an output of a forward primary transform before applying a quantization process, or applying, during decoding, an inverse secondary transform to an output of a dequantization process before applying an inverse primary transform.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2021
From: ZHANG, KAI; ZHANG, LI; XU, JIZHENG
To: BYTEDANCE INC.
Reel/Frame 056489/0926 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2021
From: DENG, ZHIPIN; LIU, HONGBIN
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 056489/0988 →
Priority Claims (1)
WO PCT/CN2019/082424 · Apr 12, 2019 · international
Continuity (2)
Continuation PCTCN2020084472 · Apr 13, 2020
Related Publication 20210297672A1 · Sep 23, 2021