IP Library › Granted Patent US 12,003,766
Granted Patent B2
US 12,003,766 · App. 17/942,552 · Granted Jun 4, 2024

Context coding for matrix-based intra prediction

Inventors: Zhipin Deng (Beijing, CN); Kai Zhang (San Diego, CA); Li Zhang (San Diego, CA); Hongbin Liu (Beijing, CN); Jizheng Xu (San Diego, CA)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/593H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,003,766
App. No.
17/942,552
Granted
Jun 4, 2024
Kind
B2
Abstract

Devices, systems and methods for digital video coding, which includes matrix-based intra prediction methods for video coding, are described. In a representative aspect, a method for video processing includes encoding a current video block of a video using a matrix intra prediction (MIP) mode in which a prediction block of the current video block is determined by performing, on previously coded samples of the video, a boundary downsampling operation, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation; and adding, to a coded representation of the current video block, a syntax element indicative of applicability of the MIP mode to the current video block using arithmetic coding in which a context for the syntax element is derived based on a rule.

Claims (51)

1. A method of processing video data, comprising:

determining, for a conversion between a current video block of a video and a bitstream of the video, whether a matrix intra prediction (MIP) mode is applied on the current video block based on a syntax element, wherein in the MIP mode, prediction samples of the current video block are determined by performing a matrix vector multiplication operation; and

performing the conversion based on the determining,

wherein at least one bin of the syntax element is context coded, and an increasement value of the context is determined based on characteristics of a neighboring block of the current video block,

wherein a boundary down-sampling operation on reference samples of the current video block and an up-sampling operation are included in the MIP mode based on a size of the current video block, and

wherein in the boundary down-sampling operation, reduced boundary samples are directly generated from the reference samples and a downscaling factor without generating intermediate samples.

2. The method of claim 1 , wherein the increasement value of the context is determined further based on the size of the current video block.

3. The method of claim 2 , wherein in response to a width-height ratio of the current video block being greater than 2, a context with a first predefined increasement value is used for coding the at least one bin of the syntax element.

4. The method of claim 3 , wherein in response to the width-height ratio of the current video block being smaller than or equal to 2, a context with a second increasement value is used for coding the at least one bin of the syntax element, wherein the second increasement value is not identical to the first predefined increasement value.

5. The method of claim 4 , wherein the second increasement value is selected from an increasement value group based on the characteristics of the neighboring block.

6. The method of claim 5 , wherein the characteristics of the neighboring block include coding mode of the neighboring block and availability of the neighboring block.

7. The method of claim 6 , wherein a left neighboring video block of the current video block and a top neighboring video block of the current video block are used to select the second increasement value.

8. The method of claim 7 , wherein the second increasement value is derived based on a following equation:

ctxInc=(cond L && available L )+(cond A && available A ),

wherein condL is a first MIP syntax element of the left neighboring video block of the current video block,

wherein condA is a second MIP syntax element of the top neighboring video block of the current video block,

wherein availableL and availableA indicate the availability of the left neighboring video block and the top neighboring video block, respectively, and

wherein && indicates a logical And operation.

9. The method of claim 1 , wherein the downscaling factor is calculated based on a width or height of the current video block and a boundary size value.

10. The method of claim 1 , wherein the MIP mode includes multiple types, and a type index for the current video block is derived excluding referring to type indices of previous video blocks.

11. The method of claim 10 , wherein the type index for the current video block is explicitly included in the bitstream.

12. The method of claim 1 , wherein the conversion includes encoding the current video block into the bitstream.

13. The method of claim 1 , wherein the conversion includes decoding the current video block from the bitstream.

14. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

determine, for a conversion between a current video block of a video and a bitstream of the video, whether a matrix intra prediction (MIP) mode is applied on the current video block based on a syntax element, wherein in the MIP mode, prediction samples of the current video block are determined by performing a matrix vector multiplication operation; and

perform the conversion based on the determining,

wherein at least one bin of the syntax element is context coded, and an increasement value of the context is determined based on characteristics of a neighboring block of the current video block,

wherein a boundary down-sampling operation on reference samples of the current video block and an up-sampling operation are included in the MIP mode based on a size of the current video block, and

wherein in the boundary down-sampling operation, reduced boundary samples are directly generated from the reference samples and a downscaling factor without generating intermediate samples.

15. The apparatus of claim 14 , wherein the increasement value of the context is determined further based on the size of the current video block, and

wherein in response to a width-height ratio of the current video block being greater than 2, a context with a first predefined increasement value is used for coding the at least one bin of the syntax element, and in response to the width-height ratio of the current video block being smaller than or equal to 2, a context with a second increasement value is used for coding the at least one bin of the syntax element, wherein the second increasement value is not identical to the first predefined increasement value.

16. The apparatus of claim 15 , wherein the second increasement value is derived based on a following equation:

ctxInc=(cond L && available L )+(cond A && available A ),

wherein condL is a first MIP syntax element of a left neighboring video block of the current video block,

wherein condA is a second MIP syntax element of a top neighboring video block of the current video block,

wherein availableL and availableA indicate the availability of the left neighboring video block and the top neighboring video block, respectively, and

wherein && indicates a logical And operation.

17. A non-transitory computer-readable storage medium storing instructions that cause a processor to:

determine, for a conversion between a current video block of a video and a bitstream of the video, whether a matrix intra prediction (MIP) mode is applied on the current video block based on a syntax element, wherein in the MIP mode, prediction samples of the current video block are determined by performing a matrix vector multiplication operation; and

perform the conversion based on the determining,

wherein at least one bin of the syntax element is context coded, and an increasement value of the context is determined based on characteristics of a neighboring block of the current video block,

wherein a boundary down-sampling operation on reference samples of the current video block and an up-sampling operation are included in the MIP mode based on a size of the current video block, and

wherein in the boundary down-sampling operation, reduced boundary samples are directly generated from the reference samples and a downscaling factor without generating intermediate samples.

18. The non-transitory computer-readable storage medium of claim 17 , wherein the increasement value of the context is determined further based on the size of the current video block, and

wherein in response to a width-height ratio of the current video block being greater than 2, a context with a first predefined increasement value is used for coding the at least one bin of the syntax element, and in response to the width-height ratio of the current video block being smaller than or equal to 2, a context with a second increasement value is used for coding the at least one bin of the syntax element, wherein the second increasement value is not identical to the first predefined increasement value.

19. A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by a video processing apparatus, wherein the method comprises:

determining whether a matrix intra prediction (MIP) mode is applied on a current video block of the video based on a syntax element, wherein in the MIP mode, prediction samples of the current video block are determined by performing a matrix vector multiplication operation; and

generating the bitstream based on the determining,

wherein at least one bin of the syntax element is context coded, and an increasement value of the context is determined based on characteristics of a neighboring block of the current video block,

wherein a boundary down-sampling operation on reference samples of the current video block and an up-sampling operation are included in the MIP mode based on a size of the current video block, and

wherein in the boundary down-sampling operation, reduced boundary samples are directly generated from the reference samples and a downscaling factor without generating intermediate samples.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2023
From: DENG, ZHIPIN; LIU, HONGBIN
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 062643/0802 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2023
From: ZHANG, KAI; ZHANG, LI; XU, JIZHENG
To: BYTEDANCE INC.
Reel/Frame 062643/0911 →
Priority Claims (1)
WO PCT/CN2019/085399 · May 1, 2019 · international
Continuity (3)
Continuation 17479338 · Sep 20, 2021
Continuation PCTCN2020088583 · May 5, 2020
Related Publication 20230037931A1 · Feb 9, 2023