IP Library › Granted Patent US 12,389,037
Granted Patent B2
US 12,389,037 · App. 18/193,131 · Granted Aug 12, 2025

Use of secondary transform in coded video

Inventors: Kai Zhang (San Diego, CA); Li Zhang (San Diego, CA); Hongbin Liu (Beijing, CN); Jizheng Xu (San Diego, CA); Yue Wang (Beijing, CN)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/61H04N19/11H04N19/124H04N19/132H04N19/159H04N19/176H04N19/186H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,389,037
App. No.
18/193,131
Granted
Aug 12, 2025
Kind
B2
Abstract

A video processing method includes determining, for a conversion between a block of a video and a bitstream representation of the video, that a secondary transform with a reduced dimension dimension (e.g., an inverse low frequency non-separable transform) is applicable to a single sub-block of the block in case a dimension of the block satisfies a condition. The secondary transform is performed between a forward primary transform and a quantization step or between a de-quantization step and an inverse primary transform. The reduced dimension is reduced from a dimension of the block. The method also includes performing the conversion based on the determining.

Claims (67)

1. A method of processing video data, comprising:

determining, for a conversion between a current block of a video and a bitstream of the video, that a secondary transform is applicable to the current block, wherein the secondary transform comprises at least one of a forward secondary transform and an inverse secondary transform, wherein the forward secondary transform is performed between a forward primary transform and a quantization, and the inverse secondary transform is performed between a de-quantization and an inverse primary transform;

determining, in response to a dimension of the current block satisfying a first condition, that the secondary transform with an 8×8 secondary transform size is applicable to a first single top-left sub-block of the current block with a dimension of 8×8, and wherein the first condition requires that the dimension of the current block is W1×H1, and wherein H1≥8 and W1≥8;

determining, in response to the dimension of the current block satisfying a second condition, that the secondary transform with a 4×4 secondary transform size is applicable to a second single top-left sub-block of the current block with a dimension of 4×4 and that no secondary transform is applied to a sub-block having a dimension of 4×4 and adjacent to the second single top-left sub-block, and wherein the second condition requires that the dimension of the current block is 4×H1 or W1×4, wherein H1>8 and W1>8; and

performing the conversion based on the determining,

wherein an output value from the inverse secondary transform is constrained within a first range of [min, max] inclusively, wherein min and max are integer values, and wherein min is equal to −(1<<15) and max is equal to (1<<15)−1, and

wherein coefficients after the de-quantization are constrained to a second range of [qmin, qmax] inclusively, wherein qmin is equal to the min of the first range and qmax is equal to the max of the first range, qmin and qmax being integers.

2. The method of claim 1 , wherein a matrix for the secondary transform is selected from four transform sets, and each of the four transform sets consists of two transform matrices.

3. The method of claim 1 , wherein whether to apply the secondary transform depends on a coding mode of a block,

wherein in response to the block being coded with a transform skip mode, the secondary transform is not applied to the block, or

wherein in response to the block being coded with a non intra prediction mode, the secondary transform is not applied to the block.

4. The method of claim 1 , wherein a quantization matrix used in the de-quantization is determined based on whether the inverse secondary transform is applied or not.

5. The method of claim 1 , wherein a clipping operation is applied to clip the output value within the first range of [min, max] inclusively.

6. The method of claim 1 , wherein the inverse secondary transform is applied to de-quantized transformed coefficients of the current block.

7. The method of claim 1 , wherein the secondary transform comprises a low frequency non-separable transform.

8. The method of claim 1 , wherein the conversion includes encoding the video into the bitstream.

9. The method of claim 1 , wherein the conversion includes decoding the video from the bitstream.

10. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

determine, for a conversion between a current block of a video and a bitstream of the video, that a secondary transform is applicable to the current block, wherein the secondary transform comprises at least one of a forward secondary transform and an inverse secondary transform, wherein the forward secondary transform is performed between a forward primary transform and a quantization, and the inverse secondary transform is performed between a de-quantization and an inverse primary transform;

determine, in response to a dimension of the current block satisfying a first condition, that the secondary transform with an 8×8 secondary transform size is applicable to a first single top-left sub-block of the current block with a dimension of 8×8, and wherein the first condition requires that the dimension of the current block is W1×H1, and wherein H1≥8 and W1≥8;

determine, in response to the dimension of the current block satisfying a second condition, that the secondary transform with a 4×4 secondary transform size is applicable to a second single top-left sub-block of the current block with a dimension of 4×4 and that no secondary transform is applied to a sub-block having a dimension of 4×4 and adjacent to the second single top-left sub-block, and wherein the second condition requires that the dimension of the current block is 4×H1 or W1×4, wherein H1>8 and W1>8; and

perform the conversion based on the determining,

wherein an output value from the inverse secondary transform is constrained within a first range of [min, max] inclusively, wherein min and max are integer values, and wherein min is equal to −(1<<15) and max is equal to (1<<15)−1, and

wherein coefficients after the de-quantization are constrained to a second range of [qmin, qmax]inclusively, wherein qmin is equal to the min of the first range and qmax is equal to the max of the first range, qmin and qmax being integers.

11. The apparatus of claim 10 , wherein a matrix for the secondary transform is selected from four transform sets, and each of the four transform sets consists of two transform matrices,

whether to apply the secondary transform depends on a coding mode of a block,

wherein in response to the block being coded with a transform skip mode, the secondary transform is not applied to the block, or

wherein in response to the block being coded with a non intra prediction mode, the secondary transform is not applied to the block, and

wherein the secondary transform comprises a low frequency non-separable transform.

12. The apparatus of claim 10 , wherein a quantization matrix used in the de-quantization is determined based on whether the inverse secondary transform is applied or not,

wherein a clipping operation is applied to clip the output value within the first range of [min, max] inclusively, and

wherein the inverse secondary transform is applied to de-quantized transformed coefficients of the current block.

13. The apparatus of claim 10 , wherein the conversion includes encoding the video into the bitstream.

14. The apparatus of claim 10 , wherein the conversion includes decoding the video from the bitstream.

15. A non-transitory computer-readable storage medium storing instructions that cause a processor to:

determine, for a conversion between a current block of a video and a bitstream of the video, that a secondary transform is applicable to the current block, wherein the secondary transform comprises at least one of a forward secondary transform and an inverse secondary transform, wherein the forward secondary transform is performed between a forward primary transform and a quantization, and the inverse secondary transform is performed between a de-quantization and an inverse primary transform;

determine, in response to a dimension of the current block satisfying a first condition, that the secondary transform with an 8×8 secondary transform size is applicable to a first single top-left sub-block of the current block with a dimension of 8×8, and wherein the first condition requires that the dimension of the current block is W1×H1, and wherein H1≥8 and W1≥8;

determine, in response to the dimension of the current block satisfying a second condition, that the secondary transform with a 4×4 secondary transform size is applicable to a second single top-left sub-block of the current block with a dimension of 4×4 and that no secondary transform is applied to a sub-block having a dimension of 4×4 and adjacent to the second single top-left sub-block, and wherein the second condition requires that the dimension of the current block is 4×H1 or W1×4, wherein H1>8 and W1>8; and

perform the conversion based on the determining,

wherein an output value from the inverse secondary transform is constrained within a first range of [min, max] inclusively, wherein min and max are integer values, and wherein min is equal to −(1<<15) and max is equal to (1<<15)−1, and

wherein coefficients after the de-quantization are constrained to a second range of [qmin, qmax] inclusively, wherein qmin is equal to the min of the first range and qmax is equal to the max of the first range, qmin and qmax being integers.

16. The non-transitory computer-readable storage medium of claim 15 , wherein a matrix for the secondary transform is selected from four transform sets, and each of the four transform sets consists of two transform matrices,

whether to apply the secondary transform depends on a coding mode of a block,

wherein in response to the block being coded with a transform skip mode, the secondary transform is not applied to the block, or

wherein in response to the block being coded with a non intra prediction mode, the secondary transform is not applied to the block, and

wherein the secondary transform comprises a low frequency non-separable transform.

17. The non-transitory computer-readable storage medium of claim 15 ,

wherein a quantization matrix used in the de-quantization is determined based on whether the inverse secondary transform is applied or not,

wherein a clipping operation is applied to clip the output value within the first range of [min, max] inclusively, and

wherein the inverse secondary transform is applied to de-quantized transformed coefficients of the current block.

18. A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by a video processing apparatus, wherein the method comprises:

determining that a secondary transform is applicable to a current block of the video, wherein the secondary transform comprises at least one of a forward secondary transform and an inverse secondary transform, wherein the forward secondary transform is performed between a forward primary transform and a quantization, and the inverse secondary transform is performed between a de-quantization and an inverse primary transform;

determining, in response to a dimension of the current block satisfying a first condition, that the secondary transform with an 8×8 secondary transform size is applicable to a first single top-left sub-block of the current block with a dimension of 8×8, and wherein the first condition requires that the dimension of the current block is W1×H1, and wherein H1≥8 and W1≥8;

determining, in response to the dimension of the current block satisfying a second condition, that the secondary transform with a 4×4 secondary transform size is applicable to a second single top-left sub-block of the current block with a dimension of 4×4 and that no secondary transform is applied to a sub-block having a dimension of 4×4 and adjacent to the second single top-left sub-block, and wherein the second condition requires that the dimension of the current block is 4×H1 or W1×4, wherein H1>8 and W1>8; and

generating the bitstream of the video based on the determining,

wherein an output value from the inverse secondary transform is constrained within a first range of [min, max] inclusively, wherein min and max are integer values, and wherein min is equal to −(1<<15) and max is equal to (1<<15)−1, and

wherein coefficients after the de-quantization are constrained to a second range of [qmin, qmax] inclusively, wherein qmin is equal to the min of the first range and qmax is equal to the max of the first range, qmin and qmax being integers.

19. The non-transitory computer-readable recording medium of claim 18 ,

wherein a matrix for the secondary transform is selected from four transform sets, and each of the four transform sets consists of two transform matrices,

whether to apply the secondary transform depends on a coding mode of a block,

wherein in response to the block being coded with a transform skip mode, the secondary transform is not applied to the block, or

wherein in response to the block being coded with a non intra prediction mode, the secondary transform is not applied to the block, and

wherein the secondary transform comprises a low frequency non-separable transform.

20. The non-transitory computer-readable recording medium of claim 18 ,

wherein a quantization matrix used in the de-quantization is determined based on whether the inverse secondary transform is applied or not,

wherein a clipping operation is applied to clip the output value within the first range of [min, max] inclusively, and

wherein the inverse secondary transform is applied to de-quantized transformed coefficients of the current block.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2023
From: ZHANG, KAI; ZHANG, LI; XU, JIZHENG
To: BYTEDANCE INC.
Reel/Frame 063884/0332 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2023
From: LIU, HONGBIN; WANG, YUE
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 063884/0403 →
Priority Claims (1)
CN 2019/083853 · Apr 23, 2019 · national
Continuity (3)
Continuation 17406242 · Aug 19, 2021
Continuation PCTCN2020086444 · Apr 23, 2020
Related Publication 20230262263A1 · Aug 17, 2023
References Cited (92)
US 5822004A · Crocitti · 1998 [cited by applicant]
US 6389072B1 · Tzou · 2002 [cited by applicant]
US 10666976B2 · Huang et al. · 2020 [cited by applicant]
US 11647229B2 · Zhang · 2023 [cited by examiner]
US 20020044605A1 · Nakamura · 2002 [cited by applicant]
US 20040264571A1 · Zhang · 2004 [cited by applicant]
US 20130003856A1 · Saxena · 2013 [cited by applicant]
US 20140254676A1 · Jiang · 2014 [cited by applicant]
US 20160050422A1 · Rosewarne · 2016 [cited by applicant]
US 20170094313A1 · Zhao · 2017 [cited by applicant]
US 20180041776A1 · Kim · 2018 [cited by applicant]
US 20180103252A1 · Hsieh · 2018 [cited by applicant]
US 20180302631A1 · Chiang · 2018 [cited by applicant]
US 20180338143A1 · Fracastoro · 2018 [cited by applicant]
US 20180367814A1 · Seregin · 2018 [cited by applicant]
US 20190028701A1 · Yu · 2019 [cited by applicant]
US 20190141334A1 · Lim · 2019 [cited by applicant]
US 20190356915A1 · Jang · 2019 [cited by examiner]
US 20200288134A1 · Lim · 2020 [cited by applicant]
US 20200288172A1 · Huang · 2020 [cited by applicant]
US 20200322617A1 · Zhao · 2020 [cited by examiner]
US 20210321134A1 · Koo · 2021 [cited by examiner]
US 20210352326A1 · Lim · 2021 [cited by applicant]
US 20220109876A1 · Zhang · 2022 [cited by applicant]
US 20220159300A1 · Chiang · 2022 [cited by applicant]
US 20220182675A1 · Zhang · 2022 [cited by applicant]
US 20220201335A1 · Chiang · 2022 [cited by applicant]
US 20220345744A1 · Leleannec · 2022 [cited by applicant]
CN 104025589A · 2014 [cited by applicant]
CN 105516730A · 2016 [cited by applicant]
CN 108141594A · 2018 [cited by applicant]
CN 108141596A · 2018 [cited by applicant]
CN 108141597A · 2018 [cited by applicant]
CN 108322745A · 2018 [cited by applicant]
CN 108632611A · 2018 [cited by applicant]
CN 108712649A · 2018 [cited by applicant]
CN 109076222A · 2018 [cited by applicant]
CN 109076226A · 2018 [cited by applicant]
CN 109076230A · 2018 [cited by applicant]
CN 109076242A · 2018 [cited by applicant]
CN 109076243A · 2018 [cited by applicant]
CN 109644269A · 2019 [cited by applicant]
CN 110636313B · 2022 [cited by applicant]
EP 3349451A1 · 2018 [cited by applicant]
EP 3506634A4 · 2019 [cited by applicant]
JP 7509944B2 · 2024 [cited by applicant]
KR 20180041578A · 2018 [cited by applicant]
KR 20210120802A · 2021 [cited by applicant]
KR 102321394B1 · 2021 [cited by applicant]
WO 2017195555A1 · 2017 [cited by applicant]
WO 2017195666A1 · 2017 [cited by applicant]
WO 2018037737A1 · 2018 [cited by applicant]
WO 2018166429A1 · 2018 [cited by applicant]
WO 2018174402A1 · 2018 [cited by applicant]
WO 2019022099A1 · 2019 [cited by applicant]
WO 2020046092A1 · 2020 [cited by applicant]
WO 2021194052A1 · 2021 [cited by applicant]
Siekmann et al. “CE6-Related: Simplification of the Reduced Secondary Transform,” Joint Video Experts Team JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting, Geneva, CH, Mar. 19-27, 2019, document JV… [cited by applicant]
Koo et al. “CE6-5.1: Reduced Secondary Transform (RST),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting: Marrakech, MA, Jan. 9-18, 2019, document JVET-M0292, 2019. [cited by applicant]
Zhao et al. “TU-Level Non-Separable Secondary Transform,” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/1/VG 11, 2nd Meeting, San Diego, USA, Feb. 20-26, 2016, document JVET-B0059, 2016. [cited by applicant]
Xinzzhao, From HEVC to VVC Evolution of Transform Technology (2)—Secondary Transform; retrieved from the Internet on Jun. 24, 2021 <URL: https:/lcloud.tencent.com/developer/article/1427150>, May 16, 2019. [cited by applicant]
Koo et al. “Low Frequency Non-Separable Transform (LFNST),” 2019 Picture Coding Symposium (PCS), Jan. 9, 2020. [cited by applicant]
Zhao et al. “NSST: Non-Separable Secondary Transforms for Next Generation Video Coding,” 2016 Picture Coding Symposium (PCS), Apr. 24, 2017. [cited by applicant]
Bross et al. “Versatile Video Coding (Draft 5),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting, Geneva, CH, Mar. 19-27, 2019, documentJVET-N1001, 2019. [cited by applicant]
Retrieved from the internet: vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/tags/VTM-4.0, Nov. 12, 2021. [cited by applicant]
De-Luxan-Hernandez et al. “CE3: Intra Sub-Partitions Coding Mode {Tests 1.1.1 and 1.1.2),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting, Marrakech, MA, Jan. 9-18, 2019,… [cited by applicant]
Abdoli et al. ““CEB: BDPCM with HorizontalNertical Predictor and Independently Decodable Areas (test 8.3.1b),”” Joint Video Experts Team (JVET)of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 1113th Meeting: Marrakech, MA… [cited by applicant]
Karczewicz et al. “CE8-Related: Quantized Residual BDPCM,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting, Geneva, CH, Mar. 19-27, 2019, document JVET-N0413, 2019. [cited by applicant]
Koo et al. “CE6: Reduced Secondary Transform (RST) {CE6-3.1),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting, Geneva, CH, Mar. 19-27, 2019, document JVET-N0193, 019. [cited by applicant]
Salehifar et al. “CE 6.2.6: Reduced Secondary Transform (RST),” Joint Video Experts Team (JVET) of ITU-T SG 6 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 11th Meeting, Ljubljana, SI, Jul. 10-18, 2018, document JVET-K0099, 2018. [cited by applicant]
Koo et al. “CE6-2.1: Reduced Secondary Transform (RST),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macao, CN, Oct. 3-12, 2018, document JVET-L0133, 2018. [cited by applicant]
Pfaff et al. “CE3: Affine Linear Weighted Intra Prediciton {CE3-4.1, CE3-4.2),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting, Geneva, CH, Mar. 19-27, 2019, document JVE… [cited by applicant]
Zhang et al. “CE4-Related: Interweaved Prediction for Affine Motion Compensation,” Joint Video Experts Team JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/1/VG 11, 11th Meeting, Ljubljana, SI, Jul. 10-18, 2018. docum… [cited by applicant]
Chen et al. “Algorithm description for Versatile Video Coding and Test Model 5 (VTM 5),” Joint Video Experts Team JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 1114th Meeting: Geneva, CH, Mar. 19-27, 2019, docume… [cited by applicant]
Fan et al. “Non-CE6: A Unified Zero-Out Range for 4x4 LFNST,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 16th Meeting, Geneva, CH Oct. 1-11, 2019, document JVET-P0379, 2019. [cited by applicant]
Non Final Office Action from U.S. Appl. No. 17/406,260 dated Dec. 29, 2021. [cited by applicant]
Internnational Search Report and Written Opinion from International Patent Application No. PCT/CN2020/086421 dated Jul. 22, 2020 (11 pages). [cited by applicant]
Internnational Search Report and Written Opinion from International Patent Application No. PCT/CN2020/086444 dated Aug. 11, 2020 (11 pages). [cited by applicant]
International Search Report and Written Opinion from International Patent Application No. PCT/CN2020/086458 dated Jul. 6, 2020 (9 pages). [cited by applicant]
International Search Report and Written Opinion from International Patent Application No. PCT/CN2020/096046 dated Sep. 24, 2020 (11 pages). [cited by applicant]
International Search Report and Written Opinion from International Patent Application No. PCT/CN2020/133273 dated Feb. 18, 2021 (10 pages). [cited by applicant]
International Search Report and Written Opinion from International Patent Application No. PCT/CN2020/082962 dated Jul. 1, 2021 (10 pages). [cited by applicant]
Non Final Office Action from U.S. Appl. No. 17/406,242 dated Dec. 13, 2021. [cited by applicant]
Non Final Office Action from U.S. Appl. No. 17/406,242 dated Aug. 4, 2022. [cited by applicant]
Document: JVET-N0555-v3, Siekmann, M., et al., “CE6-related: Simplification of the Reduced Secondary Transform,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 14th Meeting: Geneva, CH… [cited by applicant]
JVET-N0555-v1, Siekmann, M., et al., “CE6—related: Simplification of the Reduced Secondary Transform,” Fraunhofer HHI, Mar. 24, 2019, 10 pages. [cited by applicant]
Document: JVET-E0380, Alshina, E., et al., “CE7: Experimental results of ROT by Samsung,” Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11 5th Meeting: Geneva, CH, Mar. 16-2… [cited by applicant]
Document: JCTVC-10408, Lan, C., et al., “Intra transform skipping,” Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 9th Meeting: Geneva, CH, 27 Apr.-May 7, 2012, 6 pag… [cited by applicant]
Document: JVET-N0193, Koo, M., et al., “CE6: Reduced Secondary Transform (RST) (CE6-3.1),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 14th Meeting: Geneva, CH, Mar. 19-27, 2019, 18… [cited by applicant]
Document: JVET-M0292, Koo., M., et al., “CE6: Reduced Secondary Transform (RST) (test 6.5.1),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting: Marrakech, MA, Jan. 9-18, 2… [cited by applicant]
Columbia TriStar Home Video DVD, optical disc storing video bitstream of motion picture “Anatomy of a Murder,” Year 2000, 3 pages. [cited by applicant]
Non-Final Office Action from U.S. Appl. No. 17/953,045 dated Mar. 29, 2024, 26 pages. [cited by applicant]