IP Library › Granted Patent US 12,744,906
Granted Patent B2
US 12,744,906 · App. 18/668,876 · Granted Sep 22, 2026

Method and apparatus for secondary transform with adaptive kernel options

Inventors: Madhu Peringassery Krishnan (Mountain View, CA); Xin Zhao (San Jose, CA); Shan Liu (San Jose, CA)
Assignee: Tencent America LLC
H04N19/13H04N19/159H04N19/176H04N19/18H04N19/91
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,906
App. No.
18/668,876
Granted
Sep 22, 2026
Kind
B2
Abstract

This disclosure relates to secondary transform of video blocks with adaptive kernel options. For example, a method for decoding a video block in a video stream is disclosed. The method may include comprising parsing and processing the video stream to generate: a set of secondary transform coefficients associated with the video block; an intra prediction mode associated with the video block; and a kernel index indicating a secondary transform kernel among a group of secondary transform kernels. The method may further include identifying the group of secondary transform kernels based on the intra prediction mode; and performing an inverse secondary transform of the set of secondary transform coefficients to generate primary transform coefficients of the video block based on the secondary transform kernel among the group of secondary transform kernels identified by the kernel index. The quantity of kernels in the group of secondary transform kernels depends on at least one of: the intra prediction mode associated with the video block; a size of the video block; or a primary transform type associated with the video block.

Claims (40)

1 . A method for decoding a video block in a video bitstream, comprising:

receiving a set of secondary transform coefficients associated with the video block;

determining an intra prediction mode associated with the video block; and

determining a first group of secondary transform kernels when the intra prediction mode is one of Vertical mode (V_PRED), Horizontal mode (H_PRED), Smooth horizontal mode (SMOOTH_H_PRED) and Smooth Vertical mode (SMOOTH_V_PRED), and when that a size of the video block meet a predetermined block size threshold condition, the first group of secondary transform kernels having N number of secondary transform kernels;

determining a second group of secondary transform kernels otherwise, the second group of secondary transform kernels having K number of secondary transform kernels;

selecting a secondary transform kernel from the first or second group of secondary transform kernels based on at least the intra prediction mode after a group of transform kernels are selected from the first or secondar group of secondary transform kernels according to a size of the video block; and

performing an inverse secondary transform of the set of secondary transform coefficients to generate primary transform coefficients of the video block based on the selected secondary transform kernel.

2 . The method of claim 1 , wherein selecting a secondary transform kernel from the first or second group of secondary transform kernels based on the intra prediction mode comprises:

selecting a group of secondary transform kernels from the first or second group of secondary transform kernels based on the intra prediction mode; and

selecting the secondary transform kernel from the selected group of secondary transform kernels based on a kernel index received from the video block in the video bitstream.

3 . The method of claim 2 , wherein a bit size of the kernel index depends on the intra prediction mode.

4 . The method of claim 2 , wherein the kernel index is entropy-coded in the video bitstream using different context models depending on which of the first group and second group of secondary transform kernel is used.

5 . The method of claim 2 , wherein, when numbers of kernels in the selected group of secondary transform kernels used for different video blocks in the video bitstream are different, entropy coding of binarized codewords of the kernel indexes of the different video blocks share context modeling for at least one bin of the binarized codewords.

6 . The method of claim 1 , wherein N and K are different non-negative integers between 0 and 6.

7 . A video encoder comprising a memory for storing computer code and at least one processor for executing the computer code to cause the video encoder to:

select an intra prediction mode for a video block;

transform a residual block of the video block using at least one primary transform kernel to generate a set of primary transform coefficients;

identify a first group of secondary transform kernels when the intra prediction mode is one of Vertical mode (V_PRED), Horizontal mode (H_PRED), Smooth horizontal mode (SMOOTH_H_PRED) and Smooth Vertical mode (SMOOTH_V_PRED), and when that a size of the video block meet a predetermined block size threshold condition, the first group of secondary transform kernels having N number of secondary transform kernels;

identify a second group of K secondary transform kernels otherwise, the second group of secondary transform kernels having K number of secondary transform kernels;

select a group of secondary transform kernels among the first group or second group of secondary transform kernels according to the selected intra prediction mode for the video block;

select a secondary transform kernel among the selected group of secondary transform kernels based on the intra prediction mode;

determine a kernel index of the selected secondary transform kernel within the selected group of secondary transform kernels;

transform the set of primary transform coefficients to generate a set of secondary transform coefficients using the selected secondary transform kernel; and

encode the kernel index and the set of secondary transform coefficients into an encoded video bitstream of the video block.

8 . The video encoder of claim 7 , wherein the at least one processor is configured to execute the computer code to further determine a range of the kernel index based on the intra prediction mode.

9 . The video encoder of claim 8 , wherein the at least one processor is configured to execute the computer code to further determine a bit size of the kernel index for encoding based on the range.

10 . The video encoder of claim 7 , wherein N and K are different non-negative integers between 0 and 6.

11 . The video encoder of claim 7 , wherein the kernel index is entropy-coded in the encoded video bitstream using different context models depending on which of the first group and second group of secondary transform kernel is selected.

12 . The video encoder of claim 7 , wherein, when numbers of kernels in the selected group of secondary transform kernels used for different video blocks in the encoded video bitstream are different, entropy coding of binarized codewords of the kernel indexes of the different video blocks share context modeling for at least one bin of the binarized codewords.

13 . A method for processing a video block, comprising converting the video block to a bitstream, wherein the bitstream comprises:

an encoded syntax element for indicating an intra prediction mode associated with the video block, wherein the intra prediction mode, when being one of Vertical mode (V_PRED), Horizontal mode (H_PRED), Smooth horizontal mode (SMOOTH_H_PRED) and Smooth Vertical mode, and when that a size of the video block meet a predetermined block size threshold condition, indicates a group of secondary transform kernels as comprising a first group of secondary transform kernels having N second transform kernels, and otherwise indicates the group of secondary transform kernels as comprising a second group of secondary transform kernels having K secondary transform kernels;

an encoded kernel index to identify a secondary transform kernel among of the group of secondary transform kernels determined according to the intra prediction mode; and

encoded secondary transform coefficients of the video block generated based on the secondary transform kernel.

14 . The method of claim 13 , wherein the encoded syntax element for indicating the intra prediction mode associated with the video block enables a video decoder to select the group of secondary transform kernels from the first and second group of secondary transform kernels.

15 . The method of claim 13 , wherein the encoded kernel index enables a video decoder to select the secondary transform kernel from the group of secondary transform kernels.

16 . The method of claim 13 , wherein a bit size of the kernel index is determined based on the intra prediction mode.

17 . The method of claim 13 , wherein N and K are different non-negative integers between 0 and 6.

18 . The method of claim 13 , wherein the kernel index is entropy-coded in the bitstream using different context models depending on which of the first group and second group of secondary transform kernel is used to generate the encoded secondary transform coefficients.

19 . The method of claim 13 , wherein, when numbers of kernels in the groups of secondary transform kernels used for different video blocks in the bitstream are different, entropy coding of binarized codewords of the kernel indexes of the different video blocks share context modeling for at least one bin of the binarized codewords.

20 . A video decoder comprising a memory for storing computer code and at least one processor for executing the computer code to cause the video decoder to perform the method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 20, 2024
From: PERINGASSERY KRISHNAN, MADHU; ZHAO, XIN; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 067467/0332 →
Continuity (3)
Continuation 17897815 · Aug 29, 2022
Provisional Application 63238635 · Aug 30, 2021
Related Publication 20240314320A1 · Sep 19, 2024
References Cited (55)
US 11290747B2 · Koo · 2022 [cited by examiner]
US 11425421B1 · Koo · 2022 [cited by examiner]
US 20170094314A1 · Zhao et al. · 2017 [cited by applicant]
US 20200177889A1 · Kim · 2020 [cited by examiner]
US 20200304818A1 · Koo et al. · 2020 [cited by applicant]
US 20200322623A1 · Chiang et al. · 2020 [cited by applicant]
US 20210076070A1 · Jung · 2021 [cited by examiner]
US 20220159300A1 · Chiang · 2022 [cited by examiner]
US 20220201335A1 · Chiang · 2022 [cited by examiner]
US 20220329819A1 · Kerofsky · 2022 [cited by examiner]
US 20220417529A1 · Zhang · 2022 [cited by examiner]
US 20230037302A1 · Rosewarne · 2023 [cited by applicant]
US 20240364903A1 · Rosewarne · 2024 [cited by examiner]
BR 122022006290B1 · 2024 [cited by examiner]
CN 111919450A · 2020 [cited by applicant]
CN 112400322A · 2021 [cited by applicant]
CN 113661710A · 2021 [cited by applicant]
CN 114009022A · 2022 [cited by applicant]
EP 3588952A1 · 2020 [cited by applicant]
EP 3790275A1 · 2021 [cited by applicant]
JP 2019517221A · 2019 [cited by applicant]
KR 1020190125482A · 2019 [cited by applicant]
WO WO2020116961A1 · 2020 [cited by applicant]
WO WO2020206286A1 · 2020 [cited by applicant]
WO WO2020259891A1 · 2020 [cited by applicant]
WO WO2021051156A1 · 2021 [cited by applicant]
Office Action mailed May 29, 2024 for Japanese Patent Application No. 2023-528040. [cited by applicant]
Chinese Office Action with English Translation, Dec. 30, 2024, pp. 1-15, issued in Chinese Application No. 202280006610.5, State Intellectual Property Office, Beijing, China. [cited by applicant]
Chinese Office Action with English translation, dated Jun. 28, 2024, pp. 1-26, issued in Chinese Patent Application No. 202280006610.5, State Intellectual Property Office, Beijing, China. [cited by applicant]
Extended European Search Report issued in European Patent Application No. 22865386.1 dated Nov. 18, 2024, 11 pages. [cited by applicant]
Abe et al. “CE6: AMT and NSST complexity reduction (CE6-3.3)” Joint Video Exploration Team (JVET) of ITU-T SG 16 WG 3 and 1S0/IEC JTC 1/SC 29/WG 11, Jul. 10, 2018, XP030198678, 6 pages. [cited by applicant]
Zhao et al. Unified Secondary Transform for Intra for Intra Coding Beyong Avl, 2020 IEEE International Conference on Image Processing (ICIP), IEEE, Oct. 25, 2020, XP033869355, 5 pages. [cited by applicant]
International Search Report mailed Dec. 29, 2022 for International Application No. PCT/US2022/041913. [cited by applicant]
Written Opinion mailed Dec. 29, 2022 for International Application No. PCT/US2022/041913. [cited by applicant]
Bross et al.; “General Video Coding Technology in Responses to the Joint Call for Proposals on Video Compression with Capability beyond HEVC”; IEEE Transactions on Circuits and Systems for Video Technology; 2019; 16 pag… [cited by applicant]
Bross et al.; “Versatile Video Coding (Draft 2)”; JVET-K1001-v6; Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 11th Meeting: Ljubljana, SI, Jul. 10-18, 2018; 139 pages. [cited by applicant]
Bross et al.; “CE3: Multiple reference line intra prediction (Test 1.1.1, 1.1.2, 1.1.3 and 1.1 .4)”, JVET-L0283-v2; Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting: Macao,… [cited by applicant]
Bross et al.; “Versatile Video Coding (Draft 6)”, JVET-O2001-vE; Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 15th Meeting: Gothenburg, SE, Jul. 3-12, 2019; 455 pages. [cited by applicant]
Chang et al.; “Intra prediction using multiple reference lines for the versatile video coding standard”; InApplications of Digital Image Processing XLII, vol. 11137; Sep. 6, 2019; 8 pages. [cited by applicant]
Racapé et al.; “CE3-related: Wide-angle intra prediction for non-square blocks”; JVET-K0500; Joint Video Experts Team (JVET) of ITU-T SG; Jul. 2018, 10 pages. [cited by applicant]
Racapé et al.; “CE3-related: Wide-angle intra prediction for non-square blocks”; JVET-K0500_r4; Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11; 11th Meeting: Ljubjana, SI, Jul. 10-18, … [cited by applicant]
De Rivaz et al.; “AV1 Bitstream & Decoding Process Specification”; The Alliance for Open Media; Jan. 8, 2019, 681 pages. [cited by applicant]
Zhang et al.; “Fast Adaptive Multiple Transform for Versatile Video Coding”; 2019 Data Compression Conference (DCC); IEEE; Mar. 26, 2019; 10 pages. [cited by applicant]
Zhang et al.; “Fast DST-7/DCT-8 with Dual Implementation Support for Versatile Video Coding”; IEEE Transactions on Circuits and Systems for Video Technology; IEEE; 17 pages. [cited by applicant]
Zhao et al.; “Novel Statistical Modeling, Analysis and Implementation of Rate-Distortion Estimation for H.264/AVC Coders”; IEEE Transactions on Circuits and Systems for Video Technology, vol. 20, No. 5; May 2010; 14 pag… [cited by applicant]
Zhao et al.; “NSST: Non-Separable Secondary Transforms for Next Generation Video Coding”; In2016 Picture Coding Symposium (PCS); IEEE; Dec. 4, 2016; 5 pages. [cited by applicant]
Zhao et al.; “Low-Complexity Intra Prediction Refinements for Video Coding”; IEEE; In2018 Picture Coding Symposium (PCS); Jun. 24, 2018; 5 pages. [cited by applicant]
Zhao et al.; “Joint Separable and Non-Separable Transforms for Next-Generation Video Coding”; IEEE Transactions on Image Processing, Feb. 5, 2018; 13 pages. [cited by applicant]
Zhao et al.; “Coupled Primary and Secondary Transform for Next Generation Video Coding”; In2018 IEEE Visual Communications and Image Processing (VCIP); Dec. 9, 2018; 4 pages. [cited by applicant]
Zhao et al.; “Wide Angular Intra Prediction for Versatile Video Coding”; 2019 Data Compression Conference (DCC); 10 pages. [cited by applicant]
Zhao et al.; “CE6: On 8-bit primary transform core (Test 6.1.3)”, JVET-L0285-r1; Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting; Macao, CN, Oct. 3-12, 2018; 17 pages. [cited by applicant]
Zhao et al.; “CE6: Fast DST-7/DCT-8 with dual implementation support (Test 6.2.3)”, JVET-M0497; Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting: Marrakech, MA, Jan. 9-18, … [cited by applicant]
Zhao et al.; “CE6-related: Unified LFNST using block size independent kernel”, JVET-O0539-v2; Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 15th Meeting: Gothenburg, SE, Jul. 3-12, 2… [cited by applicant]
Zhao et al.; “Non-CE6: Configurable max transform size in VVC”, JVET-O0545-v2; Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 15th Meeting: Gothenburg, SE, Jul. 3-12; 6 pages. [cited by applicant]
Korean-language Office Action issued in Korean Application No. 10-2023-7010445 dated Oct. 27, 2025 with English translation (8 pages). [cited by applicant]