IP Library › Granted Patent US 12,615,365
Granted Patent B2
US 12,615,365 · App. 17/658,803 · Granted Apr 28, 2026

Intra-mode dependent multiple transform selection for video coding

Inventors: Bappaditya Ray (San Diego, CA); Muhammed Zeyd Coban (Carlsbad, CA); Louis Joseph Kerofsky (San Diego, CA); Vadim Seregin (San Diego, CA); Marta Karczewicz (San Diego, CA); Keming Cao (San Diego, CA)
Assignee: QUALCOMM Incorporated
H04N19/12H04N19/159H04N19/176H04N19/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,615,365
App. No.
17/658,803
Granted
Apr 28, 2026
Kind
B2
Abstract

An example device for decoding video data includes a memory configured to store video data; and one or more processors implemented in circuitry and configured to: determine a size of a current block of video data; determine an intra-prediction mode for the current block of video data; determine a mode group including the determined intra-prediction mode, the mode group being one of a plurality of mode groups, each including respective sets of intra-prediction modes; determine a set of available multiple transform selection (MTS) schemes for the current block according to the size and the intra-prediction mode for the current block; determine an MTS scheme from the set of available MTS schemes according to the determined mode group; apply transforms of the MTS scheme to a transform block of the current block to produce a residual block for the current block; and decode the current block using the residual block.

Claims (146)

1 . A method of decoding video data, the method comprising:

determining a size of a current block of video data;

determining an intra-prediction mode for the current block of video data, wherein the determined intra-prediction mode comprises a matrix intra-prediction (MIP) mode having a transpose flag value;

determining a mode group including the determined intra-prediction mode, the mode group being one of a plurality of mode groups, each of the mode groups in the plurality of mode groups including respective sets of intra-prediction modes such that each possible intra-prediction mode is included in no more than one of the mode groups;

determining a set of available multiple transform selection (MTS) schemes for the current block according to the size and the MIP intra-prediction mode for the current block, the set of available MTS schemes being one set of available MTS schemes of a plurality of sets of MTS schemes, each of the sets of MTS schemes of the plurality of sets of MTS schemes including a common number of MTS schemes, the common number being greater than one;

determining an MTS scheme from the set of available MTS schemes according to the determined mode group, wherein each of the MTS schemes of the plurality of sets of MTS schemes includes a respective transform pair including a respective horizontal transform and a respective vertical transform, wherein the current block has a size of W×H, wherein W is not equal to H;

applying transforms of the MTS scheme to a transform block of the current block to produce a residual block for the current block, comprising:

based on the transpose flag value being a first value:

applying the respective horizontal transform of the MTS scheme and the respective vertical transform of the MTS scheme to the transform block; or

based on the transpose flag value being a second value, different from the first value:

transposing the respective horizontal transform of the MTS scheme to form a transposed vertical transform;

transposing the respective vertical transform of the MTS scheme to form a transposed horizontal transform; and

applying the transposed horizontal transform and the transposed vertical transform to the transform block; and

decoding the current block using the residual block.

2 . The method of claim 1 , wherein the plurality of mode groups includes a first mode group including intra-prediction modes 0 and 1, a second group including intra-prediction modes 2 to 12, a third group including intra-prediction modes 13 to 23, a fourth group including intra-prediction modes 24 to 34, and a fifth group including matrix intra-prediction (MIP) mode.

3 . The method of claim 1 , wherein the size of the current block comprises a width of the current block and a height of the current block, and wherein the size of the current block is included in a size group.

4 . The method of claim 3 , wherein the size group of the current block is selected from one of a plurality of size groups including 4×4, 4×8, 4×16, 4×N, 8×4, 8×8, 8×16, 8×N, 16×4, 16×8, 16×16, 16×N, N×4, N×8, N×16, N×N, wherein N is an integer power of 2 and greater than 16.

5 . The method of claim 4 , wherein determining the set of available MTS schemes according to the size of the current block comprises determining the set of available MTS according to the size group for the current block.

6 . The method of claim 1 , further comprising decoding an MTS index value representing the MTS scheme of the set of available MTS schemes, wherein determining the MTS scheme comprises determining the MTS scheme using the MTS index value.

7 . The method of claim 6 , wherein the MTS index value has a value between 0 and 3, inclusive, wherein the plurality of sets of MTS schemes comprises:

{17, 18, 23, 24},

{3, 7, 18, 22},

{2, 17, 18, 22},

{3, 15, 17, 18},

{3, 12, 18, 19},

{12, 18, 19, 23},

{2, 12, 17, 18},

{2, 17, 18, 22},

{2, 11, 17, 18},

{12, 18, 19, 23},

{12, 13, 16, 24},

{2, 11, 16, 23},

{2, 13, 17, 22},

{2, 11, 17, 21},

{13, 16, 19, 22},

{7, 12, 13, 18},

{1, 11, 12, 16},

{3, 13, 17, 22},

{1, 6, 12, 22},

{12, 13, 15, 16},

{18, 19, 23, 24},

{2, 17, 18, 24},

{3, 4, 17, 22},

{12, 18, 19, 23},

{12, 18, 19, 23},

{6, 12, 18, 24},

{2, 6, 12, 21},

{1, 11, 17, 22},

{3, 11, 16, 17},

{8, 12, 19, 23},

{7, 13, 16, 23},

{1, 6, 11, 12},

{1, 11, 17, 21},

{6, 11, 17, 21},

{8, 11, 14, 17},

{6, 11, 12, 21},

{1, 6, 11, 12},

{2, 6, 11, 12},

{1, 6, 11, 21},

{7, 11, 12, 16},

{8, 12, 19, 24},

{1, 13, 18, 22},

{2, 6, 17, 21},

{11, 12, 16, 19},

{8, 12, 17, 24},

{6, 12, 19, 21},

{6, 12, 13, 21},

{2, 16, 17, 21},

{6, 17, 19, 23},

{6, 12, 14, 17},

{6, 7, 11, 21},

{1, 11, 12, 16},

{1, 6, 11, 12},

{6, 11, 12, 21},

{7, 8, 9, 11},

{6, 7, 11, 12},

{6, 7, 11, 12},

{1, 11, 12, 16},

{6, 11, 17, 21},

{6, 7, 11, 12},

{12, 14, 18, 21},

{1, 11, 16, 22},

{1, 11, 16, 22},

{7, 13, 15, 16},

{1, 8, 12, 19},

{6, 7, 9, 12},

{2, 6, 12, 13},

{1, 12, 16, 21},

{7, 11, 16, 19},

{7, 8, 11, 12},

{6, 7, 11, 12},

{6, 7, 11, 12},

{1, 6, 11, 12},

{6, 7, 11, 16},

{6, 7, 11, 12},

{6, 7, 11, 12},

{6, 11, 12, 21},

{1, 6, 11, 12},

{6, 7, 11, 12},

{6, 7, 11, 12},

and wherein the MTS index indicates a transform pair of the set of available MTS schemes according to:

{DCT8, DCT8}, {DCT8, DST7}, {DCT8, DCT5}, {DCT8, DST4}, {DCT8, DST1},

{DST7, DCT8}, {DST7, DST7}, {DST7, DCT5}, {DST7, DST4}, {DST7, DST1},

{DCT5, DCT8}, {DCT5, DST7}, {DCT5, DCT5}, {DCT5, DST4}, {DCT5, DST1},

{DST4, DCT8}, {DST4, DST7}, {DST4, DCT5}, {DST4, DST4}, {DST4, DST1},

{DST1, DCT8}, {DST1, DST7}, {DST1, DCT5}, {DST1, DST4}, {DST1, DST1}.

8 . The method of claim 1 , wherein the common number of MTS schemes of each of the sets of MTS schemes is four.

9 . The method of claim 1 , wherein decoding the current block comprises:

forming a prediction block for the current block using the intra-prediction mode; and

adding samples of the prediction block to corresponding samples of the residual block.

10 . The method of claim 1 , further comprising encoding the current block prior to decoding the current block.

11 . A device for decoding video data, the device comprising:

a memory configured to store video data; and

one or more processors implemented in circuitry and configured to:

determine a size of a current block of video data;

determine an intra-prediction mode for the current block of video data, wherein the determined intra-prediction mode comprises a matrix intra-prediction (MIP) mode having a transpose flag value;

determine a mode group including the determined intra-prediction mode, the mode group being one of a plurality of mode groups, each of the mode groups in the plurality of mode groups including respective sets of intra-prediction modes such that each possible intra-prediction mode is included in no more than one of the mode groups;

determine a set of available multiple transform selection (MTS) schemes for the current block according to the size and the intra-prediction mode for the current block, the set of available MTS schemes being one set of available MTS schemes of a plurality of sets of MTS schemes, each of the sets of MTS schemes of the plurality of sets of MTS schemes including a common number of MTS schemes, the common number being greater than one;

determine an MTS scheme from the set of available MTS schemes according to the determined mode group, wherein each of the MTS schemes of the plurality of sets of MTS schemes includes a respective transform pair including a respective horizontal transform and a respective vertical transform, wherein the current block has a size of W×H, wherein W is not equal to H;

apply transforms of the MTS scheme to a transform block of the current block to produce a residual block for the current block, wherein to apply the transforms, the processors are further configured to:

based on the transpose flag value being a first value:

apply the respective horizontal transform of the MTS scheme and the respective vertical transform of the MTS scheme to the transform block; or

based on the transpose flag value being a second value, different than the first value:

transpose the respective horizontal transform of the MTS scheme to form a transposed vertical transform;

transpose the respective vertical transform of the MTS scheme to form a transposed horizontal transform; and

apply the transposed horizontal transform and the transposed vertical transform to the transform block; and

decode the current block using the residual block.

12 . The device of claim 11 , wherein the plurality of mode groups includes a first mode group including intra-prediction modes 0 and 1, a second group including intra-prediction modes 2 to 12, a third group including intra-prediction modes 13 to 23, a fourth group including intra-prediction modes 24 to 34, and a fifth group including matrix intra-prediction (MIP) mode.

13 . The device of claim 11 , wherein the size of the current block comprises a width of the current block and a height of the current block, and wherein the size of the current block is included in a size group.

14 . The device of claim 11 , wherein the one or more processors are further configured to decode an MTS index value representing the MTS scheme of the set of available MTS schemes, and wherein the one or more processors are configured to determine the MTS scheme using the MTS index value.

15 . The device of claim 11 , wherein the one or more processors are further configured to encode the current block prior to decoding the current block.

16 . The device of claim 11 , further comprising a display configured to display the decoded video data.

17 . The device of claim 11 , wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

18 . A non-transitory computer-readable medium having stored thereon instructions that when executed cause one or more processors to:

determine a size of a current block of video data;

determine an intra-prediction mode for the current block of video data wherein the determined intra-prediction mode comprises a matrix intra-prediction (MIP) mode having a transpose flag value;

determine a mode group including the determined intra-prediction mode, the mode group being one of a plurality of mode groups, each of the mode groups in the plurality of mode groups including respective sets of intra-prediction modes such that each possible intra-prediction mode is included in no more than one of the mode groups;

determine a set of available multiple transform selection (MTS) schemes for the current block according to the size and the intra-prediction mode for the current block, the set of available MTS schemes being one set of available MTS schemes of a plurality of sets of MTS schemes, each of the sets of MTS schemes of the plurality of sets of MTS schemes including a common number of MTS schemes, the common number being greater than one;

determine an MTS scheme from the set of available MTS schemes according to the determined mode group, wherein each of the MTS schemes of the plurality of sets of MTS schemes includes a respective transform pair including a respective horizontal transform and a respective vertical transform, wherein the current block has a size of W×H, wherein W is not equal to H;

apply transforms of the MTS scheme to a transform block of the current block to produce a residual block for the current block, wherein to apply the transforms, the instructions further cause the processors to:

based on the transpose flag value being a first value:

apply the respective horizontal transform of the MTS scheme and the respective vertical transform of the MTS scheme to the transform block; or based on the transpose flag value being a second value, different than the first value:

transpose the respective horizontal transform of the MTS scheme to form a transposed vertical transform;

transpose the respective vertical transform of the MTS scheme to form a transposed horizontal transform; and

apply the transposed horizontal transform and the transposed vertical transform to the transform block; and

decode the current block using the residual block.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2022
From: RAY, BAPPADITYA; COBAN, MUHAMMED ZEYD; KEROFSKY, LOUIS JOSEPH; SEREGIN, VADIM; KARCZEWICZ, MARTA; CAO, KEMING
To: QUALCOMM INCORPORATED
Reel/Frame 060018/0602 →
Continuity (3)
Provisional Application 63223377 · Jul 19, 2021
Provisional Application 63173884 · Apr 12, 2021
Related Publication 20220329800A1 · Oct 13, 2022
References Cited (37)
US 20180020218A1 · Zhao et al. · 2018 [cited by applicant]
US 20190215521A1 · Chuang · 2019 [cited by examiner]
US 20210211729A1 · Koo · 2021 [cited by examiner]
US 20210235119A1 · Kim · 2021 [cited by examiner]
US 20220038740A1 · Zhao · 2022 [cited by examiner]
US 20220060700A1 · Kang · 2022 [cited by examiner]
US 20220060751A1 · Nam et al. · 2022 [cited by applicant]
US 20220150504A1 · Koo · 2022 [cited by examiner]
US 20220224922A1 · Wang · 2022 [cited by examiner]
US 20220264151A1 · Lim · 2022 [cited by examiner]
US 20220329862A1 · Huo · 2022 [cited by examiner]
US 20220360785A1 · Huo · 2022 [cited by examiner]
US 20220385906A1 · Zhao et al. · 2022 [cited by applicant]
US 20230024223A1 · Le Leannec · 2023 [cited by examiner]
US 20230328287A1 · Pfaff · 2023 [cited by examiner]
US 20240205392A1 · Wang · 2024 [cited by examiner]
JP 2022532114A · 2022 [cited by applicant]
WO 2020226424A1 · 2020 [cited by applicant]
WO 2020242183A1 · 2020 [cited by applicant]
Abdoli M., et al., “Non-CE3: Decoder-Side Intra Mode Derivation with Prediction Fusion Using Planar”, JVET-00449-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 15th Meeting: Gothe… [cited by applicant]
Bross B., et al., “Versatile Video Coding (Draft 10)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 131, MPEG Meeting, 19th Meeting, by Teleconference, Jun. 22-Jul. 1, 2020, Jun. 29… [cited by applicant]
Cao K., et al., “EE2-Related: Fusion for Template-Based Intra Mode Derivation”, JVET-W0123-V2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 23rd Meeting, by teleconference, Jul. 7-16, 202… [cited by applicant]
Chang Y-J., et al., “Compression Efficiency Methods Beyond VVC”, 21. JVET Meeting, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, JVET-U0100, 133. MPEG Meeting, 21st Meeting, by teleconfere… [cited by applicant]
Chen J., et al., “Algorithm Description for Versatile Video Coding and Test Model 10 (VTM 10)”, JVET-S2002-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 131. MPEG Meeting, 19th M… [cited by applicant]
ITU-T H.265: “Series H: Audiovisual and Multimedia Systems Infrastructure of Audiovisual Services—Coding of Moving Video”, High Efficiency Video Coding, The International Telecommunication Union, Jun. 2019, 696 Pages. [cited by applicant]
ITU-T H.266: “Series H: Audiovisual and Multimedia Systems Infrastructure of Audiovisual Services—Coding of Moving Video”, Versatile Video Coding, The International Telecommunication Union, Aug. 2020, 516 pages. [cited by applicant]
Karczewicz M., et al., “Common Test Conditions and Evaluation Procedures for Enhanced Compression Tool Testing”, JVET-V2017-v1, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 22nd Meeting, … [cited by applicant]
Ray B., et al., “EE2: Enhanced Intra MTS and LFNST (Tests 4.1, 4.2, and 4.4)”, JVET-W0103-v4, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 23rd Meeting, by teleconference, Jul. 7-16, 2021… [cited by applicant]
Ray B., et al., “Enhanced Intra MTS and LFNST for Compression Beyond VVC”, JVET-V0116-v3, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 22nd Meeting, by teleconference, Apr. 20-28, 2021, p… [cited by applicant]
Said A., et al., “CE6.1.1: Extended AMT”, JVET-K0375-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 11th Meeting: Ljubljana, SI, Jul. 10-18, 2018, pp. 1-11. [cited by applicant]
Seregin V., et al., “Exploration Experiment on Enhanced Compression Beyond VVC capability (EE2)”, JVET-V2024-v3, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 22nd Meeting, by teleconferen… [cited by applicant]
Seregin V., et al., “Exploration Experiment on Enhanced Compression Beyond VVC Capability”, JVET-U2024-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 21st Meeting, by teleconference, Ja… [cited by applicant]
Wang Y., et al.,“EE2-Related: Template-Based Intra Mode Derivation Using MPMs”, JVET-V0098, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 22nd Meeting, by teleconference, Apr. 20-28, 2021,… [cited by applicant]
Zhao X., et al., “Six Tap Intra Interpolation Filter,” JVET Meeting, (The Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11), 4th Meeting, Chengdu, CN, Oct. 15-21, 2016, No. JVET-D011… [cited by applicant]
Chen J., et al., “Algorithm Description for Versatile Video Coding and Test Model 11 (VTM 11)”, JVET-T2002-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 20th Meeting, by teleconference… [cited by applicant]
International Search Report and Written Opinion—PCT/US2022/071669—ISA/EPO—Aug. 10, 2022, 17 Pages. [cited by applicant]
Abdoli M., et al., “Non-CE3: Decoder-side Intra Mode Derivation with Prediction Fusion Using Planar”, JVET-O0449-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 1115th Meeting: Gothenb… [cited by applicant]