IP Library Granted Patent US 12,335,524
Granted Patent B2
US 12,335,524 · App. 18/416,187 · Granted Jun 17, 2025

Frequency-dependent joint component secondary transform

Inventors: Xin Zhao (San Diego, CA); Madhu Peringassery Krishnan (Mountain View, CA); Shan Liu (San Jose, CA)
Assignee: TENCENT AMERICA LLC
H04N19/61H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,335,524
App. No.
18/416,187
Granted
Jun 17, 2025
Kind
B2
Abstract

A method and apparatus for performing a frequency-dependent joint component secondary transform (FD-JCST). The method includes obtaining a plurality of transform coefficients in a transform coefficient block; determining whether at least one of the plurality of transform coefficients is a low-frequency coefficient; based on determining that the at least one of the plurality of transform coefficients is the low-frequency coefficient, determining whether the low-frequency coefficient is a non-zero value; and based on determining that the low-frequency coefficient is the non-zero value, performing a joint component secondary transform (JCST) on the low-frequency coefficient and signaling a related syntax to indicate that the JCST is performed.

Claims (45)

1. A method of processing visual media data based on performing a joint component secondary transform (JCST), the method comprising performing a conversion between a visual media file and a bitstream of a visual media data according to a format rule, the format rule indicating:

obtaining residual components among a plurality of transform coefficients;

performing a transformation on the residual components; and

performing the JCST on the transformed residual components.

2. The method of claim 1 , wherein the residual components comprise a residual component of 0 and a residual component of 1 that are Cb and Cr transform coefficients, respectively.

3. The method of claim 2 , wherein the performing the JCST comprises performing the JCST element-wise on the residual component of 0 and the residual component of 1.

4. The method of claim 1 , further comprising:

determining whether at least one of the plurality of transform coefficients is a low-frequency coefficient;

based on determining that the at least one of the plurality of transform coefficients is the low-frequency coefficient, determining whether the low-frequency coefficient has a non-zero value; and

based on determining that the low-frequency coefficient has the non-zero value, performing the JCST on the low-frequency coefficient.

5. The method of claim 4 , wherein the determining whether the at least one of the plurality of transform coefficients is the low-frequency coefficient comprises determining whether the at least one of the plurality of transform coefficients is the low-frequency coefficient based on a coordinate (x, y) of a transform coefficient block including the plurality of transform coefficients.

6. The method of claim 1 , further comprising signaling a related syntax to indicate that the JCST is performed.

7. The method of claim 6 , wherein the related syntax comprises a high level syntax (HLS), and

wherein the signaling the related syntax further comprises signaling a transform kernel that is used in performing the JCST at the HLS.

8. The method of claim 7 , wherein the HLS indicates at least one of the transform kernel used for each frequency of the plurality of transform coefficients, the transform kernel used for each prediction mode, or the transform kernel used for each transform type.

9. The method of claim 4 , wherein an output of the JCST is divided by a factor of N, where N is a power of 2, or

wherein the output of the JCST is clipped in a predetermined data range.

10. An apparatus for processing visual media data based on performing a joint component secondary transform (JCST) comprising:

at least one memory storing computer program code; and

at least one processor configured to access the at least one memory and operate as instructed by the computer program code, and the at least one processor is configured to perform a conversion between a visual media file and a bitstream of a visual media data according to a format rule, the format rule indicating:

obtain residual components among a plurality of transform coefficients;

perform a transformation on the residual components; and

perform the JCST on the transformed residual components.

11. The apparatus of claim 10 , wherein the residual components comprise a residual component of 0 and a residual component of 1 that are Ch and Cr transform coefficients, respectively.

12. The apparatus of claim 11 , wherein the at least one processor is further configured to perform the JCST element-wise on the residual component of 0 and the residual component of 1.

13. The apparatus of claim 10 , wherein the at least one processor is further configured to:

determine whether at least one of the plurality of transform coefficients is a low-frequency coefficient;

based on determining that the at least one of the plurality of transform coefficients is the low-frequency coefficient, determine whether the low-frequency coefficient has a non-zero value; and

based on determining that the low-frequency coefficient has the non-zero value, perform the JCST on the low-frequency coefficient.

14. The apparatus of claim 13 , wherein the at least one processor is further configured to:

determine whether the at least one of the plurality of transform coefficients is the low-frequency coefficient based on a coordinate (x, y) of a transform coefficient block including the plurality of transform coefficients.

15. The apparatus of claim 10 , wherein the at least one processor is further configured to signal a related syntax to indicate that the JCST is performed.

16. The apparatus of claim 15 , wherein the related syntax comprises a high level syntax (HLS), and

wherein the at least one processor is further configured to signal a transform kernel that is used in performing the JCST at the HLS.

17. The apparatus of claim 16 , wherein the HLS indicates at least one of the transform kernel used for each frequency of the plurality of transform coefficients, the transform kernel used for each prediction mode, or the transform kernel used for each transform type.

18. The apparatus of claim 13 , wherein an output of the JCST is divided by a factor of N, where N is a power of 2, or

wherein the output of the JCST is clipped in a predetermined data range.

19. A non-transitory computer-readable recording medium storing computer program code for processing visual media data based on performing a joint component secondary transform (JCST), the computer program code when executed by at least one processor, causes the at least one processor to perform a conversion between a visual media file and a bitstream of a visual media data according to a format rule, the format rule indicating to:

obtain residual components among a plurality of transform coefficients;

perform a transformation on the residual components; and

perform the JCST on the transformed residual components.

20. The non-transitory computer-readable recording medium of claim 19 , wherein the at least one processor is further configured to:

determine whether at least one of the plurality of transform coefficients is a low-frequency coefficient;

based on determining that the at least one of the plurality of transform coefficients is the low-frequency coefficient, determine whether the low-frequency coefficient has a non-zero value; and

based on determining that the low-frequency coefficient has the non-zero value, perform the JCST on the low-frequency coefficient.

Continuity (3)
Continuation 17495202 · Oct 6, 2021
Continuation 16928760 · Jul 14, 2020
Related Publication 20240155159A1 · May 9, 2024
References Cited (51)
US 8238442B2 · Liu · 2012 [cited by applicant]
US 8526495B2 · Liu et al. · 2013 [cited by applicant]
US 9049452B2 · Liu et al. · 2015 [cited by applicant]
US 9788019B2 · Liu et al. · 2017 [cited by applicant]
US 10404980B1 · Zhao et al. · 2019 [cited by applicant]
US 10419754B1 · Zhao et al. · 2019 [cited by applicant]
US 10432929B2 · Zhao et al. · 2019 [cited by applicant]
US 10462486B1 · Zhao et al. · 2019 [cited by applicant]
US 10491893B1 · Zhao et al. · 2019 [cited by applicant]
US 10536720B2 · Zhao et al. · 2020 [cited by applicant]
US 10547854B2 · Seregin et al. · 2020 [cited by applicant]
US 10567801B2 · Zhao et al. · 2020 [cited by applicant]
US 10609384B2 · Chen et al. · 2020 [cited by applicant]
US 10609402B2 · Zhao et al. · 2020 [cited by applicant]
US 20050078869A1 · Kim · 2005 [cited by applicant]
US 20130182758A1 · Seregin et al. · 2013 [cited by applicant]
US 20140198857A1 · Deshpande · 2014 [cited by applicant]
US 20170094314A1 · Zhao et al. · 2017 [cited by applicant]
US 20180302631A1 · Chiang et al. · 2018 [cited by applicant]
US 20190373276A1 · Hu et al. · 2019 [cited by applicant]
US 20200396487A1 · Nalci et al. · 2020 [cited by applicant]
US 20210120252A1 · Koo et al. · 2021 [cited by applicant]
B. Bross et al., “General Video Coding Technology in Responses to the Joint Call for Proposals on Video Compression with Capability beyond HEVC,” in IEEE Transactions on Circuits and Systems for Video Technology, 2019, … [cited by applicant]
Benjamin Bross et al., “Versatile Video Coding (Draft 6)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 15th Meeting: Gothenburg, SE, Jul. 3-12, 2019, JVET-O2001-VE (455 pages). [cited by applicant]
Bross et al., “CE3: Multiple reference line intra prediction (Test 1.1.1, 1.1.2, 1.1.3 and 1.1.4)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting: Macao, CN, Oct. 3-12, … [cited by applicant]
Bross et al., “Versatile Video Coding (Draft 2)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 11th Meeting: Ljubljana, SI, Jul. 10-18, 2018, JVET-K1001-v6 (141 pages total). [cited by applicant]
Christian Rudat et al., “Inter-Component Transform for Color Video Coding”, 2019 Picture Coding Symposium (PCS) ,IEEE, Nov. 12-15, 2019, pp. 1-5 (5 pages total). [cited by applicant]
De Rivaz et al., “AV1 Bitstream & Decoding Process Specification”, Version 1.0.0 with Errata 1, 2018, The Alliance for Open Media, https://aomediacodec.github.io/av1-spec/av1-spec.pdf (681 pages total). [cited by applicant]
Dong Liu et al., “Deep Learning-Based Technology in Responses to the Joint Call for Proposals on Video Compression with Capability beyond HEVC”, IEEE Transactions on Circuits and Systems for Video Technology, 2019, pp. … [cited by applicant]
Extended European Search Report dated Apr. 24, 2023 in European Application No. 21842772.2. [cited by applicant]
International Search Report dated Aug. 24, 2021 in International Application No. PCT/US2021/031944. [cited by applicant]
Jianle Chen et al., “Algorithm description for Versatile Video Coding and Test Model 9 (VTM 9)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-R2002-v2, Apr. 15-24, 2020, 18th M… [cited by applicant]
L. Zhao et al., “Wide Angular Intra Prediction for Versatile Video Coding,” 2019 Data Compression Conference (DCC), Snowbird, UT, USA, 2019, pp. 53-62 (10 pages total). [cited by applicant]
Racape et al., “CE3-related: Wide-angle intra prediction for non-square blocks”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 11th Meeting: Ljubljana, SI, Jul. 10-18, 2018, JVET-K05… [cited by applicant]
S. Liu et al., “Joint temporal-spatial bit allocation for video coding with dependency,” in IEEE Transactions on Circuits and Systems for Video Technology, vol. 15, No. 1, pp. 15-26, Jan. 2005 (12 pages total). [cited by applicant]
Shan Liu et al., “Bit-depth Scalable Coding for High Dynamic Range Video”, SPIE-IS&T Electronic Imaging, 2008, vol. 6822, pp. 1-10 (10 pages total). [cited by applicant]
Shan Liu et al., “JVET AHG Report: Neural Networks in Video Coding (AHG9)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 10th Meeting: San Diego, US, Apr. 10-20, 2018, JVET-J0009-v1… [cited by applicant]
Written Opinion of the International Searching Authority dated Aug. 24, 2021 in International Application No. PCT/US2021/031944. [cited by applicant]
X. Zhao et al., “Coupled Primary and Secondary Transform for Next Generation Video Coding,” 2018 IEEE Visual Communications and Image Processing (VCIP), Taichung, Taiwan, 2018, pp. 1-4 (4 pages total). [cited by applicant]
X. Zhao et al., “Low-Complexity Intra Prediction Refinements for Video Coding,” 2018 Picture Coding Symposium (PCS), San Francisco, CA, 2018, pp. 139-143 (5 pages total). [cited by applicant]
X. Zhao et al., “Joint separable and non-separable transforms for next-generation video coding,” IEEE Transactions on Image Processing, vol. 27, No. 5, pp. 2514-2525, May 2018 (13 pages total). [cited by applicant]
X. Zhao et al., “Novel Statistical Modeling, Analysis and Implementation of Rate-Distortion Estimation for H.264/AVC Coders,” in IEEE Transactions on Circuits and Systems for Video Technology, vol. 20, No. 5, pp. 647-66… [cited by applicant]
X. Zhao et al., “NSST: Non-Separable Secondary Transforms for Next Generation Video Coding,” in Proc. Picture Coding Symposium, 2016 (5 pages total). [cited by applicant]
Y.-J. Chang et al., “Intra prediction using multiple reference lines for the versatile video coding standard,” Proc. SPIE 11137, Applications of Digital Image Processing XLII, 1113716, Sep. 2019 (8 pages total). [cited by applicant]
Yiming Li et al., “Methodology and reporting template for neural network coding tool testing”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting: Marrakech, MA, Jan. 9-18, 2… [cited by applicant]
Z. Zhang et al., “Fast Adaptive Multiple Transform for Versatile Video Coding,” 2019 Data Compression Conference (DCC), Snowbird, UT, USA, 2019, pp. 63-72 (10 pages total). [cited by applicant]
Z. Zhang et al., “Fast DST-7/DCT-8 with Dual Implementation Support for Versatile Video Coding,” in IEEE Transactions on Circuits and Systems for Video Technology, 2020, pp. 1-17 (17 pages total). [cited by applicant]
Zhao et al., “CE6: Fast DST-7/DCT-8 with dual implementation support (Test 6.2.3)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting: Marrakech, MA, Jan. 9-18, 2019, JVET-M… [cited by applicant]
Zhao et al., “CE6: On 8-bit primary transform core (Test 6.1.3)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting: Macao, CN, Oct. 3-12, 2018, JVET-L0285 (35 pages total). [cited by applicant]
Zhao et al., “CE6-related: Unified LFNST using block size independent kernel”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 15th Meeting: Gothenburg, SE, Jul. 3-12, 2019, JVET-00539… [cited by applicant]
Zhao et al., “Non-CE6: Configurable max transform size in VVC”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 15th Meeting: Gothenburg, SE, Jul. 3-12, 2019, JVET-O0545-v2 (6 pages to… [cited by applicant]