IP Library › Granted Patent US 12,256,092
Granted Patent B2
US 12,256,092 · App. 18/192,343 · Granted Mar 18, 2025

Multiple transforms adjustment stages for video coding

Inventors: Amir Said (San Diego, CA); Hilmi Enes Egilmez (San Diego, CA); Marta Karczewicz (San Diego, CA); Vadim Seregin (San Diego, CA)
Assignee: QUALCOMM INCORPORATED
H04N19/45G06F17/16H04N19/117H04N19/12H04N19/147H04N19/176H04N19/18H04N19/463H04N19/61H04N19/625H04N19/96
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,256,092
App. No.
18/192,343
Granted
Mar 18, 2025
Kind
B2
Abstract

A device may perform a first prediction process for a first block of video data to produce a first residual. The device may apply a first transform process to the first residual to generate first transform coefficients for the first block of video data and encode the first transform coefficients. The device may perform a second prediction process for a second block of video data to produce a second residual. The device may determine that a second transform process, which includes the first transform process and at least one of a pre-adjustment operation or a post-adjustment operation, is to be applied to the second residual. The device may apply the first transform process and the pre- or post-adjustment operation to the second residual to generate second transform coefficients for the second block. The coding device may code the first and second transform coefficients.

Claims (72)

1. A method of encoding video data comprising:

performing a first prediction process for a first block of video data to produce a first residual;

determining that a first transform process of a plurality of transform processes is to be applied to the first residual, wherein the first inverse transform process comprises one of a DCT-2 matrix, a DCT-3 matrix, a DST-2 matrix, or a DST-3 matrix;

applying the first transform process to the first residual to generate first transform coefficients for the first block of video data;

encoding the first transform coefficients;

performing a second prediction process for a second block of video data to produce a second residual;

determining that a second transform process is to be applied to the second residual, wherein the second transform process comprises the first transform process and at least one of a first pre-adjustment operation or a first post-adjustment operation to apply to the second residual in addition to the first transform process;

applying the first transform process and at least one of the first pre-adjustment operation or the first post-adjustment operation to the second residual to generate second transform coefficients for the second block of video data, wherein the first pre-adjustment operation, if applied, is applied prior to applying the first transform process, and wherein the first post- adjustment operation, if applied, is applied after applying the first transform process, and wherein the application of the first inverse transform process and the at least one of the pre- adjustment operation or the post-adjustment operation approximates an inverse transform process of a type different from the first inverse transform process; and

encoding the second transform coefficients.

2. The method of claim 1 , further comprising:

performing a third prediction process on the third block of video data to produce a third residual;

determining a subset of transform processes from the plurality of transform processes, wherein the subset of transform processes includes fewer transform processes than the plurality of transform processes, and wherein the subset of transform processes includes the first transform process;

determining a set of adjustment operations, wherein each adjustment operation comprises at least one of a pre-adjustment operation or a post-adjustment operation to be applied to the third residual, and wherein each adjustment operation, when applied in conjunction with a transform process of the subset of transform processes, results in transform coefficients that are approximately equal to transform coefficients resulting from the application of a transform process in the plurality of transform processes that is not included in the subset of transform processes, and wherein each adjustment operation of the set of adjustment operations is associated with a transform process of the subset of transform processes;

determining, for each transform process of the subset of transform processes, rate- distortion characteristics for the respective transform process;

determining, for each adjustment operation of the set of adjustment operations, rate- distortion characteristics for the respective adjustment operation and the associated transform process for the respective adjustment operation;

selecting, based on the determined rate-distortion characteristics, either the transform process from the subset of transform processes or the adjustment operation from the set of adjustment operations and the transform process from the subset of transform processes associated with the selected adjustment operation as a complete transform process to apply to the third residual;

applying the complete transform process to the third residual to generate third transform coefficients for the third block of video data; and

encoding the third transform coefficients.

3. The method of claim 2 , wherein the subset of transform processes comprises one or more of a DCT-2 matrix, a DCT-3 matrix, a DST-2 matrix, or a DST-3 matrix.

4. The method of claim 2 , further comprising:

encoding, in a bitstream, one or more indexes indicating the selected transform process from the subset of transform processes; and

if an adjustment operation is selected, encoding, in the bitstream, one or more indexes indicating the selected adjustment operation from the set of adjustment operations.

5. The method of claim 1 , wherein encoding the first transform coefficients and encoding the second transform coefficients comprises entropy encoding the first transform coefficients and entropy encoding the second transform coefficients.

6. The method of claim 5 , wherein each transform process of the plurality of transform processes comprises a discrete trigonometrical transform matrix.

7. The method of claim 1 , wherein the first transform process comprises a discrete trigonometrical transform matrix.

8. The method of claim 1 , wherein the first transform process and the at least one of the first pre-adjustment operation or the first post-adjustment operation each comprise a respective sparse matrix.

9. The method of claim 8 , wherein the sparse matrix of the first transform process comprises a band diagonal matrix.

10. The method of claim 8 , wherein the sparse matrix of the first transform process comprises a block diagonal matrix.

11. The method of claim 1 , wherein the at least one of the first pre-adjustment operation or the first post-adjustment operation comprises a set of one or more Givens rotations.

12. The method of claim 1 , wherein the first block of video data comprises a first coding unit in a first row of a coding tree unit, wherein the second block of video data comprises a second coding unit in a second row of the coding tree unit, and wherein the first row is different than the second row.

13. The method of claim 1 , wherein the first block of video data comprises a first coding unit in a first column of a coding tree unit, wherein the second block of video data comprises a second coding unit in a second column of the coding tree unit, and wherein the first column is different than the second column.

14. The method of claim 1 , wherein the first block of video data comprises a first coding unit in a first row of a coding tree unit, wherein the second block of video data comprises a second coding unit in a second row of the coding tree unit, wherein the first row is different than the second row, and wherein the method further comprises:

performing a third prediction process for a third coding unit of the coding tree to produce a third residual, wherein the third coding unit is in a third row of the coding tree unit, wherein the second row and the third row are contiguous rows;

determining that the second transform process is to be applied to the third residual;

applying the first transform process and at least one of the first pre-adjustment operation or the first post-adjustment operation to the third residual to generate third transform coefficients for the third coding unit; and

encoding the third transform coefficients.

15. The method of claim 1 , wherein the first block of video data comprises a first coding tree unit, and wherein the second block of video data comprises a second coding tree unit different than the first coding tree unit.

16. A video encoding device comprising:

a memory configured to store video data; and

one or more processors implemented in circuitry and configured to:

perform a first prediction process for a first block of video data to produce a first residual;

determine that a first transform process of a plurality of transform processes is to be applied to the first residual, wherein the first inverse transform process comprises one of a DCT-2 matrix, a DCT-3 matrix, a DST-2 matrix, or a DST-3 matrix;;

apply the first transform process to the first residual to generate first transform coefficients for the first block of video data;

encode the first transform coefficients;

determine that a second transform process is to be applied to the second residual, wherein the second transform process comprises the first transform process and at least one of a first pre-adjustment operation or a first post-adjustment operation to apply to the second residual in addition to the first transform process;

apply the first transform process and at least one of the first pre-adjustment operation or the first post-adjustment operation to the second residual to generate second transform coefficients for the second block of video data, wherein the first pre-adjustment operation, if applied, is applied prior to applying the first transform process, and wherein the first post-adjustment operation, if applied, is applied after applying the first transform process and wherein the application of the first inverse transform process and the at least one of the pre-adjustment operation or the post-adjustment operation approximates an inverse transform process of a type different from the first inverse transform process; and

encode the second transform coefficients.

17. The video encoding device of claim 16 , wherein the one or more processors are further configured to:

perform a third prediction process on the third block of video data to produce a third residual;

determine a subset of transform processes from the plurality of transform processes, wherein the subset of transform processes includes fewer transform processes than the plurality of transform processes, and wherein the subset of transform processes includes the first transform process;

determine a set of adjustment operations, wherein each adjustment operation comprises at least one of a pre-adjustment operation or a post-adjustment operation to be applied to the third residual, and wherein each adjustment operation, when applied in conjunction with a transform process of the subset of transform processes, results in transform coefficients that are approximately equal to transform coefficients resulting from the application of a transform process in the plurality of transform processes that is not included in the subset of transform processes, and wherein each adjustment operation of the set of adjustment operations is associated with a transform process of the subset of transform processes;

determine, for each transform process of the subset of transform processes, rate-distortion characteristics for the respective transform process;

determine, for each adjustment operation of the set of adjustment operations, rate- distortion characteristics for the respective adjustment operation and the associated transform process for the respective adjustment operation;

select, based on the determined rate-distortion characteristics, either the transform process from the subset of transform processes or the adjustment operation from the set of adjustment operations and the transform process from the subset of transform processes associated with the selected adjustment operation as a complete transform process to apply to the third residual;

apply the complete transform process to the third residual to generate third transform coefficients for the third block of video data; and encode the third transform coefficients.

18. The video encoding device of claim 17 , wherein the subset of transform processes comprises one or more of a DCT-2 matrix, a DCT-3 matrix, a DST-2 matrix, or a DST-3 matrix, wherein the one or more processors are further configured to:

encode, in a bitstream, one or more indexes indicating the selected transform process from the subset of transform processes; and

if an adjustment operation is selected, encode, in the bitstream, one or more indexes indicating the selected adjustment operation from the set of adjustment operations.

19. The video encoding device of claim 16 , wherein encoding the first transform coefficients and encoding the second transform coefficients comprises entropy encoding the first transform coefficients and entropy encoding the second transform coefficients.

20. The video encoding device of claim 16 , wherein the first transform process comprises a discrete trigonometrical transform matrix, and wherein each transform process of the plurality of transform processes comprises a discrete trigonometrical transform matrix.

21. The video encoding device of claim 16 , wherein the first transform process and the at least one of the first pre-adjustment operation or the first post-adjustment operation each comprise a respective sparse matrix, wherein the sparse matrix of the first transform process comprises one of a band diagonal matrix or a block diagonal matrix.

22. The video encoding device of claim 16 , wherein the at least one of the first pre-adjustment operation or the first post-adjustment operation comprise a set of one or more Givens rotations.

23. The video encoding device of claim 16 , wherein the first block of video data comprises a first coding unit in a first row of a coding tree unit, wherein the second block of video data comprises a second coding unit in a second row of the coding tree unit, and wherein the first row is different than the second row.

24. The video encoding device of claim 16 , wherein the first block of video data comprises a first coding unit in a first column of a coding tree unit, wherein the second block of video data comprises a second coding unit in a second column of the coding tree unit, and wherein the first column is different than the second column.

25. The video encoding device of claim 16 , wherein the first block of video data comprises a first coding unit in a first row of a coding tree unit, wherein the second block of video data comprises a second coding unit in a second row of the coding tree unit, wherein the first row is different than the second row, and wherein the one or more processors are further configured to:

predict a third residual for a third coding unit of the coding tree unit, wherein the third coding unit is in a third row of the coding tree unit, wherein the second row and the third row are contiguous rows;

determine that the second transform process is to be applied to the third residual;

apply the first transform process and at least one of the first pre-adjustment operation or the first post-adjustment operation to the third residual to generate third transform coefficients for the third coding unit; and

encode the third transform coefficients.

26. The video encoding device of claim 16 , wherein the first block of video data comprises a first coding tree unit, and wherein the second block of video data comprises a second coding tree unit different than the first coding tree unit.

27. The video encoding device of claim 16 , further comprising:

a camera configured to capture the video data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2023
From: SAID, AMIR; EGILMEZ, HILMI ENES; KARCZEWICZ, MARTA; SEREGIN, VADIM
To: QUALCOMM INCORPORATED
Reel/Frame 063159/0613 →
Continuity (4)
Division 16368455 · Mar 28, 2019
Provisional Application 62668097 · May 7, 2018
Provisional Application 62650953 · Mar 30, 2018
Related Publication 20240244246A1 · Jul 18, 2024
References Cited (52)
US 5347305A · Bush et al. · 1994 [cited by applicant]
US 5867416A · Feldmann · 1999 [cited by examiner]
US 10194158B2 · Karczewicz et al. · 2019 [cited by applicant]
US 10531123B2 · Kim · 2020 [cited by examiner]
US 20030206582A1 · Srinivasan et al. · 2003 [cited by applicant]
US 20080243971A1 · Po et al. · 2008 [cited by applicant]
US 20100057822A1 · Manolescu · 2010 [cited by examiner]
US 20100312811A1 · Reznik · 2010 [cited by applicant]
US 20100329352A1 · DeCegama · 2010 [cited by examiner]
US 20110268183A1 · Sole · 2011 [cited by examiner]
US 20120121167A1 · Atoyan · 2012 [cited by examiner]
US 20140079135A1 · Van Der Auwera et al. · 2014 [cited by applicant]
US 20140254661A1 · Saxena · 2014 [cited by examiner]
US 20140294086A1 · Oh et al. · 2014 [cited by applicant]
US 20150016542A1 · Rosewarne · 2015 [cited by examiner]
US 20160171641A1 · Jeon · 2016 [cited by examiner]
US 20160255371A1 · Heo et al. · 2016 [cited by applicant]
US 20170094313A1 · Zhao et al. · 2017 [cited by applicant]
US 20170134732A1 · Chen · 2017 [cited by examiner]
US 20170155906A1 · Puri · 2017 [cited by examiner]
US 20170238013A1 · Said et al. · 2017 [cited by applicant]
US 20180288439A1 · Hsu · 2018 [cited by examiner]
US 20190306522A1 · Said et al. · 2019 [cited by applicant]
US 20200186835A1 · Said · 2020 [cited by applicant]
WO 2010087807A1 · 2010 [cited by applicant]
WO 2014039398 · 2014 [cited by applicant]
Biatek T., et al., “Low-Complexity Adaptive Multiple Transforms for post-HEVC Video Coding”, 2016 Picture Coding Symposium (PCS), IEEE, Dec. 4, 2016, XP033086858, DOI: 10.1109/PCS.2016.7906348 [retrieved on Apr. 19, 201… [cited by applicant]
Bossen F., et al., “JEM Software Manual”, Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, Document: JCTVC-Software Manual, Retrieved on Aug. 3, 2016, pp. 1-29. [cited by applicant]
Britanak V., et al., “Discrete Cosine and Sine Transforms: General Properties, Fast Algorithms and Integer Approximations”, Acadmemic Press, 2007, pp. 16-38. [cited by applicant]
Bross B., et al., “Versatile Video Coding (Draft 1)”, JVET-J1001-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 10th Meeting, Apr. 2018, JVET-J1001-v2, 43 pages. [cited by applicant]
Chen H., et al., “New Transforms Tightly Bounded by DCT and KLT”, IEEE Signal Processing Letters, IEEE Service Center, Piscataway, NJ, US, vol. 19, No. 6, Jun. 1, 2012 (Jun. 1, 2012), pp. 344-347, XP011442399, ISSN: 107… [cited by applicant]
Chen J., et al., “Algorithm Description for Versatile Video Coding and Test Model 1 (VTM 1)”, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 10th Meeting: San Diego, US, Apr. 10-2… [cited by applicant]
Chen J., et al., “Algorithm Description of Joint Exploration Test Model 1”, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 1st Meeting: Geneva, CH, Oct. 19-21, 2015, No. H.266, JV… [cited by applicant]
Egilmez H.E., et al., “Row-Column Transforms: Low-complexity Approximation of Optimal Non-Separable Transform”, 2016 IEEE International Conference on Image Processing (ICIP), Sep. 2016, pp. 2385-2389. [cited by applicant]
Guerreiro R.F.C, et al., “Maximizing Compression Efficiency Through Block Rotation”, Nov. 16, 2014 (Nov. 16, 2014), 4 Pages, XP055284160, Retrieved from the Internet: URL: http://arxiv.org/pdf/1411.4290v1.pdf the whole … [cited by applicant]
Han J., et al., “Jointly Optimized Spatial Prediction and Block Transform for Video and Image Coding”, IEEE Transactions on Image Processing, Apr. 2012, vol. 21, No. 4, pp. 1874-1884. [cited by applicant]
International Preliminary Report on Patentability—PCT/US2019/024899, The International Bureau of WIPO—Geneva, Switzerland, Oct. 15, 2020. [cited by applicant]
International Search Report and Written Opinion—PCT/US2019/024899—ISA/EPO—Jun. 18, 2019. [cited by applicant]
ITU-T H.233, Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Transmission multiplexing and synchronization; Multiplexing protocol for low bit rate multimedia communication, The Inter… [cited by applicant]
“ITU-T H.265, Series H: Audiovisual and Multimedia Systems, Infrastructure of Audiovisual Services—Coding of Moving Video, High Efficiency Video Coding”, The International Telecommunication Union, Apr. 2015, 634 Pages, … [cited by applicant]
Lorcy V., et al.,“CE6: Further simplification of AMT with adjustment stages (Test CE6.1.6b)”, JVET-L0135-v1, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Oct. 3-12, 2018, 8 pages. [cited by applicant]
Philippe P., “CE6: Mts simplification with TAF (tests 1.5a-d)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-M0080-v4, Jan. 9-18, 2019, 17 pages. [cited by applicant]
Philippe P., et al., “CE6-Related: Further Simplification for AMT Complexity Reduction (CE6.1.2)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-K0000-v1, Jul. 10-18, 2018, 14 p… [cited by applicant]
Said A., et al., “CE6.1.2: Efficient Implementations of AMT with Transform Adjustment Stages”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Jul. 10-18, 2018, JVET-K0272-v2, 6 pages. [cited by applicant]
Said A., et al., “CE6-1.4: Efficient Implementations of MTS with Transform Adjustments”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-M0538-v2, Jan. 9-18, 2019, 26 pages. [cited by applicant]
Said A., et al., “CE6.1.6: Efficient Implementations of AMT with Transform Adjustment Filters (TAF)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Oct. 3-12, 2018, JVET-L0386-v4, 8 … [cited by applicant]
Said A., et al., “CE6-Related: Efficient Computation of MTS Transform Combinations”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Jan. 9-18, 2019, JVET-M0539-v2, 6 pages. [cited by applicant]
Said A., et al., “Non-CE6: Efficient separable Multiple-Transform-Selection (MTS) without Zero-Out”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Mar. 19-27, 2019, JVET-N0485-v4, 14… [cited by applicant]
Said A., et al., “Complexity Reduction for Adaptive Multiple Transforms (AMTs) using Adjustment Stages”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 10th Meeting: San Diego, US, Ap… [cited by applicant]
Said (QUALCOMM) A., et al., “Description of Core Experiment 6 (CE6): Transforms and Transform Signalling”, 11. JVET Meeting; Jul. 11, 2018-Jul. 18, 2018; Ljubljana; (The Joint Video Exploration team of ISO/IEC JTC1/SC29… [cited by applicant]
Wien M., “High Efficiency Video Coding, Coding Tools and Specification”, Chapter 5, Springer-Verlag, Berlin, 2015, 30 Pages. [cited by applicant]
Zhao X., et al., “Enhanced Multiple Transform for Video Coding”, Data Compression Conference, Mar. 30, 2016, XP033027689, pp. 73-82, DOI: 10.1109/DCC.2016.9 [retrieved on Dec. 15, 2016]. [cited by applicant]