IP Library Granted Patent US 12,445,651
Granted Patent B2
US 12,445,651 · App. 17/956,752 · Granted Oct 14, 2025

Residual sign prediction of transform coefficients in video coding

Inventors: Jie Chen (Beijing, CN); Mohammad Golam Sarwer (Cupertino, CA); Yan Ye (San Diego, CA); Ru-Ling Liao (Beijing, CN); Xinwei Li (Beijing, CN)
Assignee: Alibaba Innovation Private Limited
H04N19/70H04N19/176H04N19/18H04N19/88
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,445,651
App. No.
17/956,752
Granted
Oct 14, 2025
Kind
B2
Abstract

A VVC-standard encoder and a VVC-standard decoder are provided, implementing a residual sign prediction method utilizing a sorting order of residual coefficients and an expanded region of a TB. As sign prediction accuracy is higher for larger transform coefficient levels, a VVC-standard encoder and a VVC-standard decoder sort transform coefficient signs of a TB in a one-dimensional array, based on corresponding QIdx values instead of the residual coefficient level value. The first n signs according to corresponding QIdx values, ordered from largest to smallest, are predicted using a residual sign prediction method, and the rest of the signs are signaled by EP bins. A sign prediction area is also extended, without limitation to an upper-left 4×4 region within the transform block, but to a region up to 32×32 in size; a VVC-standard encoder signals the maximum dimensions of the region to a VVC-standard decoder in syntax structures of the block.

Claims (56)

1. A method for encoding a video sequence, the method comprising:

receiving a video sequence;

encoding the video sequence by:

predicting a plurality of transform coefficient signs of a current block corresponding to largest respective quantization indices of the transform coefficients of the current block in sorted order, wherein the plurality of transform coefficient signs are predicted up to a maximum allowable limit;

encoding residues of the predicted plurality of transform coefficient signs and the maximum allowable limit in a syntax structure of the current block.

2. The method of claim 1 , further comprising:

bypassing coding of transform coefficient signs of the current block other than the predicted plurality of transform coefficient signs; and

outputting a coded block comprising the encoded residues of predicted plurality of transform coefficient signs and equiprobable binary string signaling of coding-bypassed transform coefficient signs.

3. The method of claim 2 , further comprising:

encoding a residue for each predicted transform coefficient sign in a bitstream before signaling coding-bypassed transform coefficient signs; and

encoding coding-bypassed transform coefficient signs in the bitstream trailing the encoded residue.

4. The method of claim 1 , further comprising:

recording the quantization indices of the transform coefficients of the current block as elements of an array in raster scanning order; and

sorting the elements of the array by size to yield a sorted array.

5. The method of claim 4 , wherein predicting a plurality of transform coefficient signs of the current block corresponding to largest quantization indices comprises predicting each transform coefficient corresponding to quantization indices recorded in a plurality of elements at an end of the sorted array.

6. The method of claim 1 , wherein:

the predicted plurality of transform coefficient signs are limited to a region of the current block;

a width of the region of the current block is a lesser value among a fixed width and a width of the current block; and

a height of the region of the current block is a lesser value among a fixed height and a height of the current block.

7. The method of claim 6 , further comprising:

encoding the fixed width and the fixed height in the syntax structure of the current block.

8. The method of claim 1 , wherein

the maximum allowable limit is greater than 8.

9. A method for decoding a bitstream, the method comprising:

receiving a bitstream; and

decoding the bitstream to output a video sequence, the decoding comprising:

decoding a plurality of entropy-coded transform coefficients of a current block to yield decoded transform coefficients of the current block;

decoding a maximum allowable limit of a predicted plurality of transform coefficient signs; and

reconstructing the current block based on residues of the predicted plurality of transform coefficient signs of reconstructed residuals of the current block corresponding to largest respective quantization indices of the transform coefficients of reconstructed residuals of the current block in sorted order; wherein the plurality of transform coefficient signs are predicted up to the maximum allowable limit.

10. The method of claim 9 , further comprising:

parsing, by the entropy decoder, a plurality of residues of predicted transform coefficient signs from a syntax structure of the current block; and

inverse quantizing and inverse transforming the decoded transform coefficients and the plurality of residues of transform coefficient signs to yield reconstructed residuals of the current block.

11. The method of claim 10 , further comprising:

recording quantization indices of transform coefficients of reconstructed residuals of the current block as elements of an array in raster scanning order; and

sorting the elements of the array by size to yield a sorted array.

12. The method of claim 11 , wherein yielding reconstructed residuals of the current block comprises selecting a plurality of transform coefficient signs of reconstructed residuals of the current block corresponding to largest quantization indices comprises predicting each transform coefficient corresponding to quantization indices recorded in a plurality of elements at an end of the sorted array.

13. The method of claim 11 , further comprising adding the reconstructed residuals of the current block and a prediction signal comprising the plurality of transform coefficient signs to yield a reconstructed block.

14. The method of claim 9 , wherein:

the predicted plurality of transform coefficient signs are limited to a region of the current block;

a width of the region of the current block is a lesser value among a fixed width and a width of the current block; and

a height of the region of the current block is a lesser value among a fixed height and a height of the current block.

15. The method of claim 14 , further comprising: parsing the fixed width and the fixed height from a syntax structure of the current block.

16. A non-transitory computer-readable storage medium storing a bitstream, the bitstream generated by receiving a video sequence, encoding the video sequence to generate coded information included in the bitstream, and transmit the bitstream, wherein the encoding comprises:

predicting a plurality of transform coefficient signs of a current block corresponding to largest respective quantization indices of the transform coefficients of the current block in sorted order, wherein the plurality of transform coefficient signs are predicted up to a maximum allowable limit; and

encoding residues of the predicted plurality of transform coefficient signs and the maximum allowable limit in a syntax structure of the current block.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the encoding further comprises:

bypassing coding of transform coefficient signs of the current block other than the predicted plurality of transform coefficient signs; and

outputting a coded block comprising the encoded residues of predicted plurality of transform coefficient signs and equiprobable binary string signaling of coding-bypassed transform coefficient signs.

18. The non-transitory computer-readable storage medium of claim 16 , wherein the encoding further comprises:

recording the quantization indices of the transform coefficients of the current block as elements of an array in raster scanning order; and

sorting the elements of the array by size to yield a sorted array.

19. The non-transitory computer-readable storage medium of claim 16 , wherein the predicted plurality of transform coefficient signs are limited to a region of the current block;

a width of the region of the current block is a lesser value among a fixed width and a width of the current block; and

a height of the region of the current block is a lesser value among a fixed height and a height of the current block.

20. The non-transitory computer-readable storage medium of claim 19 , wherein the encoding further comprising:

encoding the fixed width and the fixed height in the syntax structure of the current block.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2024
From: ALIBABA SINGAPORE HOLDING PRIVATE LIMITED
To: ALIBABA INNOVATION PRIVATE LIMITED
Reel/Frame 066348/0252 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2023
From: SARWER, MOHAMMED GOLAM; CHEN, JIE; LI, XINWEI; LIAO, RU-LING; YE, YAN
To: ALIBABA SINGAPORE HOLDING PRIVATE LIMITED
Reel/Frame 063178/0039 →
Continuity (3)
Provisional Application 63296370 · Jan 4, 2022
Provisional Application 63250202 · Sep 29, 2021
Related Publication 20230093994A1 · Mar 30, 2023
References Cited (9)
US 11856216B2 · Filippov · 2023 [cited by examiner]
US 20240298009A1 · Xiu · 2024 [cited by examiner]
EP 4503608A1 · 2025 [cited by examiner]
Coban et al., Algorithm description of Enhanced Compression Model 2 (ECM 2), JVET-W2025, 23 [cited by applicant]
ECM, https://vogit.hhi.fraunhofer.de/ecm/ECM. [cited by applicant]
Henry et al., “Residual Coefficient Sign Prediction,” JVET-D0031, 4 [cited by applicant]
International Telecommunications Union “Series H: Audiovisual and Multimedia Systems Infrastructure of audiovisual services—Coding of moving video”, ITU-T Telecommunication Standardization Sector of ITU, Apr. 2013, 317 … [cited by applicant]
Schwarz et al., “EE2-4.1: Results for dependent quantization with 8 states,” JVET-V0082-v1, 22nd Meeting, by teleconference, Apr. 20-28, 2021, 6 pages. [cited by applicant]
Sullivan et al., “Overview of the High Efficiency Video Coding (HEVC) Standard,” IEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, pp. 1649-1668 (2012). [cited by applicant]