IP Library Granted Patent US 12,563,234
Granted Patent B2
US 12,563,234 · App. 18/606,849 · Granted Feb 24, 2026

Sign prediction for block-based video coding

Inventors: Xiaoyu Xiu (San Diego, CA); Ning Yan (San Diego, CA); Yi-Wen Chen (San Diego, CA); Che-Wei Kuo (San Diego, CA); Wei Chen (San Diego, CA); Hong-Jheng Jhu (San Diego, CA); Xianglin Wang (San Diego, CA); Bing Yu (Beijing, CN)
Assignee: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
H04N19/61H04N19/132H04N19/176H04N19/18H04N19/196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,563,234
App. No.
18/606,849
Filed
Mar 15, 2024
Granted
Feb 24, 2026
Kind
B2
Art Unit
2487
USPC
375/240.12
Abstract

Implementations of the disclosure provide a video decoding apparatus and method for transform coefficient sign prediction on a video decoder side. The method may include generating a plurality of candidate hypotheses for a set of candidate transform coefficients associated with a transform block of a video frame from a video. The method may further include selecting a hypothesis from the plurality of candidate hypotheses as a set of predicted signs for the set of candidate transform coefficients, wherein the hypothesis is selected based on a cost function calculated by extrapolating neighboring samples of the transform block in an extrapolation direction determined based on a dominant gradient direction. The method may also include estimating original signs for the set of candidate transform coefficients based on the set of predicted signs and a sequence of sign signaling bits received from a video encoder.

Claims (61)

1 . A video decoding method for transform coefficient sign prediction, comprising:

generating, by one or more processors, a plurality of candidate hypotheses for a set of candidate transform coefficients associated with a transform block of a video frame from a video;

selecting, by the one or more processor, a hypothesis from the plurality of candidate hypotheses as a set of predicted signs for the set of candidate transform coefficients, wherein the hypothesis is selected based on a cost function calculated by extrapolating neighboring samples of the transform block in an extrapolation direction, wherein the extrapolation direction is determined based on a dominant gradient direction, wherein the dominant gradient direction is a direction with the largest magnitude in a histogram of gradient obtained based on an accumulated magnitude of gradients for each of a plurality of angular directions; and

estimating, by the one or more processors, original signs for the set of candidate transform coefficients based on the set of predicted signs and a sequence of sign signaling bits received from a video encoder.

2 . The video decoding method of claim 1 , further comprising:

receiving a bit stream comprising the sequence of sign signaling bits and quantized transform coefficients associated with the transform block;

generating dequantized transform coefficients from the quantized transform coefficients; and

updating the dequantized transform coefficients based on the estimated original signs for the set of candidate transform coefficients.

3 . The video decoding method of claim 2 , further comprising:

applying an inverse primary transform and an inverse low-frequency non-separable transform (LFNST) to the dequantized transform coefficients to generate residual samples in a residual block corresponding to the transform block.

4 . The video decoding method of claim 1 , further comprising:

determining the dominant gradient direction for extrapolating neighboring samples of the transform block.

5 . The video decoding method of claim 4 , wherein determining the dominant gradient direction further comprises:

selecting templates of neighboring samples for the transform block;

applying a gradient filter window to the templates to calculate gradients for the respective templates; and

selecting the dominant gradient direction based on the calculated gradients.

6 . The video decoding method of claim 5 , wherein selecting the dominant gradient direction based on the calculated gradients further comprises:

determining the histogram of gradient based on the calculated gradients, wherein each entry of the histogram of gradient corresponds to an accumulated magnitude of the gradients at a predefined angular direction; and

selecting the angular direction in the histogram of gradient that corresponds to the maximum accumulated magnitude as the dominant gradient direction.

7 . The video decoding method of claim 6 , wherein determining the histogram of gradient based on the calculated gradients further comprises:

calculating an angle and a magnitude of the gradient of each template;

converting the angle of the gradient to one of the predefined angular directions in the histogram of gradient; and

adding the magnitude of the gradient to the accumulated magnitude corresponding to the predefined angular direction converted from the angle of the gradient.

8 . The video decoding method of claim 5 , wherein the dominant gradient direction is determined to be the extrapolation direction for extrapolating the neighboring samples of the transform block in calculating the cost function, when a percentage of the neighboring samples of the transform block that belong to the dominant gradient direction exceeds a first predetermined threshold; or

the dominant gradient direction is determined to be the extrapolation direction for extrapolating the neighboring samples of the transform block in calculating the cost function, when a ratio between a magnitude of the gradient of the dominant gradient direction and a sum of magnitudes of all gradients exceeds a second predetermined threshold; or

the dominant gradient direction is determined to be the extrapolation direction for extrapolating the neighboring samples of the transform block in calculating the cost function, when a percentage of the neighboring samples of the transform block that belong to the dominant gradient direction exceeds a first predetermined threshold and that a ratio between a magnitude of the gradient of the dominant gradient direction and a sum of magnitudes of all gradients exceeds a second predetermined threshold.

9 . The video decoding method of claim 5 , wherein an orthogonal direction of the dominant gradient direction is determined to be the extrapolation direction for extrapolating the neighboring samples of the transform block in calculating the cost function.

10 . The video decoding method of claim 2 , wherein generating the dequantized transform coefficients from the quantized transform coefficients further comprises:

determining quantization indices of the dequantized transform coefficients based on transform coefficient levels of the quantized transform coefficients.

11 . The video decoding method of claim 10 , wherein determining quantization indices of the dequantized transform coefficients further comprises: for each quantized transform coefficient,

selecting a quantizer between two predefined scalar quantizers based on the transform coefficient level of a quantized transform coefficient preceding the quantized transform coefficient; and

determining the quantization index of the corresponding dequantized transform coefficient based on the selected quantizer applied to the quantized transform coefficient.

12 . The video decoding method of claim 11 , wherein the quantizer is selected between the two predefined scalar quantizers according to a state transition machine, wherein each state of the state transition machine corresponds to one of the two predefined scaler quantizers,

wherein the video decoding method further comprises:

determining a current state of a state transition machine based on a previous state of a state transition machine and a parity of the transform coefficient level of a quantized transform coefficient preceding the quantized transform coefficient; and

selecting the predefined scaler quantizer corresponding to the current state as the quantizer.

13 . The video decoding method of claim 10 , wherein:

the set of candidate transform coefficients are selected from the dequantized transform coefficients based on magnitudes of the quantization indices of the dequantized transform coefficients.

14 . The video decoding method of claim 10 , wherein the set of candidate transform coefficients are selected from the dequantized transform coefficients based on influence scores of the dequantized transform coefficients on reconstructed border samples of the transform block,

wherein the influence scores of the dequantized transform coefficients on the reconstructed border samples are measured as an L1 norm or an L2 norm of a variation of the quantization index of each dequantized transform coefficient on the reconstructed border samples.

15 . The video decoding method of claim 1 , wherein generating the plurality of candidate hypotheses for the set of candidate transform coefficients further comprises:

determining a plurality of combinations of sign candidates for the set of candidate transform coefficients based on a total number of candidate transform coefficients in the set of candidate transform coefficients; and

applying a template-based hypothesis generation scheme to generate the plurality of candidate hypotheses for the plurality of combinations of sign candidates, respectively.

16 . The video decoding method of claim 1 , wherein the sequence of sign signaling bits for the set of candidate transform coefficients indicate whether original signs of the candidate transform coefficients are identical to predicted signs of the candidate transform coefficients.

17 . The video decoding method of claim 16 , wherein the sequence of sign signaling bits includes a first bin with a value of zero,

wherein estimating the original signs for the set of candidate transform coefficients based on the set of predicted signs and the sequence of sign signaling bits further comprises:

adopting the predicted signs of a first group of candidate transform coefficients corresponding to the first bin in the sequence of sign signaling bits as the original signs of the first group of candidate transform coefficients.

18 . The video decoding method of claim 17 , wherein the sequence of sign signaling bits further includes a second bin with a value of one and a set of additional bins for informing corresponding correctness of the predicted signs of a second group of candidate transform coefficients,

wherein estimating the original signs for the set of candidate transform coefficients based on the set of predicted signs and the sequence of sign signaling bits further comprises:

correcting the predicted signs of the second group of candidate transform coefficients according to the additional set of bins before adopting the predicted signs as the original signs of the second group of candidate transform coefficients.

19 . A video decoding apparatus for transform coefficient sign prediction, comprising:

a memory configured to store a video comprising a plurality of video frames; and

one or more processors coupled to the memory and configured to:

generate a plurality of candidate hypotheses for a set of candidate transform coefficients associated with a transform block of a video frame from a video;

select a hypothesis from the plurality of candidate hypotheses as a set of predicted signs for the set of candidate transform coefficients, wherein the hypothesis is selected based on a cost function calculated by extrapolating neighboring samples of the transform block in an extrapolation direction, wherein the extrapolation direction is determined based on a dominant gradient direction, wherein the dominant gradient direction is a direction with the largest magnitude in a histogram of gradient obtained based on an accumulated magnitude of gradients for each of a plurality of angular directions; and

estimate original signs for the set of candidate transform coefficients based on the set of predicted signs and a sequence of sign signaling bits received from a video encoder.

20 . A non-transitory computer-readable storage medium having stored therein a bitstream comprising video information to be decoded by acts comprising:

generating a plurality of candidate hypotheses for a set of candidate transform coefficients associated with a transform block of a video frame from a video;

selecting a hypothesis from the plurality of candidate hypotheses as a set of predicted signs for the set of candidate transform coefficients, wherein the hypothesis is selected based on a cost function calculated by extrapolating neighboring samples of the transform block in an extrapolation direction, wherein the extrapolation direction is determined based on a dominant gradient direction, wherein the dominant gradient direction is a direction with the largest magnitude in a histogram of gradient obtained based on an accumulated magnitude of gradients for each of a plurality of angular directions; and

estimating original signs for the set of candidate transform coefficients based on the set of predicted signs and a sequence of sign signaling bits received through a bit stream from a video encoder,

wherein the bit stream is stored in the non-transitory computer-readable storage medium.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2024
From: CHEN, YI-WEN
To: KWAI, INC.
Reel/Frame 066818/0029 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2024
From: KWAI, INC.
To: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 066818/0046 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2024
From: XIU, XIAOYU; YAN, NING; KUO, CHE-WEI; CHEN, WEI; JHU, HONG-JHENG; WANG, XIANGLIN; YU, BING
To: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 066978/0106 →
Continuity (7)
Continuation In Part 18435883 · Feb 7, 2024
Continuation In Part PCTUS2022043607 · Sep 15, 2022
Continuation PCTUS2022040442 · Aug 16, 2022
Provisional Application 63250797 · Sep 30, 2021
Provisional Application 63244317 · Sep 15, 2021
Provisional Application 63233940 · Aug 17, 2021
Related Publication 20240223811A1 · Jul 4, 2024
References Cited (53)
US 8130828B2 · Hsu · 2012 [cited by examiner]
US 10609367B2 · Zhao · 2020 [cited by examiner]
US 10713523B2 · Paschalakis et al. · 2020 [cited by applicant]
US 20100166074A1 · Ho et al. · 2010 [cited by applicant]
US 20130039423A1 · Helle · 2013 [cited by examiner]
US 20150221068A1 · Martensson et al. · 2015 [cited by applicant]
US 20160014421A1 · Cote et al. · 2016 [cited by applicant]
US 20170142444A1 · Henry · 2017 [cited by applicant]
US 20180176556A1 · Zhao · 2018 [cited by examiner]
US 20180176563A1 · Zhao et al. · 2018 [cited by applicant]
US 20190208225A1 · Chen · 2019 [cited by examiner]
US 20200396487A1 · Nalci et al. · 2020 [cited by applicant]
US 20200404311A1 · Filippov et al. · 2020 [cited by applicant]
US 20210014509A1 · Filippov · 2021 [cited by examiner]
US 20210067807A1 · Lainema · 2021 [cited by applicant]
US 20210297703A1 · Li · 2021 [cited by examiner]
US 20220174281A1 · Jiang · 2022 [cited by examiner]
KR 1020200064171A · 2020 [cited by applicant]
WO 2017115028A1 · 2017 [cited by applicant]
WO 2019135930A1 · 2019 [cited by applicant]
WO 2019172797A1 · 2019 [cited by applicant]
WO 2019172802A1 · 2019 [cited by applicant]
WO 20190172798A1 · 2019 [cited by applicant]
WO 2023055300A2 · 2023 [cited by applicant]
Maryam Mokhtari et al., (hereinafter Mokhtari); “Texture Classification using Dominant Gradient Descriptors”, Conference of Machine Vision and Image Processing (MVIP), IEEE, 2013 (Year: 2013). [cited by examiner]
Anthony Nasrallah et al., (hereinafter Nasrallah) “Decoder-Side Intra Mode Derivation With Texture Analysis in VVC Test Model” *Ateme, France, 978-1-5386-6249-6 @ 2019 IEEE (Year: 2019). [cited by examiner]
Alexander Alshin et al., :Description of SDR, HDR and 360 video coding technology proposal considering mobile application scenario by Samsung, Huawei, GoPro, and HiSilicon JVET-J0024-v2, San Diego, US Apr. 10-20, 2018 (… [cited by examiner]
International Search Report and Written Opinion in related PCT Application No. PCT/US2022/040442 dated Nov. 25, 2022 (11 pages). [cited by applicant]
International Search Report and Written Opinion in related PCT Application No. PCT/US2023/010901 dated May 10, 2023 (9 pages). [cited by applicant]
International Search Report and Written Opinion in related PCT Application No. PCT/US2022/049264 dated Apr. 5, 2023 (9 pages). [cited by applicant]
International Search Report and Written Opinion in related PCT Application No. PCT/US2022/053150 dated Apr. 14, 2023 (12 pages). [cited by applicant]
Jianle Chen et al., Algorithm Description for Versatile Video Coding and Test Model 13 (VTM 13), Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC I/SC 29, [Document: JVETV2002-v1 (Version 1 )], 22nd M… [cited by applicant]
Joel Sole et al., Transform Coefficient Coding in HEVC, IEEE Transactions on Circuits and Systems for Video Technology (vol. 22, No. 12), pp. 1765-1777, Dec. 12, 2012. [cited by applicant]
Sehwan Ki et al., Learning-Based JND-Directed HDR Video Preprocessing for Perceptually Lossless Compression With HEVC, IEEE Access (vol. 8), pp. 228605-228618, Dec. 31, 2020. [cited by applicant]
Mohammed Golam Sarwer et al., AHG12: On Sign Pediction, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29 24th Meeting, by teleconference, JVET-X0120-v2, Oct. 7, 2021. [cited by applicant]
Jie Chen et al., EE2-4.3 related: More Combined Test Results for Sign Prediction, JVET-Y0141-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 25th Meeting, by teleconference, pp. 1-5, Jan… [cited by applicant]
International Search Report and Written Opinion in related PCT Application No. PCT/US2022/043607 dated Dec. 12, 2022 (11 pages). [cited by applicant]
Heiko Schwarz et al., “Additional Support of Dependent Quantization with 8 States”, JVET-Q0243-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 17th Meeting: Brussels, BE, pp. 1-12,… [cited by applicant]
Office Action in related Japanese Application No. 2024-516981, dated May 27, 2025 (8 pgs.). [cited by applicant]
Supplementary European Search Report in related European Application No. 22870667.7 dated Jul. 7, 2025 (13 pgs.). [cited by applicant]
Mohsen Abdoli, et al. , Non-CE3: Decoder-side Intra Mode Derivation with Prediction Fusion Using Planar , Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 , JVET-00449-v21, 15th Meeting:… [cited by applicant]
Felix Henry and Gordon Clare , Residual Coefficient Sign Prediction , Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 , JVET-D0031 , 4th Meeting: Chengdu, CN , Oct. 15-21, 2016 (8 p… [cited by applicant]
Xiaoyu Xiu, et al. , AHG12: Enhanced sign prediction , Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29 , JVET-X0150-v2 , 24th Meeting, by teleconference , Oct. 6-15, 2021 (5 pgs.). [cited by applicant]
Nasrallah Anthony et al: “Decoder-Side Intra Mode Derivation with Texture Analysis in VVC Test Model”, 2019 IEEE International Conference On Image Processing (ICIP), IEEE, Sep. 22, 2019 (Sep. 22, 2019), pp. 3153-3157, X… [cited by applicant]
Coban M et al: “Algorithm description of Enhanced Compression Model 2 (ECM 2)”, 135. Mpeg Meeting; Jul. 12, 2021-Jul. 16, 2021; Online; (Motion Picture Expert Group or Iso/Iec JTC1/SC29/WG11),, No. m57745; JVET-W2025 Se… [cited by applicant]
Extended European Search report in related European Application No. 22893519.3, dated Oct. 27, 2025 (17 pages). [cited by applicant]
Final Office Action in related U.S. Appl. No. 18/658,783, filed Nov. 14, 2025 (24 pages). [cited by applicant]
Extended European Search Report in related European Application No. 22908485.0, dated Nov. 14, 2025 (13 pages). [cited by applicant]
Notice of Allowance in related U.S. Appl. No. 18/743,673, dated Dec. 16, 2025 (16 pages). [cited by applicant]
Extended European Search Report in related European Application No. 23740726.7, dated Nov. 14, 2025 (13 pages). [cited by applicant]
Alshin A et al: “Description of SDR, HDR and 360° video coding technology proposal by Samsung, Huawei, GoPro, and HiSilicon ” mobile application scenario, 10. Jvet Meeting; Oct. 4, 2018-Apr. 20, 2018; San Diego; (The Jo… [cited by applicant]
Office Action in related U.S. Appl. No. 18/775,207 dated Jan. 12, 2026 (21 pages). [cited by applicant]
Office Action in related Indian Application No. 202417028954 dated Dec. 19, 2025 (8 pages). [cited by applicant]