IP Library Granted Patent US 12,604,039
Granted Patent B2
US 12,604,039 · App. 18/743,673 · Granted Apr 14, 2026

Sign prediction for block-based video coding

Inventors: Xiaoyu Xiu (San Diego, CA); Ning Yan (San Diego, CA); Yi-Wen Chen (San Diego, CA); Che-Wei Kuo (San Diego, CA); Wei Chen (San Diego, CA); Hong-Jheng Jhu (San Diego, CA); Han Gao (San Diego, CA); Xianglin Wang (San Diego, CA); Bing Yu (Beijing, CN)
Assignee: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
H04N19/60H04N19/124H04N19/18H04N19/91
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,604,039
App. No.
18/743,673
Filed
Jun 14, 2024
Granted
Apr 14, 2026
Kind
B2
Art Unit
2487
USPC
375/240.12
Abstract

Implementations of the disclosure provide a video decoding apparatus and method for transform coefficient sign prediction on a video decoder side. The method may include receiving a bitstream including a sequence of sign signaling bits for a set of candidate transform coefficients. The method may further include generating a set of predicted signs for the set of candidate transform coefficients associated with a transform block of a video frame from a video. The method may also include decoding the sequence of sign signaling bits based on one or more contexts used to entropy-encode the sequence of sign signaling bits to obtain an indication of correctness of the predicted signs of the respective candidate transform coefficients. The method may additionally include estimating original signs for the set of candidate transform coefficients based on the set of predicted signs and the decoded sequence of sign signaling bits.

Claims (52)

1 . A video decoding method for sign prediction of transform coefficients, comprising:

receiving a bitstream comprising a sequence of sign signaling bits for a set of candidate transform coefficients;

generating a set of predicted signs for the set of candidate transform coefficients associated with a transform block of a video frame;

decoding the sequence of sign signaling bits based on one or more contexts used to entropy-encode the sequence of sign signaling bits to obtain an indication of correctness of the predicted signs of the respective candidate transform coefficients, wherein a context of the one or more contexts used to entropy-encode a sign signaling bit is determined based on a scan position of a candidate transform coefficient corresponding to the sign signaling bit; and

estimating original signs for the set of candidate transform coefficients based on the set of predicted signs and the sequence of decoded sign signaling bits.

2 . The video decoding method of claim 1 , wherein estimating the original signs for the set of candidate transform coefficients based on the set of predicted signs and the sequence of sign signaling bits comprises:

updating the predicted signs of the candidate transform coefficients according to the correctness indicated by the respective sign signaling bits to determine the original signs of the candidate transform coefficients.

3 . The video decoding method of claim 1 , wherein decoding the sequence of sign signaling bits comprises:

determining a context used to entropy-encode the sign signaling bit based on a magnitude of the candidate transform coefficient corresponding to the sign signaling bit.

4 . The video decoding method of claim 3 , wherein magnitudes of the candidate transform coefficients fall within a plurality of magnitude segments, wherein a plurality of contexts are assigned to entropy-encode sign signaling bits for transform coefficients belonging to the respective magnitude segments,

wherein determining the context used to entropy-encode the sign signaling bit comprises:

determining a magnitude segment, among the plurality of magnitude segments, to which the magnitude of the candidate transform coefficient corresponding to the sign signaling bit belongs; and

determining the context assigned to the magnitude segment as the context used to entropy-encode the sign signaling bit.

5 . The video decoding method of claim 4 , wherein every two consecutive magnitude segments, among the plurality of magnitude segments, are separated by a threshold.

6 . The video decoding method of claim 3 , wherein the bitstream further comprises encoded data of quantized levels of the set of candidate transform coefficients,

wherein the video decoding method further comprises:

determining the magnitude of the candidate transform coefficient by directly parsing a quantized level of the candidate transform coefficient from the bitstream without dequantization.

7 . The video decoding method of claim 3 , wherein the bitstream further comprises encoded data of quantized levels of the set of candidate transform coefficients,

wherein the video decoding method further comprises:

parsing a quantized level of the candidate transform coefficient from the bitstream;

determining a quantization index of the candidate transform coefficient based on the quantized level of the candidate transform coefficient and a transition state of the candidate transform coefficient; and

determining the magnitude of the candidate transform coefficient based on the quantization index of the candidate transform coefficient.

8 . The video decoding method of claim 1 , wherein the set of candidate transform coefficients are selected from transform coefficients of the transform block that are reordered according to their magnitudes,

wherein the video decoding method further comprises:

reordering the set of candidate transform coefficients parsed from the bitstream.

9 . The video decoding method of claim 8 , wherein the contexts used to entropy-encode the sign signaling bits are determined based on the magnitudes of reordered transform coefficients.

10 . The video decoding method of claim 8 , wherein the contexts used to entropy-encode the sign signaling bits are determined based on the magnitudes of transform coefficients before the reordering.

11 . The video decoding method of claim 1 , wherein scan positions of the transform coefficients are classified into a plurality of groups, wherein a plurality of contexts are assigned to entropy-encode sign signaling bits for transform coefficients belonging to the respective groups,

wherein the context of the one or more contexts used to entropy-encode a sign signaling bit is determined by:

determining a group, among the plurality of groups, to which the scan position of the candidate transform coefficient corresponding to the sign signaling bit belongs; and

determining the context assigned to the group as the context used to entropy-encode the sign signaling bit.

12 . The video decoding method of claim 1 , wherein decoding the sequence of sign signaling bits comprises:

determining a context used to entropy-encode the sign signaling bit based on coding mode, a block size or component channel information of the candidate transform coefficient corresponding to the sign signaling bit.

13 . The video decoding method of claim 1 , wherein the set of transform coefficients in the bitstream are generated by applying a plurality of different transform cores,

wherein decoding the sequence of sign signaling bits comprises:

determining a context used to entropy-encode the sign signaling bit based on the transform core applied to generate the candidate transform coefficient corresponding to the sign signaling bit.

14 . The video decoding method of claim 13 , wherein the transform cores comprise a multiple transform selection (MTS) and a low-frequency non-separable transform (LFNST).

15 . The video decoding method of claim 1 , wherein generating the set of predicted signs for the set of candidate transform coefficients comprises:

generating a plurality of candidate hypotheses for the set of candidate transform coefficients; and

selecting a hypothesis from the plurality of candidate hypotheses as the set of predicted signs for the set of candidate transform coefficients.

16 . A video decoding apparatus, comprising:

a memory configured to store a bitstream comprising a set of candidate transform coefficients and a sequence of sign signaling bits for the set of candidate transform coefficients; and

a processor coupled to the memory and configured to perform operations comprising:

receiving a bitstream comprising a sequence of sign signaling bits for a set of candidate transform coefficients;

generating a set of predicted signs for the set of candidate transform coefficients associated with a transform block of a video frame;

decoding the sequence of sign signaling bits based on one or more contexts used to entropy-encode the sequence of sign signaling bits to obtain an indication of correctness of the predicted signs of the respective candidate transform coefficients, wherein a context of the one or more contexts used to entropy-encode a sign signaling bit is determined based on a scan position of a candidate transform coefficient corresponding to the sign signaling bit; and

estimating original signs for the set of candidate transform coefficients based on the set of predicted signs and the sequence of decoded sign signaling bits.

17 . A non-transitory computer-readable storage medium having stored therein a bitstream comprising a sequence of sign signaling bits for a set of candidate transform coefficients, wherein the bitstream is decodable by performing operations comprising:

receiving a bitstream comprising a sequence of sign signaling bits for a set of candidate transform coefficients;

generating a set of predicted signs for the set of candidate transform coefficients associated with a transform block of a video frame;

decoding the sequence of sign signaling bits based on one or more contexts used to entropy-encode the sequence of sign signaling bits to obtain an indication of correctness of the predicted signs of the respective candidate transform coefficients, wherein a context of the one or more contexts used to entropy-encode a sign signaling bit is determined based on a scan position of a candidate transform coefficient corresponding to the sign signaling bit; and

estimating original signs for the set of candidate transform coefficients based on the set of predicted signs and the sequence of decoded sign signaling bits.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2024
From: GAO, HAN
To: KWAI, INC.
Reel/Frame 067751/0908 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: CHEN, YI-WEN
To: KWAI, INC.
Reel/Frame 067731/0375 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: KWAI, INC.
To: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 067731/0381 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: XIU, XIAOYU; YAN, NING; KUO, CHE-WEI; CHEN, WEI; JHU, HONG-JHENG; WANG, XIANGLIN; YU, BING
To: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 067736/0136 →
Continuity (13)
Continuation In Part 18658783 · May 8, 2024
Continuation In Part 18606849 · Mar 15, 2024
Continuation In Part 18435883 · Feb 7, 2024
Continuation In Part PCTUS2022053150 · Dec 16, 2022
Continuation PCTUS2022049264 · Nov 8, 2022
Continuation PCTUS2022043607 · Sep 15, 2022
Continuation PCTUS2022040442 · Aug 16, 2022
Provisional Application 63290307 · Dec 16, 2021
Provisional Application 63277705 · Nov 10, 2021
Provisional Application 63250797 · Sep 30, 2021
Provisional Application 63244317 · Sep 15, 2021
Provisional Application 63233940 · Aug 17, 2021
Related Publication 20240414375A1 · Dec 12, 2024
References Cited (50)
US 10609367B2 · Zhao et al. · 2020 [cited by applicant]
US 10713523B2 · Paschalakis et al. · 2020 [cited by applicant]
US 11223849B2 · Lainema · 2022 [cited by examiner]
US 11856216B2 · Filippov · 2023 [cited by examiner]
US 20100166074A1 · Ho et al. · 2010 [cited by applicant]
US 20130039423A1 · Helle et al. · 2013 [cited by applicant]
US 20150221068A1 · Martensson et al. · 2015 [cited by applicant]
US 20160014421A1 · Cote et al. · 2016 [cited by applicant]
US 20180176556A1 · Zhao · 2018 [cited by examiner]
US 20180176563A1 · Zhao · 2018 [cited by examiner]
US 20190208225A1 · Chen · 2019 [cited by examiner]
US 20200396487A1 · Nalci et al. · 2020 [cited by applicant]
US 20200404311A1 · Filippov et al. · 2020 [cited by applicant]
US 20210014509A1 · Filippov · 2021 [cited by examiner]
US 20210297703A1 · Li et al. · 2021 [cited by applicant]
US 20250080735A1 · Haase · 2025 [cited by examiner]
KR 1020200064171A · 2020 [cited by applicant]
WO WO2017115028A1 · 2017 [cited by examiner]
WO WO2019135930A1 · 2019 [cited by examiner]
WO WO2019172798A1 · 2019 [cited by examiner]
WO WO2019172802A1 · 2019 [cited by examiner]
WO WO2019172797A1 · 2019 [cited by examiner]
WO 2020242238A1 · 2020 [cited by applicant]
WO WO2020244238A1 · 2020 [cited by examiner]
Felix Henry et al., “Residual Coefficient Sign Prediction”; JVET-D0031, Chengdu, CN Oct. 15-21, 2016 (Year: 2016). [cited by examiner]
International Search Report and Written Opinion in related PCT Application No. PCT/US22/40442 dated Nov. 25, 2022 (11 pages). [cited by applicant]
International Search Report and Written Opinion in related PCT Application No. PCT/US22/43607 dated Dec. 12, 2022 (11 pages). [cited by applicant]
International Search Report and Written Opinion in related PCT Application No. PCT/US22/49264 dated Apr. 5, 2023 (9 pages). [cited by applicant]
International Search Report and Written Opinion in related PCT Application No. PCT/US22/53150 dated Apr. 14, 2023 (12 pages). [cited by applicant]
Jianle Chen et al., Algorithm Description for Versatile Video Coding and Test Model 13 (VTM 13), Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC I/SC 29, [Document: JVETV2002-v1 (Version 1 )], 22nd M… [cited by applicant]
Joel Sole et al., Transform Coefficient Coding in HEVC, IEEE Transactions on Circuits and Systems for Video Technology (vol. 22, No. 12), pp. 1765-1777, Dec. 12, 2012. [cited by applicant]
Sehwan Ki et al., Learning-Based JND-Directed HDR Video Preprocessing for Perceptually Lossless Compression With HEVC, IEEE Access (vol. 8), pp. 228605-228618, Dec. 31, 2020. [cited by applicant]
Heiko Schwarz et al., Additional Support of Dependent Quantization with 8 states, JVET-Q0243-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 17th Meeting: Brussels, BE, pp. 1-12, J… [cited by applicant]
Mohammed Golam Sarwer et al., AHG12: On Sign Pediction, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29 24th Meeting, by teleconference, JVET-X0120-v2, Oct. 7, 2021. [cited by applicant]
Office Action in related Japanese Application No. 2024-516981 dated May 27, 2025 (8 pages). [cited by applicant]
Office Action in related Japanese Application No. 2024-527601 dated May 27, 2025 (10 pages). [cited by applicant]
Extended European Search Report in related European Application No. EP22870667.7 dated Jul. 7, 2025 (13 pages). [cited by applicant]
Henry et al Residual Coefficient Sign Prediction, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-D0031, 4th Meeting: Chengdu, CN, Oct. 2016, pp. 1-6. [cited by applicant]
Coban M et al: “Algorithm description of Enhanced Compression Model 2 (ECM 2)”, 135. MPEG Meeting; Jul. 12, 2021-Jul. 16, 2021; Online; (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11),, No. m57745; JVET-W2025 Se… [cited by applicant]
“Test Model 13 for Versatile Video Coding 7-9 (VTM 13)”, 134. MPEG Meeting; Apr. 26, 2021-Apr. 30, 2021; Online; (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11),, No. n20241 Jul. 7, 2021 (Jul. 7, 2021), XP030297… [cited by applicant]
Extended European Search Report in related European Application No. 22908485.0 dated Nov. 14, 2025 (13 pages). [cited by applicant]
Notice of Allowance in related U.S. Appl. No. 18/606,849 dated Oct. 24, 2025 (18 pages). [cited by applicant]
Extended European Search Report in related European Application No. 22893519.3 dated Oct. 27, 2025 (17 pages). [cited by applicant]
Office Action in related U.S. Appl. No. 18/658,783 dated Nov. 14, 2025 (24 pages). [cited by applicant]
Extended European Search Report in related European Application No. 23740726.7 dated Nov. 14, 2025 (13 pages). [cited by applicant]
Office Action in related Indian Application No. 202417028954 dated Dec. 19, 2025 (8 pages). [cited by applicant]
Office Action in related U.S. Appl. No. 18/775,207 dated Jan. 12, 2026 (21 pages). [cited by applicant]
Alshin A et al: “Description of SDR, HDR and 360° video coding technology proposal by Samsung, Huawei, GoPro, and HiSilicon mobile application scenario”, 10. JVET Meeting; Oct. 4, 2018-Apr. 20, 2018; San Diego; (The Joi… [cited by applicant]
Maryam Mokhtari et al., (hereinafter Mokhtari); “Texture Classification using Dominant Gradient Descriptors”, Conference of Machine Vision and Image Processing (MVIP), IEEE, 2013 (Year:2013). [cited by applicant]
Xiu (Kwai) X et al: “AHG12: Enhanced sign prediction”, 24. JVET Meeting; Oct. 6, 2021-Oct. 15, 2021; Teleconference; (The Joint Video Exploration Team of ISO/IEC JTC1/SC29/WG11 and ITU-T SG.16 ), , No. JVET-X0150 ; m579… [cited by applicant]