IP Library Granted Patent US 12,335,460
Granted Patent B2
US 12,335,460 · App. 17/838,729 · Granted Jun 17, 2025

Systems and methods for generalized multi-hypothesis prediction for video coding

Inventors: Chun-Chi Chen (Hsinchu, TW); Xiaoyu Xiu (San Diego, CA); Yuwen He (San Diego, CA); Yan Ye (San Diego, CA)
Assignee: INTERDIGITAL VC HOLDINGS, INC.
H04N19/105H04N19/139H04N19/176H04N19/463H04N19/573H04N19/96H04N19/577H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,335,460
App. No.
17/838,729
Granted
Jun 17, 2025
Kind
B2
Abstract

Systems and methods are described for video coding using generalized bi-prediction. In an exemplary embodiment, to code a current block of a video in a bitstream, a first reference block is selected from a first reference picture and a second reference block is selected from a second reference picture. Each reference block is associated with a weight, where the weight may be an arbitrary weight ranging, e.g., between 0 and 1. The current block is predicted using a weighted sum of the reference blocks. The weights may be selected from among a plurality of candidate weights. Candidate weights may be signaled in the bitstream or may be derived implicitly based on a template. Candidate weights may be pruned to avoid out-of-range or substantially duplicate candidate weights. Generalized bi-prediction may additionally be used in frame rate up conversion.

Claims (45)

1. A video encoding method comprising:

for at least a current block in a current picture, encoding block-level information identifying a first weight and a second weight, wherein at least one of the first and second weights has a value not equal to 0, 0.5 or 1; and

for each sub-block in the current block:

obtaining a first sub-block motion vector and a second sub-block motion vector based on an affine motion model of the current block; and

predicting the sub-block as a weighted sum of a first reference block in a first reference picture pointed to by the first sub-block motion vector and a second reference block in a second reference picture pointed to by the second sub-block motion vector, wherein the first reference block is weighted by the first weight and the second reference block is weighted by the second weight;

wherein the first weight and the second weight are shared across each of the sub-blocks in the current block.

2. The method of claim 1 , wherein encoding of the block-level information comprises mapping a weight index to a codeword using a truncated unary code and entropy encoding the codeword in a bitstream.

3. The method of claim 1 , wherein the second weight is identified by subtracting the first weight from one.

4. The method of claim 1 , wherein the first weight and the second weight are identified from among a predetermined set of weights.

5. A video decoding method comprising:

for at least a current block in a current picture, decoding block-level information identifying a first weight and a second weight, wherein at least one of the first and second weights has a value not equal to 0, 0.5 or 1; and

for each sub-block in the current block:

obtaining a first sub-block motion vector and a second sub-block motion vector based on an affine motion model of the current block; and

predicting the sub-block as a weighted sum of a first reference block in a first reference picture pointed to by the first sub-block motion vector and a second reference block in a second reference picture pointed to by the second sub-block motion vector, wherein the first reference block is weighted by the first weight and the second reference block is weighted by the second weight;

wherein the first weight and the second weight are shared across each of the sub-blocks in the current block.

6. The method of claim 5 , wherein the decoding of the block-level information comprises entropy decoding a codeword from a bitstream and recovering a weight index from the codeword using a truncated unary code.

7. The method of claim 5 , wherein the second weight is identified by subtracting the first weight from one.

8. The method of claim 5 , wherein the first weight and the second weight are identified from among a predetermined set of weights.

9. A video encoding apparatus comprising a processor configured to perform at least:

for at least a current block in a current picture, encoding block-level information identifying a first weight and a second weight, wherein at least one of the first and second weights has a value not equal to 0, 0.5 or 1; and

for each sub-block in the current block:

obtaining a first sub-block motion vector and a second sub-block motion vector based on an affine motion model of the current block; and

predicting the sub-block as a weighted sum of a first reference block in a first reference picture pointed to by the first sub-block motion vector and a second reference block in a second reference picture pointed to by the second sub-block motion vector, wherein the first reference block is weighted by the first weight and the second reference block is weighted by the second weight;

wherein the first weight and the second weight are shared across each of the sub-blocks in the current block.

10. The apparatus of claim 9 , wherein encoding of the block-level information comprises mapping a weight index to a codeword using a truncated unary code and entropy encoding the codeword in a bitstream.

11. The apparatus of claim 9 , wherein the second weight is identified by subtracting the first weight from one.

12. The apparatus of claim 9 , wherein the first weight and the second weight are identified from among a predetermined set of weights.

13. A video decoding apparatus comprising a processor configured to perform at least:

for at least a current block in a current picture, decoding block-level information identifying a first weight and a second weight, wherein at least one of the first and second weights has a value not equal to 0, 0.5 or 1; and

for each sub-block in the current block:

obtaining a first sub-block motion vector and a second sub-block motion vector based on an affine motion model of the current block; and

predicting the sub-block as a weighted sum of a first reference block in a first reference picture pointed to by the first sub-block motion vector and a second reference block in a second reference picture pointed to by the second sub-block motion vector, wherein the first reference block is weighted by the first weight and the second reference block is weighted by the second weight;

wherein the first weight and the second weight are shared across each of the sub-blocks in the current block.

14. The apparatus of claim 13 , wherein the decoding of the block-level information comprises entropy decoding a codeword from a bitstream and recovering a weight index from the codeword using a truncated unary code.

15. The apparatus of claim 13 , wherein the second weight is identified by subtracting the first weight from one.

16. The apparatus of claim 13 , wherein the first weight and the second weight are identified from among a predetermined set of weights.

17. A non-transitory computer-readable medium including instructions for causing one or more processors to perform a method comprising:

for at least a current block in a current picture, decoding block-level information identifying a first weight and a second weight, wherein at least one of the first and second weights has a value not equal to 0, 0.5 or 1; and

for each sub-block in the current block:

obtaining a first sub-block motion vector and a second sub-block motion vector based on an affine motion model of the current block; and

predicting the sub-block as a weighted sum of a first reference block in a first reference picture pointed to by the first sub-block motion vector and a second reference block in a second reference picture pointed to by the second sub-block motion vector, wherein the first reference block is weighted by the first weight and the second reference block is weighted by the second weight;

wherein the first weight and the second weight are shared across each of the sub-blocks in the current block.

18. The computer-readable medium of claim 17 , wherein the decoding of the block-level information comprises entropy decoding a codeword from a bitstream and recovering a weight index from the codeword using a truncated unary code.

19. The computer-readable medium of claim 17 , wherein the second weight is identified by subtracting the first weight from one.

20. The computer-readable medium of claim 17 , wherein the first weight and the second weight are identified from among a predetermined set of weights.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: VID SCALE, INC.
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068284/0031 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2023
From: CHEN, CHUN-CHI; XIU, XIAOYU; HE, YUWEN; YE, YAN
To: VID SCALE, INC.
Reel/Frame 062292/0170 →
Continuity (6)
Continuation 16099392
Provisional Application 62415187 · Oct 31, 2016
Provisional Application 62399234 · Sep 23, 2016
Provisional Application 62342772 · May 27, 2016
Provisional Application 62336227 · May 13, 2016
Related Publication 20220312001A1 · Sep 29, 2022
References Cited (106)
US 5832124A · Sato · 1998 [cited by applicant]
US 7801217B2 · Boyce · 2010 [cited by applicant]
US 10298939B2 · Coban · 2019 [cited by applicant]
US 10805631B2 · Lee · 2020 [cited by applicant]
US 20030215014A1 · Koto · 2003 [cited by applicant]
US 20040008782A1 · Boyce · 2004 [cited by applicant]
US 20040008786A1 · Boyce · 2004 [cited by applicant]
US 20040141615A1 · Chujoh · 2004 [cited by applicant]
US 20040184544A1 · Kondo · 2004 [cited by applicant]
US 20060268166A1 · Bossen · 2006 [cited by examiner]
US 20070110390A1 · Toma · 2007 [cited by applicant]
US 20100215095A1 · Hayase · 2010 [cited by applicant]
US 20110007803A1 · Karczewicz · 2011 [cited by applicant]
US 20120163455A1 · Zheng · 2012 [cited by applicant]
US 20120230417A1 · Sole Rojals · 2012 [cited by applicant]
US 20130243093A1 · Chen · 2013 [cited by applicant]
US 20130259122A1 · Sugio · 2013 [cited by applicant]
US 20140105299A1 · Chen · 2014 [cited by applicant]
US 20140153647A1 · Nakamura · 2014 [cited by applicant]
US 20140161175A1 · Zhang et al. · 2014 [cited by applicant]
US 20140198846A1 · Guo · 2014 [cited by applicant]
US 20140253681A1 · Zhang · 2014 [cited by applicant]
US 20140321551A1 · Ye · 2014 [cited by applicant]
US 20140362922A1 · Puri · 2014 [cited by applicant]
US 20150195563A1 · Ramasubramonian · 2015 [cited by applicant]
US 20150319441A1 · Puri · 2015 [cited by applicant]
US 20160029035A1 · Nguyen · 2016 [cited by examiner]
US 20170034513A1 · Leontaris · 2017 [cited by applicant]
US 20170264904A1 · Koval · 2017 [cited by applicant]
US 20170280163A1 · Kao · 2017 [cited by applicant]
US 20180249171A1 · Lim · 2018 [cited by applicant]
US 20190230350A1 · Chen · 2019 [cited by applicant]
CN 101176350A · 2008 [cited by applicant]
CN 101695114A · 2010 [cited by applicant]
CN 101855910A · 2010 [cited by applicant]
CN 101902645A · 2010 [cited by applicant]
CN 102804775A · 2012 [cited by applicant]
CN 103636202A · 2014 [cited by applicant]
CN 104769948A · 2015 [cited by applicant]
CN 104798372A · 2015 [cited by applicant]
CN 104969551A · 2015 [cited by applicant]
CN 105009586A · 2015 [cited by applicant]
CN 105453570A · 2016 [cited by applicant]
CN 105830446A · 2016 [cited by applicant]
EP 2394437A1 · 2011 [cited by applicant]
EP 2394437B1 · 2011 [cited by applicant]
EP 2735151A1 · 2014 [cited by applicant]
EP 2763414 · 2014 [cited by applicant]
EP 2951996A1 · 2015 [cited by applicant]
JP 2004007377 · 2004 [cited by applicant]
JP 2005533466 · 2005 [cited by applicant]
JP 2008541502 · 2008 [cited by applicant]
JP 2010028864 · 2010 [cited by applicant]
KR 20050021487A · 2005 [cited by applicant]
WO 2004008761 · 2004 [cited by applicant]
WO 2004008761A1 · 2004 [cited by applicant]
WO 2010090749 · 2010 [cited by applicant]
WO 2014039802A2 · 2014 [cited by applicant]
WO 2014081775A1 · 2014 [cited by applicant]
WO 2017197146 · 2017 [cited by applicant]
Invitation to Pay Additional Fees, And Where Applicable, Protest Fee For PCT/US2017/032208 mailed on Aug. 14, 2017, 13 Pages. [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority, for PCT/US2017/032208 mailed Oct. 12, 2017, 19 pages. [cited by applicant]
International Preliminary Report on Patentability for PCT/US2017/032208 issued on Nov. 13, 2018. [cited by applicant]
International Telecommunication Union, “Advanced Video Coding for Generic Audiovisual Services”. Series H: Audiovisual and Multimedia System; Infrastructure of audiovisual services, Coding of moving video, ITU-T Recomme… [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority for PCT/US2019/014691 mailed May 15, 2019, 13 pages. [cited by applicant]
Liu, H., et. al., “Local Illumination Compensation”. Qualcomm Incorporated, Video Coding Experts Group (VCEG), Telecommunications Standardization Sector ITU-T SG16/Q6, Doc. VCEG-AZ06, Power Point Presentation, Jun. 2015… [cited by applicant]
Alshina, E., et. al., “Known Tools Performance Investigation for Next Generation Video Coding”. Samsung Electronics, Video Coding Experts Group (VCEG), Telecommunications Standardization Sector ITU-T SG16/Q6, Doc. VCEG-… [cited by applicant]
Chen, C.-C., et. al., “Generalized Bi-Prediction for Inter Coding”. InterDigital Communications, Inc., Joint Video Exploration Team (JVET), ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Doc. JVET-C0047, May 2016, 4 pa… [cited by applicant]
Chen, C.-C., et. al., “Generalized Bi-Prediction for Inter Coding”. InterDigital Communications, Inc., Joint Video Exploration Team (JVET), ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Doc. JVET-C0047, Power Point Pr… [cited by applicant]
Suehring, K., et. al., “JVET Common Test Conditions and Software Reference Configurations”. Joint Video Exploration Team (JVET), ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, Doc. JVET-B1010, Feb. 2016. [cited by applicant]
SMPTE 421M, “VC-1 Compressed Video Bitstream Format and Decoding Process”. SMPTE Standard, 2006, (493 pages). [cited by applicant]
Sullivan, G. J., et. al., “Overview of the High Efficiency Video Coding (HEVC) Standard”. IEEE Transaction on Circuits and Systems for Video Technology, vol. 22, No. 12, Dec. 2012, pp. 1649-1668. [cited by applicant]
Liu, H., et. al., “Local Illumination Compensation”. Qualcomm Incorporated, Video Coding Experts Group (VCEG), Telecommunications Standardization Sector ITU-T SG16/Q6, Doc. VCEG-AZ06, Jun. 2015, 4 pages. [cited by applicant]
Alshina, E., et. al., “Known Tools Performance Investigation for Next Generation Video Coding”. Samsung Electronics, Video Coding Experts Group (VCEG), Telecommunications Standardization Sector ITU-T SG16/Q6, Doc. VCEG-… [cited by applicant]
Wikipedia “Exponential-Golomb Coding”. Wikipedia article modified on Jan. 30, 2016, available at: https://en.wikipedia.org/w/index.php?title=Exponential-Golomb_coding&oldid=702406490, 2 pages. [cited by applicant]
Benjamin, et. al., “High Efficiency Video Coding (HEVC) Text Specification Draft 10 (for FDIS and Last Call)”. Joint Collaborative Team on Video Coding (JCT-VC), Document No. JCTVC-L1003, Jan. 2013, 310 pages. [cited by applicant]
He, Y., et. al., “CE4-related: Encoder Speed Up and Bug Fix for Generalized Bi-Prediction in BMS-2.1”. The Joint Video Exploration Team (JVET) Meeting, Oct. 3-12, 2018, pp. 1-5. [cited by applicant]
Chen , Jianle, et. al., “Coding Tools Investigation for Next Generation Video Coding”. ITU-Telecommunication Standardization Sector, Study Group 16, Contribution 806, COM16-C806, Jan. 2015, pp. 1-7. [cited by applicant]
An, J., et. al., “Block Partitioning Structure for Next Generation Video Coding”. MediaTek Inc., Telecommunication Standardization Sector ITU-T SG16/Q6, Doc. COM16-C966, Sep. 2015, 8 pages. [cited by applicant]
International Telecommunication Union, “Affine Transform Prediction for Next Generation Video Coding”. Huawei Technologies Co., Ltd., Telecommunication Standardization Sector, ITU-T SG16/Q6 Doc. COM16-C1016, Sep. 2015, … [cited by applicant]
International Telecommunication Union, “Advanced Video Coding for Generic Audiovisual Services”. Series H: Audiovisual and Multimedia System; Infrastructure of audiovisual services, Coding of moving video, ITU-T Recomme… [cited by applicant]
Kamikura, Kazuto, et. al., “Global Brightness-Variation Compensation for Video Coding”. IEEE Transactions on Circuits and Systems for Video Technology, vol. 8, No. 8, Dec. 1998, pp. 988-1000. [cited by applicant]
International Preliminary Report on Patentability for PCT/US2019/014691 issued on Jul. 28, 2020, 9 pages. [cited by applicant]
International Telecommunication Union, “High Efficiency Video Coding”. Series H: Audiovisual and Multimedia Systems; Infrastructure of Audiovisual Services—Coding of Moving Video, Recommendation ITU-T H.265, Telecommuni… [cited by applicant]
Sharabayko M. et al. “Iterative intra prediction search for H. 265/HEVC”. 2013 International Siberian Conference on Control and Communications (SIBCON), IEEE, Sep. 12, 2013, (4 pages). [cited by applicant]
Zhang, H., et al. “Fast intra prediction for high efficiency video coding.” In: “Advances in Databases and Information Systems”, Jan. 1, 2012, Springer International Publishing, vol. 7674, pp. 568-577 (10 pages). [cited by applicant]
Discrete Consine Transform, http://ww.mathworks.com/help/images/discrete-consine-transform.html, Nov. 2012, 3 pages. [cited by applicant]
International Telecommunication Union, “Advanced Video Coding for Generic Audiovisual Services”. In Series H: Audiovisual and Multimedia Systems; Infrastructure of audiovisual services; Coding of moving video. ITU-T Rec… [cited by applicant]
Ohm, Jens-Rainer., et. al., “Report of AHG on Future Video Coding Standardization Challenges”. International Organization for Standardization, Coding of Moving Pictures and Audio, ISO/IEC JTC1/SC29/WG11 MPEG2014/M36782,… [cited by applicant]
Alshina, E., et. al., “Known Tools Performance Investigation for Next Generation Video Coding”. ITU—Telecommunications Standardization Sector, Video Coding Experts Group (VCEG), SG16/Q6, Power Point Presentation, VCEG-A… [cited by applicant]
Karczewicz, M., et. al., “Report of AHG1 On Coding Efficiency Improvements”. ITU—Telecommunications Standardization Sector, Video Coding Experts Group (VCEG), SG16/Q6, VCEG-AZ01, Jun. 2015, 2 pages. [cited by applicant]
An, J., et. al., “Block Partitioning Structure for Next Generation Video Coding”. ITU-Telecommunication Standardization Sector, Study Group 16, Contribution 966 R3, COM16-C966R3-E, Sep. 2015, pp. 1-8. [cited by applicant]
Invitation to pay additional fees and, where applicable, protest fee for PCT/US2017/031303 mailed Jul. 11, 2017, 17 pages. [cited by applicant]
Motra, Ajit Singh, et. al., “Fast Intra Mode Decision for HEVC Video Encoder”. IEEE International Conference on Software, Telecommunications and Computer Networks (SOFTCOM), Sep. 11, 2012, 5 pages. [cited by applicant]
Xiu, Xiaoyu, et. al., “Decoder-side Intra Mode Derivation for Block-Based Video Coding”. Picture Coding Symposium (PCS), IEEE, Dec. 4, 2016, 5 pages. [cited by applicant]
Zhang, Zhenming, et. al., “Improved Intra Prediction Mode-Decision Method”. In Visual Communications and Image Processing, vol. 5960, Jul. 12, 2005, pp. 632-640. [cited by applicant]
Wiegand, Thomas, et. al., “Overview of the H.264/AVC Video Coding Standard”. IEEE Transactions on Circuits and Systems for Video Technology, vol. 13, No. 7, Jul. 2003, pp. 560-576. [cited by applicant]
International Preliminary Report on Patentability for PCT/US2017/031303 issued on Nov. 6, 2018. [cited by applicant]
Alshina, E., et al., “Known Tools Performance Investigation for Next Generation Video Coding”. ITU—Telecommunications Standardization Sector, Video Coding Experts Group (VCEG), SG16/Q6, VCEG-AZ05, Jun. 2015, 7 pages. [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority, for PCT/US2017/31303 mailed Sep. 8, 2017, 21 pages. [cited by applicant]
Chen, Jianle, et. al., “Coding Tools Investigation for Next Generation Video Coding Based on HEVC”. Applications of Digital Image Processing XXXVIII, vol. 9599, International Society for Optics and Photonics, (2015), pp… [cited by applicant]
Chen, Jianle, et al. “Description of scalable video coding technology proposal by Qualcomm (configuration 1).” Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, Doc. JCTVC-K… [cited by applicant]
Sullivan, Gary J., et. al., “Overview of The High Efficiency Video Coding (HEVC) Standard”. IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, Dec. 2012, pp. 1649-1668. [cited by applicant]
Hannuksela, M., “Generalized B/MH-Picture Averaging.” In Joint Video Team (JVT) of ISO/IEC MPEG & ITU-T VCEG (ISO/IEC JTC1/SC29/WG11 and ITU-T SG16 Q. 6), 3rd Meeting: Fairfax, VA, USA, 2002 (8 pages). [cited by applicant]
Kikuchi, Y. et al., “Multi-frame interpolative prediction with modified syntax.” Joint Video Team (JVT) of ISO/IEC MPEG & ITU-T VCEG,(ISO/IEC JTC1/SC29/WG11 and ITU-T SG16 Q. 6) Document: JVT-C066, 3rd Meeting: Fairfax,… [cited by applicant]
Huang, et al, “A New Multi-hypothesis Motion-compensated Prediction Algorithm”, Journal of Image and Graphic, vol. 13 No. 3,, Mar. 2008 (7 pages). [cited by applicant]