IP Library Granted Patent US 12,212,773
Granted Patent B2
US 12,212,773 · App. 18/097,390 · Granted Jan 28, 2025

Sub-block motion derivation and decoder-side motion vector refinement for merge mode

Inventors: Xiaoyu Xiu (San Diego, CA); Yuwen He (San Diego, CA); Yan Ye (San Diego, CA)
Assignee: InterDigital VC Holdings, Inc.
H04N19/51H04N19/105H04N19/174H04N19/176H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,212,773
App. No.
18/097,390
Granted
Jan 28, 2025
Kind
B2
Abstract

Systems, methods, and instrumentalities for sub-block motion derivation and motion vector refinement for merge mode may be disclosed herein. Video data may be coded (e.g., encoded and/or decoded). A collocated picture for a current slice of the video data may be identified. The current slice may include one or more coding units (CUs). One or more neighboring CUs may be identified for a current CU. A neighboring CU (e.g., each neighboring CU) may correspond to a reference picture. A (e.g., one) neighboring CU may be selected to be a candidate neighboring CU based on the reference pictures and the collocated picture. A motion vector (MV) (e.g., collocated MV) may be identified from the collocated picture based on an MV (e.g., a reference MV) of the candidate neighboring CU. The current CU may be coded (e.g., encoded and/or decoded) using the collocated MV.

Claims (62)

1. A device for video decoding, comprising:

a processor configured to:

obtain a collocated picture associated with a current video block;

determine a constrained region in the collocated picture based on a location of the current video block, wherein a size of the constrained region is greater than a size of the current video block;

determine a first location associated with a collocated block based on the location of the current video block and a temporal motion vector (MV);

determine whether the first location associated with the collocated block is within the constrained region in the collocated picture, wherein, based on the first location associated with the collocated block being outside the constrained region, the processor is configured to determine a second location associated with the collocated block using a clipping operation associated with the constrained region in the collocated picture, wherein, the clipping operation changes the first location associated with the collocated block based on a boundary of the constrained region;

obtain, in the collocated picture, the collocated block associated with the current video block at the second location associated with the collocated block;

obtain an MV associated with the collocated block; and

decode the current video block based on the MV associated with the collocated block.

2. The device of claim 1 , wherein the current video block is an advanced temporal motion vector prediction (ATMVP) subblock, and the collocated block is a collocated subblock associated with the ATMVP subblock, and wherein the processor is further configured to:

obtain an MV associated with the collocated subblock; and

predict the ATMVP subblock based on the MV associated with the collocated subblock, wherein the current video block is decoded based on the prediction of the ATMVP subblock.

3. The device of claim 1 , wherein the processor is further configured:

determine a picture order count (POC) difference between the collocated picture and a reference picture of a neighboring block of the current video block; and

determine the temporal MV based on the POC difference, wherein the constrained region in the collocated picture is determined further based on the temporal MV.

4. The device of claim 1 , wherein a coding tree block (CTU) comprises the current video block, and the constrained region is further determined based on a location of the CTU in the collocated picture.

5. The device of claim 1 , wherein a coding tree block (CTU) comprises the current video block, and wherein an area of the constrained region is determined to be the same size as an area of the CTU plus one column of a video block.

6. The device of claim 1 , wherein a coding tree block (CTU) comprises the current video block, and wherein an area of the constrained region is determined to be the same size as an area of the CTU.

7. The device of claim 1 , wherein the MV of the collocated block is associated with a reference picture of the collocated block.

8. The device of claim 1 , wherein the second location associated with the collocated block is on the boundary of the constrained region.

9. A method for video decoding, comprising:

obtaining a collocated picture associated with a current video block;

determining a constrained region in the collocated picture based on a location of the current video block, wherein a size of the constrained region is greater than a size of the current video block;

determining a first location associated with a collocated block based on the location of the current video block and a temporal motion vector (MV);

determining whether the first location associated with the collocated block is within the constrained region in the collocated picture, wherein, based on the first location associated with the collocated block being outside the constrained region, the method comprises determining a second location associated with the collocated block using a clipping operation associated with the constrained region in the collocated picture, wherein the clipping operation changes the first location associated with the collocated block based on a boundary of the constrained region;

obtaining, in the collocated picture, the collocated block associated with the current video block at the second location associated with the collocated block;

obtaining an MV associated with the collocated block; and

decoding the current video block based on the MV associated with the collocated block.

10. The method of claim 9 , wherein the MV of the collocated block is associated with a reference picture of the collocated block.

11. The method of claim 9 , wherein the current video block is an advanced temporal motion vector prediction (ATMVP) subblock, and the collocated block is a collocated subblock associated with the ATMVP subblock, and wherein the method further comprises:

obtaining an MV associated with the collocated subblock; and

predicting the ATMVP subblock based on the MV associated with the collocated subblock, wherein the current video block is decoded based on the prediction of the ATMVP subblock.

12. The method of claim 9 , wherein the second location associated with the collocated block is on the boundary of the constrained region.

13. A device for video encoding, comprising:

a processor configured to:

obtain a collocated picture associated with a current video block;

determine a constrained region in the collocated picture based on a location of the current video block, wherein a size of the constrained region is greater than a size of the current video block;

determine a first location associated with a collocated block based on the location of the current video block and a temporal motion vector (MV);

determine whether the first location associated with the collocated block is within the constrained region in the collocated picture, wherein, based on the first location associated with the collocated block being outside the constrained region, the processor is configured to determine a second location associated with the collocated block using a clipping operation associated with the constrained region in the collocated picture, wherein the clipping operation changes the first location associated with the collocated block based on a boundary of the constrained region;

obtain, in the collocated picture, the collocated block associated with the current video block at the second location associated with the collocated block;

obtain an MV associated with the collocated block; and

encode the current video block based on the MV associated with the collocated block.

14. The device of claim 13 , wherein the MV of the collocated block is associated with a reference picture of the collocated block.

15. The device of claim 13 , wherein the current video block is an advanced temporal motion vector prediction (ATMVP) subblock, and the collocated block is a collocated subblock associated with the ATMVP subblock, and wherein the processor is further configured to:

obtain an MV associated with the collocated subblock; and

predict the ATMVP subblock based on the MV associated with the collocated subblock, wherein the current video block is encoded based on the prediction of the ATMVP subblock.

16. The device of claim 13 , wherein the second location associated with the collocated block is on the boundary of the constrained region.

17. A method for video encoding, comprising:

obtaining a collocated picture associated with a current video block;

determining a constrained region in the collocated picture based on a location of the current video block, wherein a size of the constrained region is greater than a size of the current video block;

determining a first location associated with a collocated block based on the location of the current video block and a temporal motion vector (MV);

determining whether the first location associated with the collocated block is within the constrained region in the collocated picture, wherein, based on the first location associated with the collocated block being outside the constrained region, the method comprises determining a second location associated with the collocated block using a clipping operation associated with the constrained region in the collocated picture, wherein the clipping operation changes the first location associated with the collocated block based on a boundary of the constrained region;

obtaining, in the collocated picture, the collocated block associated with the current video block at the second location associated with the collocated block;

obtaining an MV associated with the collocated block; and

encoding the current video block based on the MV associated with the collocated block.

18. The method of claim 17 , wherein the current video block is an advanced temporal motion vector prediction (ATMVP) subblock, and the collocated block is a collocated subblock associated with the ATMVP subblock, and wherein the method further comprises:

obtaining an MV associated with the collocated subblock; and

predicting the ATMVP subblock based on the MV associated with the collocated subblock, wherein the current video block is encoded based on the prediction of the ATMVP subblock.

19. The method of claim 17 , wherein the current video block is an advanced temporal motion vector prediction (ATMVP) subblock, and the collocated block is a collocated subblock associated with the ATMVP subblock, and wherein the method further comprises:

obtaining an MV associated with the collocated subblock; and

predicting the ATMVP subblock based on the MV associated with the collocated subblock, wherein the current video block is decoded based on the prediction of the ATMVP subblock.

20. The method of claim 17 , wherein the second location associated with the collocated block is on the boundary of the constrained region.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: VID SCALE, INC.
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068284/0031 →
Continuity (6)
Continuation 16761039
Provisional Application 62690661 · Jun 27, 2018
Provisional Application 62678576 · May 31, 2018
Provisional Application 62623001 · Jan 29, 2018
Provisional Application 62580184 · Nov 1, 2017
Related Publication 20230156215A1 · May 18, 2023
References Cited (54)
US 20050152452A1 · Suzuki · 2005 [cited by applicant]
US 20100316129A1 · Zhao · 2010 [cited by examiner]
US 20130083853A1 · Coban · 2013 [cited by examiner]
US 20130163668A1 · Chen · 2013 [cited by examiner]
US 20130258052A1 · Li et al. · 2013 [cited by applicant]
US 20140044171A1 · Takehara · 2014 [cited by examiner]
US 20140064374A1 · Xiu et al. · 2014 [cited by applicant]
US 20140241436A1 · Laroche et al. · 2014 [cited by applicant]
US 20140355688A1 · Lim et al. · 2014 [cited by applicant]
US 20140376633A1 · Zhang et al. · 2014 [cited by applicant]
US 20150085929A1 · Chen · 2015 [cited by examiner]
US 20150139329A1 · Nakamura et al. · 2015 [cited by applicant]
US 20160050430A1 · Xiu et al. · 2016 [cited by applicant]
US 20160219278A1 · Chen et al. · 2016 [cited by applicant]
US 20160234492A1 · Li et al. · 2016 [cited by applicant]
US 20160330475A1 · Zhou · 2016 [cited by examiner]
US 20160344903A1 · Zhou · 2016 [cited by applicant]
US 20160360210A1 · Xiu et al. · 2016 [cited by applicant]
US 20160366435A1 · Chien et al. · 2016 [cited by applicant]
US 20170223350A1 · Xu et al. · 2017 [cited by applicant]
US 20170332099A1 · Lee et al. · 2017 [cited by applicant]
US 20180007395A1 · Ugur et al. · 2018 [cited by applicant]
US 20180070100A1 · Chen et al. · 2018 [cited by applicant]
US 20180084260A1 · Chien · 2018 [cited by examiner]
US 20180098072A1 · Zhang et al. · 2018 [cited by applicant]
US 20180199052A1 · He · 2018 [cited by examiner]
US 20180249176A1 · Lim · 2018 [cited by examiner]
US 20190182505A1 · Chuang et al. · 2019 [cited by applicant]
US 20190222837A1 · Lee et al. · 2019 [cited by applicant]
US 20190313112A1 · Han · 2019 [cited by examiner]
US 20200045307A1 · Jang · 2020 [cited by applicant]
US 20210044832A1 · Lee et al. · 2021 [cited by applicant]
US 20210092433A1 · Chen et al. · 2021 [cited by applicant]
JP 2014520477A · 2014 [cited by applicant]
JP 2016537839A · 2016 [cited by applicant]
JP 2018506908A · 2018 [cited by applicant]
KR 20170108010A · 2017 [cited by applicant]
WO 2016123081A1 · 2016 [cited by applicant]
WO 2019233423A1 · 2019 [cited by applicant]
“JEM-7.0 Reference Software”, Available at <https://jvet.hhi.fraunhofer.de/svn/svn_HMJEMSoftware/tags/HM-16.6-JEM-7.0>, 1 page. [cited by applicant]
Alshin et al., “AHG6: On Bio Memory Bandwidth”, JVET-D0042, Samsung Electronics Ltd., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 4th Meeting: Chengdu, CN, Oct. 15-21, 2016, pp… [cited by applicant]
Alshina et al., “Known Tools Performance Investigation for Next Generation Video Coding”, VCEG-AZ05, Samsung Electronics, ITU-Telecommunications Standardization Sector, Study Group 16 Question 6, Video Coding Experts Gr… [cited by applicant]
Bross et al., “High Efficiency Video Coding (HEVC) Text Specification Draft 8”, JCTVC-J1003, Editor, Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, 10th Meeting: Stockhol… [cited by applicant]
Chen et al., “Algorithm Description of Joint Exploration Test Model 7 (JEM 7)”, Editors, JVET-G1001-V1, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7th Meeting: Torino, IT, Jul… [cited by applicant]
Chen et al., “Coding Tools Investigation for Next Generation Video Coding”, Qualcomm Incorporated, COM 16-C 806-E, Jan. 2015, pp. 1-7. [cited by applicant]
ITU-T, “Advance Video Coding for Generic Audiovisual Services”, H.264, Series H: Audiovisual and Multimedia Systems, Infrastructure of Audiovisual Services—Coding of Moving Video, Nov. 2007, 564 pages. [cited by applicant]
ITU-T, “Advanced Video Coding for Generic Audiovisual Services”, H.264, Series H: Audiovisual and Multimedia Systems, Infrastructure of Audiovisual Services—Coding of Moving Video, Nov. 2007, 564 pages. [cited by applicant]
Karczewicz et al., “Report of AHG1 on Coding Efficiency Improvements”, VCEG-AZ01, Qualcomm, Samsung, ITU-Telecommunications Standardization Sector, Study Group 16 Question 6, Video Coding Experts Group (VCEG), 52nd Meet… [cited by applicant]
Li et al., “Low-Complexity Merge Candidate Decision for Fast HEVC Encoding”, 2013 IEEE International Conference on Multimedia and Expo Workshops (ICMEW), 2013, 6 pages. [cited by applicant]
Ohm et al., “Report of AHG on Future Video Coding Standardization Challenges”, AHG, ISO/IEC JTC1/SC29/WG11 MPEG2014/M36782, Warsaw, Poland, Jun. 2015, 4 pages. [cited by applicant]
SMPTE, “VC-1 Compressed Video Bitstream Format and Decoding Process”, SMPTE 421M, Feb. 24, 2006, 493 pages. [cited by applicant]
Tourapis et al., “H.264/14496-10 AVC Reference Software Manual”, JVT-AE010, Dolby Laboratories Inc., Fraunhofer-Institute HHI, Microsoft Corporation, Joint Video Team (JVT) of ISO/IEC MPEG & ITU-T VCEG (ISO/IEC JTC1/SC2… [cited by applicant]
Wikipedia, “SIMD”, Available at <https://en.wikipedia.org/wiki/SIMD>, pp. 1-8. [cited by applicant]
Xiu et al., “Description of SDR, HDR and 360° Video Coding Technology Proposal by InterDigital Communications and Dolby Laboratories”, JVET-J0015-V1, InterDigital Communications, Inc., Dolby Laboratories, Inc., Joint Vi… [cited by applicant]