IP Library › Granted Patent US 12,432,372
Granted Patent B2
US 12,432,372 · App. 18/121,438 · Granted Sep 30, 2025

Systems and methods for template matching for adaptive MVD resolution

Inventors: Liang Zhao (Palo Alto, CA); Xin Zhao (Palo Alto, CA); Shan Liu (Palo Alto, CA)
Assignee: TENCENT AMERICA LLC
H04N19/513H04N19/105H04N19/159H04N19/176H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,432,372
App. No.
18/121,438
Granted
Sep 30, 2025
Kind
B2
Abstract

The various implementations described herein include methods and systems for coding video. The methods include receiving a motion vector difference (MVD) of a video block from the video stream; in response to a determination that an adaptive MVD resolution mode of an inter prediction mode is signaled, searching for a first template of a prediction video block for the video block, and the first/second template is the neighboring reconstructed/predicted samples of the prediction video block/current block, and the prediction video block is a reconstructed/predicted forward or backward video block of the video block; locating the first template of the prediction video block that is a best match for a second template of the video block; refining a motion vector (MV) of the video block based on at least the second template, the located first template, and the MVD; and reconstructing/processing the video block based on at least the refined MV.

Claims (38)

1. A method of decoding a video stream performed at a computing system having memory and control circuitry, the method comprising:

determining, based on a value of an inter-prediction syntax element from the video stream, whether an adaptive motion vector difference (MVD) resolution mode is signaled, the adaptive MVD resolution mode being an inter-prediction mode with adaptive motion vector difference (MVD) pixel resolution;

receiving a motion vector difference (MVD) of a video block from the video stream;

in response to a determination that the adaptive MVD resolution mode is signaled, searching for a first template of a prediction video block for the video block, wherein the first template is neighboring reconstructed samples of the prediction video block, and the prediction video block is a reconstructed forward or backward video block of the video block;

locating the first template of the prediction video block that is a best match for a second template of the video block, the second template being neighboring reconstructed samples of the video block corresponding to a temporally collocated template of the first template;

refining a motion vector (MV) of the video block based on at least the second template, the located first template, and the MVD; and

reconstructing the video block based on at least the refined MV.

2. The method of claim 1 , wherein searching for the first template of the prediction video block for the video block comprises determining a search area size based on a magnitude of the MVD and searching based on the search area size.

3. The method of claim 2 , wherein determining the search area size based on the magnitude of the MVD comprises increasing the search area as the magnitude of the MVD increases or keeping the search area unchanged as the magnitude of the MVD increases.

4. The method of claim 2 , wherein determining the search area size based on the magnitude of the MVD comprises determining a same search area when the MVD is in the same MV class.

5. The method of claim 2 , wherein determining the search area size based on the magnitude of the MVD comprises determining a same search area when the magnitude of the MVD is equal to or greater than a threshold.

6. The method of claim 1 , wherein refining the MV of the video block comprises determining a refining granularity of the MV based on the magnitude of the MVD.

7. The method of claim 6 , wherein determining the refining granularity of the MV based on the magnitude of the MVD comprises implementing a fractional precision MV refinement only when the magnitude of the MVD is equal to or less than a threshold.

8. The method of claim 1 , wherein refining the MV of the video block comprises restricting the MV to one or more pre-determined directions during refining.

9. The method of claim 1 , wherein searching for the first template of the prediction video block for the video block comprises determining a search direction based on a direction of the MVD and searching based on the search direction.

10. The method of claim 1 , wherein before searching, determining, based on a second syntax element from the video stream, whether a template matching mode is signaled for the one or more video blocks, and conducting searching in response to a determination that the template matching mode is signaled.

11. The method of claim 10 , wherein the second syntax element is signaled in one or more of sequence level, frame level, and/or slice level.

12. The method of claim 10 , wherein when the adaptive MVD resolution mode is signaled, a finest allowed MVD resolution depends on whether the template matching mode is signaled.

13. A computing system comprising a memory for storing computer instructions and control circuitry in communication with the memory, wherein the control circuitry, when executing the computer instructions, is configured to cause the computing system to perform a method of decoding a video stream, the method including:

determining, based on a value of an inter-prediction syntax element from the video stream, whether an adaptive motion vector difference (MVD) resolution mode is signaled, the adaptive MVD resolution mode being an inter-prediction mode with adaptive motion vector difference (MVD) pixel resolution;

receiving a motion vector difference (MVD) of a video block from the video stream;

in response to a determination that the adaptive MVD resolution mode is signaled, searching for a first template of a prediction video block for the video block, wherein the first template is neighboring reconstructed samples of the prediction video block, and the prediction video block is a reconstructed forward or backward video block of the video block;

locating the first template of the prediction video block that is a best match for a second template of the video block, the second template being neighboring reconstructed samples of the video block corresponding to a temporally collocated template of the first template;

refining a motion vector (MV) of the video block based on at least the second template, the located first template, and the MVD; and

reconstructing the video block based on at least the refined MV.

14. The computing system of claim 13 , wherein searching for the first template of the prediction video block for the video block comprises determining a search area size based on a magnitude of the MVD and searching based on the search area size.

15. The computing system of claim 14 , wherein determining the search area size based on the magnitude of the MVD comprises increasing the search area as the magnitude of the MVD increases or keeping the search area unchanged as the magnitude of the MVD increases.

16. The computing system of claim 14 , wherein determining the search area size based on the magnitude of the MVD comprises determining a same search area when the MVD is in the same MV class.

17. The computing system of claim 14 , wherein determining the search area size based on the magnitude of the MVD comprises determining a same search area when the magnitude of the MVD is equal to or greater than a threshold.

18. The computing system of claim 13 , wherein refining the MV of the video block comprises determining a refining granularity of the MV based on the magnitude of the MVD.

19. The computing system of claim 13 , wherein searching for the first template of the prediction video block for the video block comprises determining a search direction based on a direction of the MVD and searching based on the search direction.

20. A non-transitory computer readable medium for storing computer instructions, the computer instructions, when executed by control circuitry of a computing system, cause the computing system to perform a method of decoding a video stream including:

determining, based on a value of an inter-prediction syntax element from the video stream, whether an adaptive motion vector difference (MVD) resolution mode is signaled, the adaptive MVD resolution mode being an inter-prediction mode with adaptive motion vector difference (MVD) pixel resolution;

receiving a motion vector difference (MVD) of a video block from the video stream;

in response to a determination that the adaptive MVD resolution mode is signaled, searching for a first template of a prediction video block for the video block, wherein the first template is neighboring reconstructed samples of the prediction video block, and the prediction video block is a reconstructed forward or backward video block of the video block;

locating the first template of the prediction video block that is a best match for a second template of the video block, the second template being neighboring reconstructed samples of the video block corresponding to a temporally collocated template of the first template;

refining a motion vector (MV) of the video block based on at least the second template, the located first template, and the MVD; and

reconstructing the video block based on at least the refined MV.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2023
From: ZHAO, LIANG; LIU, SHAN; ZHAO, XIN
To: TENCENT AMERICA LLC
Reel/Frame 063327/0130 →
Continuity (2)
Provisional Application 63320488 · Mar 16, 2022
Related Publication 20230300363A1 · Sep 21, 2023
References Cited (25)
US 11172216B1 · Zhang · 2021 [cited by examiner]
US 11979596B2 · Zhao · 2024 [cited by examiner]
US 20120219063A1 · Kim et al. · 2012 [cited by applicant]
US 20160337662A1 · Pang et al. · 2016 [cited by applicant]
US 20210281876A1 · Zhang · 2021 [cited by examiner]
US 20210297679A1 · Zhang · 2021 [cited by examiner]
US 20210321121A1 · Zhang · 2021 [cited by examiner]
US 20210385461A1 · Liu et al. · 2021 [cited by applicant]
US 20220103815A1 · Zhang · 2022 [cited by examiner]
WO WO2018121506A1 · 2018 [cited by applicant]
WO WO2023076787A1 · 2023 [cited by applicant]
WO WO2023137414A2 · 2023 [cited by applicant]
WO WO2023137414A3 · 2023 [cited by applicant]
Tencent Technology, ISRWO, PCT/US2023/015310, Jun. 14, 2023, 10 pgs. [cited by applicant]
Benjamin Bross et al., “Versatile Video Coding Editorial Refinements on Draft 10”, Document: JVET-T2001-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 20th Meeting, by teleconference, O… [cited by applicant]
Elliott Karpilovsky et al., “Proposal: New Inter Modes for AV2”, Document:CWG-8018 v1, Alliance for Open Media, Codec Working Group, Feb. 24, 2021, 6 pgs. [cited by applicant]
Leo Zhao et al., “Advanced Motion Vector Difference Coding”, Document: CWG-B092, Alliance for Open Media, Codec Working Group, Nov. 24, 2021, 7 pgs. [cited by applicant]
Leo Zhao et al., “Improved Adaptive MVD Resolution”, Document: CWG-C011, Alliance for Open Media, Codec Working Group, Feb. 9, 2022, 7 pgs. [cited by applicant]
Lester (Keng-Shih) Lu et al., “Optical Flow Motion Vector Refinement for AV2”, Document: CWG-B041_v3, Alliance for Open Media Codec Working Group, Google, Sep. 20, 2021, 11 pgs. [cited by applicant]
Peter de Rivaz et al., “AV1 Bitstream & Decoding Process Specification”, Jan. 8, 2019, 681 pgs. [cited by applicant]
Xin Zhao et al., “Tool Description for AV1 and Libaom”, Document: CWG-B078_v1, Alliance for Open Media Codec Working Group, Oct. 4, 2021, 41 pgs. [cited by applicant]
Yue Chen et al., “An Overview of Core Coding Tools in the AV1 Video Codec”, 2018 IEEE, 5 pgs. [cited by applicant]
Tencent America LLC, Extended European Search Report, EP Patent Application No. 23771379.7, Jun. 2, 2025, 10 pgs. [cited by applicant]
Liang Zhao et al., “Advanced Motion Vector Difference Coding Beyond AV1”, 2022 IEEE International Conference on Image Processing (ICIP), Oct. 2022, 5 pgs. [cited by applicant]
Muhammed Coban et al., “Algorithm Description of Enhanced Compression Model 3 (ECM 3)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 23rd Meeting, by teleconference, Jul. 7-16, 2021, Docu… [cited by applicant]