IP Library Granted Patent US 12,309,413
Granted Patent B2
US 12,309,413 · App. 17/490,220 · Granted May 20, 2025

Efficient affine merge motion vector derivation

Inventors: Kai Zhang (San Diego, CA); Li Zhang (San Diego, CA); Hongbin Liu (Beijing, CN); Yue Wang (Beijing, CN)
Assignees: BEJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/513H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,309,413
App. No.
17/490,220
Granted
May 20, 2025
Kind
B2
Abstract

A video processing method for efficient affine merge motion vector derivation is disclosed. In one aspect, a video processing method is provided to include partitioning a current video block into sub-blocks; deriving, for each sub-block, a motion vector, wherein the motion vector for each sub-block is associated with a position for that sub-block according to a position rule; and processing a bitstream representation of the current video block using motion vectors for the sub-blocks.

Claims (50)

1. A method of processing video data, comprising:

determining, for a conversion between a current video block of a video and a bitstream of the video, motion vectors at control points (CPMVs) of the current video block based on a rule; and

performing the conversion between the current video block and the bitstream based on the CPMVs,

wherein performing the conversion comprises:

determining a motion vector for each sub-block of a multiple sub-blocks of the current video block based on the CPMVs and a specific position of a corresponding sub-block;

rounding the motion vector of each subblock to 1/16 fraction accuracy; and

filtering the rounded motion vector of each subblock with a filter to generate a prediction of each subblock,

wherein a maximum accuracy of the filter is 1/16 fractional pel,

wherein the rule specifies that a CPMV candidate of the current video block is derived from motion vectors from top adjacent blocks coded with an affine mode, and the CPMV candidate is used to derive the CPMVs of the current video block, and

wherein the current video block belongs to a current slice, and a spatial neighboring block of the current video block is invalid for determining CPMVs of the current video block in a case that the spatial neighboring block belongs to a slice different from the current slice.

2. The method of claim 1 , wherein the specific position is a center of the corresponding sub-block.

3. The method of claim 1 , wherein the corresponding sub-block has a size M×N and a center of the corresponding sub-block is defined as [(M>>1)+a, (N>>1)+b], wherein M and N are natural numbers and a, b is 0 or −1.

4. The method of claim 1 , wherein the conversion includes encoding the current video block into the bitstream.

5. The method of claim 1 , wherein the conversion includes decoding the current video block from the bitstream.

6. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

determine, for a conversion between a current video block of a video and a bitstream of the video, motion vectors at control points (CPMVs) of the current video block based on a rule; and

perform the conversion between the current video block and the bitstream based on the CPMVs,

wherein performing the conversion comprises:

determining a motion vector for each sub-block of a multiple sub-blocks of the current video block based on the CPMVs and a specific position of a corresponding sub-block;

rounding the motion vector of each subblock to 1/16 fraction accuracy; and

filtering the rounded motion vector of each subblock with a filter to generate a prediction of each subblock,

wherein a maximum accuracy of the filter is 1/16 fractional pel,

wherein the rule specifies that a CPMV candidate of the current video block is derived from motion vectors from top adjacent blocks coded with an affine mode, and the CPMV candidate is used to derive the CPMVs of the current video block, and

wherein the current video block belongs to a current slice, and a spatial neighboring block of the current video block is invalid for determining CPMVs of the current video block in a case that the spatial neighboring block belongs to a slice different from the current slice.

7. The apparatus of claim 6 , wherein the specific position is a center of the corresponding sub-block.

8. The apparatus of claim 6 , wherein the corresponding sub-block has a size M×N and a center of the corresponding sub-block is defined as [(M>>1)+a, (N>>1)+b], wherein M and N are natural numbers and a, b is 0 or −1.

9. The apparatus of claim 6 , wherein the conversion includes encoding the current video block into the bitstream.

10. The apparatus of claim 6 , wherein the conversion includes decoding the current video block from the bitstream.

11. A non-transitory computer-readable storage medium storing instructions that cause a processor to:

determine, for a conversion between a current video block of a video and a bitstream of the video, motion vectors at control points (CPMVs) of the current video block based on a rule; and

perform the conversion between the current video block and the bitstream based on the CPMVs,

wherein performing the conversion comprises:

determining a motion vector for each sub-block of a multiple sub-blocks of the current video block based on the CPMVs and a specific position of a corresponding sub-block;

rounding the motion vector of each subblock to 1/16 fraction accuracy; and

filtering the rounded motion vector of each subblock with a filter to generate a prediction of each subblock,

wherein a maximum accuracy of the filter is 1/16 fractional pel,

wherein the rule specifies that a CPMV candidate of the current video block is derived from motion vectors from top adjacent blocks coded with an affine mode, and the CPMV candidate is used to derive the CPMVs of the current video block, and

wherein the current video block belongs to a current slice, and a spatial neighboring block of the current video block is invalid for determining CPMVs of the current video block in a case that the spatial neighboring block belongs to a slice different from the current slice.

12. A non-transitory computer-readable recording medium storing a bitstream which is generated by a method performed by a video processing apparatus, wherein the method comprises:

determining, for a conversion between a current video block of a video and the bitstream of the video, motion vectors at control points (CPMVs) of the current video block based on a rule; and

generating the bitstream from the current video block based on the determining,

wherein generating the bitstream from the current video block comprises:

determining a motion vector for each sub-block of a multiple sub-blocks of the current video block based on the CPMVs and a specific position of a corresponding sub-block;

rounding the motion vector of each subblock to 1/16 fraction accuracy; and

filtering the rounded motion vector of each subblock with a filter to generate a prediction of each subblock,

wherein a maximum accuracy of the filter is 1/16 fractional pel,

wherein the rule specifies that a CPMV candidate of the current video block is derived from motion vectors from top adjacent blocks coded with an affine mode, and the CPMV candidate is used to derive the CPMVs of the current video block, and

wherein the current video block belongs to a current slice, and a spatial neighboring block of the current video block is invalid for determining CPMVs of the current video block in a case that the spatial neighboring block belongs to a slice different from the current slice.

13. The non-transitory computer-readable storage medium of claim 11 , wherein the specific position is a center of the corresponding sub-block.

14. The non-transitory computer-readable storage medium of claim 11 , wherein the corresponding sub-block has a size M×N and a center of the corresponding sub-block is defined as [(M>>1)+a, (N>>1)+b], wherein M and N are natural numbers and a, b is 0 or −1.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2021
From: ZHANG, KAI; ZHANG, LI
To: BYTEDANCE INC.
Reel/Frame 057657/0182 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2021
From: LIU, HONGBIN; WANG, YUE
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 057657/0272 →
Priority Claims (2)
WO PCT/CN2018/093943 · Jul 1, 2018 · international
WO PCT/CN2018/095568 · Jul 13, 2018 · international
Continuity (3)
Continuation 17090122 · Nov 5, 2020
Continuation PCTIB2019055592 · Jul 1, 2019
Related Publication 20220046267A1 · Feb 10, 2022
References Cited (57)
US 10448010B2 · Chen et al. · 2019 [cited by applicant]
US 10560712B2 · Zou et al. · 2020 [cited by applicant]
US 10757417B2 · Zhang et al. · 2020 [cited by applicant]
US 10778999B2 · Li et al. · 2020 [cited by applicant]
US 10841609B1 · Liu et al. · 2020 [cited by applicant]
US 20110080954A1 · Bossen et al. · 2011 [cited by applicant]
US 20140003527A1 · Tourapis · 2014 [cited by applicant]
US 20170214932A1 · Huang · 2017 [cited by applicant]
US 20180098063A1 · Chen et al. · 2018 [cited by applicant]
US 20180359483A1 · Chen · 2018 [cited by examiner]
US 20190387250A1 · Boyce et al. · 2019 [cited by applicant]
US 20200045310A1 · Chen et al. · 2020 [cited by applicant]
US 20200145688A1 · Zou et al. · 2020 [cited by applicant]
US 20200213594A1 · Liu et al. · 2020 [cited by applicant]
US 20200213612A1 · Liu et al. · 2020 [cited by applicant]
US 20200359029A1 · Liu et al. · 2020 [cited by applicant]
US 20200382771A1 · Liu et al. · 2020 [cited by applicant]
US 20200382795A1 · Zhang et al. · 2020 [cited by applicant]
US 20200396453A1 · Zhang et al. · 2020 [cited by applicant]
US 20200396465A1 · Zhang et al. · 2020 [cited by applicant]
US 20210058637A1 · Zhang et al. · 2021 [cited by applicant]
CN 104935938A · 2015 [cited by applicant]
CN 105794210A · 2016 [cited by applicant]
CN 106303543A · 2017 [cited by applicant]
CN 106537915A · 2017 [cited by applicant]
CN 106559669A · 2017 [cited by applicant]
CN 107113446A · 2017 [cited by applicant]
CN 107211157A · 2017 [cited by applicant]
CN 107925758A · 2018 [cited by applicant]
CN 108141582A · 2018 [cited by applicant]
JP 2016536817A · 2016 [cited by applicant]
JP 2018511234A · 2018 [cited by applicant]
TW 202005389A · 2020 [cited by applicant]
WO 2017048345A1 · 2017 [cited by applicant]
WO 2017118411A1 · 2017 [cited by applicant]
WO 2018110180A1 · 2018 [cited by applicant]
Non Final Office Action from U.S. Appl. No. 17/090,122 dated Jan. 25, 2021. [cited by applicant]
Final Office Action from U.S. Appl. No. 17/090,122 dated Jun. 3, 2021. [cited by applicant]
Alshina et al. “Performance of JEM1.0 Tools Analysis by Sam sung,” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP3 and ISO/IEC JTC 1/SC 29/WG 11, 2nd Meeting, San Diego, USA, Feb. 20-26, 2016, document JVET-B0022… [cited by applicant]
Chen et al. “Algorithm Description of Joint Exploration Test Model 7 (JEM 7),” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7th Meeting: Torino, IT, Jul. 13-21, 2017, document J… [cited by applicant]
Guo et al. “CE2: Overlapped Block Motion Compensation for Geometry Partition Block,” Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, 4th Meeting, Daegu, KR, Jan. 20-28, 20… [cited by applicant]
Hsu et al. “Description of Core Experiment 10: Combined and Multihypothesis Prediction,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 10th Meeting, San Diego, US, Apr. 10-20, 2018, … [cited by applicant]
ITU-T H.265 “High efficiency video coding” Series H: Audiovisual and Multimedia Systems Infrastructure of audiovisual services—Coding of moving video,Telecommunication Standardization Sector of ITU, Available at address… [cited by applicant]
JEM-7.0: https://jvet.hhi.fraunhofer.de/svn/svn_HMJEMSoftware/tags/ HM-16.6-JEM-7.0. [cited by applicant]
Lee et al. “OBMC for IVC+,” International Organisation for Standardisation, ISO/IEC JTC1/SC29/WG11, Coding of Moving Pictures and Audio, Oct. 2016, Chengdu, China. [cited by applicant]
Li et al. “CE2-Related: Simplifications of Interweaved Affine Mode,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting, Geneva, CH, Mar. 19-27, 2019, document JVET-N0399, 20… [cited by applicant]
Zhang et al. “Merge Mode for Deformable Block Motion Information Derivation,” IEEE Transactions on Circuits and Systems for Video Technology, Nov. 2017, 27(11):2437-2449. [cited by applicant]
Zhang et al. “CE4-Related: Interweaved Prediction for Affine Motion Compensation,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 11th Meeting, Ljubljana, SI, Jul. 10-18, 2018. docume… [cited by applicant]
Zhang et al. “CE10: Interweaved Prediction for Affine Motion Compensation (Test 10.5.1 and Test 10.5.2),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macao, CN, Oct. … [cited by applicant]
Zhang et al. “Non-CE2: Interweaved Prediction for Affine Motion Compensation,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting, Marrakech, MA, Jan. 9-18, 2019, document JV… [cited by applicant]
International Search Report and Written Opinion from PCT/IB2019/055566 dated Sep. 26, 2019 (22 pages). [cited by applicant]
International Search Report and Written Opinion from PCT/IB2019/055592 dated Nov. 21, 2019 (18 pages). [cited by applicant]
Chen et al. “Algorithm Description of Joint Exploration Test Model 4,” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 4th Meeting: Chengdu, CN, Oct. 15-21, 2016, document JVET-D10… [cited by applicant]
Chen et al. “Description of SDR, HDR and 360 degree Video Coding Technology Proposed by Qualcomm and Technicolor—Low and High Complexity Versions,” Joint Video Explorations Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JT… [cited by applicant]
Communication Pursuant to Rule 164(2)(b) and Article 94(3) EPC from European Patent Application No. 19740074.0 dated Jan. 26, 2024 (10 pages). [cited by applicant]
First Office Action for Chinese Application No. 20221066664.2, mailed on Dec. 27, 2024, 23 pages. [cited by applicant]
Japanese Office Action from Japanese Patent Application No. 2023-187705 dated Apr. 1, 2025, 7 pages. [cited by applicant]