IP Library › Granted Patent US 12,726,649
Granted Patent B2
US 12,726,649 · App. 18/503,360 · Granted Sep 1, 2026

Efficient affine merge motion vector derivation

Inventors: Kai Zhang (San Diego, CA); Li Zhang (San Diego, CA); Hongbin Liu (Beijing, CN); Yue Wang (Beijing, CN)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/513H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,726,649
App. No.
18/503,360
Granted
Sep 1, 2026
Kind
B2
Abstract

A video processing method for efficient affine merge motion vector derivation is disclosed. In one aspect, a video processing method is provided to include partitioning a current video block into sub-blocks; deriving, for each sub-block, a motion vector, where the motion vector for each sub-block is associated with a position for that sub-block according to a position rule; and processing a bitstream representation of the current video block using motion vectors for the sub-blocks.

Claims (39)

1 . A method of coding video data, comprising:

determining, for a conversion between a current video block of a video and a bitstream of the video, motion vectors at control points (CPMV) of the current video block based on a rule, wherein the rule specifies to exclude using of a non-adjacent neighboring block from one or more neighboring blocks of the current video block; and

performing the conversion between the current video block and the bitstream based on the motion vectors,

wherein the rule further specifies to exclude using of an invalid neighboring block from the one or more neighboring blocks based on positions of the one or more neighboring blocks,

wherein the current video block belongs to a current tile, and a neighboring block of the one or more neighboring blocks is invalid in a case that the neighboring block belongs to a tile different from the current tile, and

wherein the rule specifies that for a merge affine mode, one CPMV candidate of the current video block is derived from motion vectors from top-left adjacent blocks coded with an affine mode without considering left adjacent blocks coded with the affine mode, and the one CPMV candidate is used to derive the CPMVs of the current video block.

2 . The method of claim 1 , wherein the current video block belongs to a current slice, and a neighboring block of the one or more neighboring blocks is invalid in a case that the neighboring block belongs to a slice different from the current slice.

3 . The method of claim 1 , wherein the current video block includes multiple sub-blocks, and performing the conversion comprises determining a motion vector for each sub-block of the multiple sub-blocks based on the CPMV and a specific position of the corresponding sub-block.

4 . The method of claim 3 , wherein the specific position is a center of the corresponding sub-block.

5 . The method of claim 4 , wherein the corresponding sub-block has a size M×N and the center is defined as [(M>>1)+a, (N>>1)+b], wherein M and N are natural numbers and a, b is 0 or −1.

6 . The method of claim 1 , wherein the conversion includes encoding the current video block into the bitstream.

7 . The method of claim 1 , wherein the conversion includes decoding the current video block from the bitstream.

8 . An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

determine, for a conversion between a current video block of a video and a bitstream of the video, motion vectors at control points (CPMV) of the current video block based on a rule, wherein the rule specifies to exclude using of a non-adjacent neighboring block from one or more neighboring blocks of the current video block; and

perform the conversion between the current video block and the bitstream based on the motion vectors,

wherein the rule further specifies to exclude using of an invalid neighboring block from the one or more neighboring blocks based on positions of the one or more neighboring blocks,

wherein the current video block belongs to a current tile, and a neighboring block of the one or more neighboring blocks is invalid in a case that the neighboring block belongs to a tile different from the current tile, and

wherein the rule specifies that for a merge affine mode, one CPMV candidate of the current video block is derived from motion vectors from top-left adjacent blocks coded with an affine mode without considering left adjacent blocks coded with the affine mode, and the one CPMV candidate is used to derive the CPMVs of the current video block.

9 . The apparatus of claim 8 , wherein the current video block belongs to a current slice, and a neighboring block of the one or more neighboring blocks is invalid in a case that the neighboring block belongs to a slice different from the current slice.

10 . The apparatus of claim 8 , wherein the current video block includes multiple sub-blocks, and performing the conversion comprises determining a motion vector for each sub-block of the multiple sub-blocks based on the CPMV and a specific position of the corresponding sub-block.

11 . The apparatus of claim 10 , wherein the specific position is a center of the corresponding sub-block.

12 . The apparatus of claim 11 , wherein the corresponding sub-block has a size M×N and the center is defined as [(M>>1)+a, (N>>1)+b], wherein M and N are natural numbers and a, b is 0 or −1.

13 . A non-transitory computer-readable storage medium storing instructions that cause a processor to:

determine, for a conversion between a current video block of a video and a bitstream of the video, motion vectors at control points (CPMV) of the current video block based on a rule, wherein the rule specifies to exclude using of non-adjacent neighboring block from one or more neighboring blocks of the current video block; and

perform the conversion between the current video block and the bitstream based on the motion vectors,

wherein the rule further specifies to exclude using of an invalid neighboring block from the one or more neighboring blocks based on positions of the one or more neighboring blocks,

wherein the current video block belongs to a current tile, and a neighboring block of the one or more neighboring blocks is invalid in a case that the neighboring block belongs to a tile different from the current tile, and

wherein the rule specifies that for a merge affine mode, one CPMV candidate of the current video block is derived from motion vectors from top-left adjacent blocks coded with an affine mode without considering left adjacent blocks coded with the affine mode, and the one CPMV candidate is used to derive the CPMVs of the current video block.

14 . The non-transitory computer-readable storage medium of claim 13 , wherein the current video block includes multiple sub-blocks, and performing the conversion comprises determining a motion vector for each sub-block of the multiple sub-blocks based on the CPMV and a specific position of the corresponding sub-block.

15 . The non-transitory computer-readable storage medium of claim 14 , wherein the specific position is a center of the corresponding sub-block.

16 . A method for storing bitstream of a video, comprising:

determining, for a current video block of a video, motion vectors at control points (CPMV) of the current video block based on a rule, wherein the rule specifies to exclude using of non-adjacent neighboring block from one or more neighboring blocks of the current video block;

generating the bitstream from the current video block based on the determining, and

storing the bitstream in a non-transitory computer-readable recording medium,

wherein the rule further specifies to exclude using of an invalid neighboring block from the one or more neighboring blocks based on positions of the one or more neighboring blocks,

wherein the current video block belongs to a current tile, and a neighboring block of the one or more neighboring blocks is invalid in a case that the neighboring block belongs to a tile different from the current tile, and

wherein the rule specifies that for a merge affine mode, one CPMV candidate of the current video block is derived from motion vectors from top-left adjacent blocks coded with an affine mode without considering left adjacent blocks coded with the affine mode, and the one CPMV candidate is used to derive the CPMVs of the current video block.

17 . The method of claim 16 , wherein the current video block includes multiple sub-blocks, and the generating comprises determining a motion vector for each sub-block of the multiple sub-blocks based on the CPMV and a specific position of the corresponding sub-block.

18 . The method of claim 17 , wherein the specific position is a center of the corresponding sub-block.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2023
From: ZHANG, KAI; ZHANG, LI
To: BYTEDANCE INC.
Reel/Frame 065638/0820 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2023
From: LIU, HONGBIN; WANG, YUE
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 065638/0859 →
Continuity (4)
Continuation 17490220 · Sep 30, 2021
Continuation 17090122 · Nov 5, 2020
Continuation PCTIB2019055592 · Jul 1, 2019
Related Publication 20240098295A1 · Mar 21, 2024
References Cited (79)
US 10448010B2 · Chen · 2019 [cited by applicant]
US 10560712B2 · Zou · 2020 [cited by applicant]
US 10757417B2 · Zhang · 2020 [cited by applicant]
US 10778999B2 · Li · 2020 [cited by applicant]
US 10841609B1 · Liu · 2020 [cited by applicant]
US 12309413B2 · Zhang et al. · 2025 [cited by applicant]
US 20030072374A1 · Sohm · 2003 [cited by applicant]
US 20110080954A1 · Bossen · 2011 [cited by applicant]
US 20140003527A1 · Tourapis · 2014 [cited by applicant]
US 20150085930A1 · Zhang · 2015 [cited by applicant]
US 20170214932A1 · Huang · 2017 [cited by applicant]
US 20170332095A1 · Zou · 2017 [cited by applicant]
US 20180098063A1 · Chen · 2018 [cited by applicant]
US 20180359483A1 · Chen · 2018 [cited by examiner]
US 20190387250A1 · Boyce · 2019 [cited by applicant]
US 20200045310A1 · Chen · 2020 [cited by applicant]
US 20200145688A1 · Zou · 2020 [cited by applicant]
US 20200213594A1 · Liu · 2020 [cited by applicant]
US 20200213612A1 · Liu · 2020 [cited by applicant]
US 20200359029A1 · Liu · 2020 [cited by applicant]
US 20200382771A1 · Liu · 2020 [cited by applicant]
US 20200382795A1 · Zhang · 2020 [cited by applicant]
US 20200396453A1 · Zhang · 2020 [cited by applicant]
US 20200396465A1 · Zhang · 2020 [cited by applicant]
US 20210058637A1 · Zhang · 2021 [cited by applicant]
CN 104322070A · 2015 [cited by applicant]
CN 104935938A · 2015 [cited by applicant]
CN 105794210A · 2016 [cited by applicant]
CN 106303543A · 2017 [cited by applicant]
CN 106537915A · 2017 [cited by applicant]
CN 106559669A · 2017 [cited by applicant]
CN 107113424A · 2017 [cited by applicant]
CN 107113446A · 2017 [cited by applicant]
CN 107211157A · 2017 [cited by applicant]
CN 107925758A · 2018 [cited by applicant]
CN 107948657A · 2018 [cited by applicant]
CN 108141582A · 2018 [cited by applicant]
CN 108141589A · 2018 [cited by applicant]
IN 538444A · 2024 [cited by applicant]
JP 2016536817A · 2016 [cited by applicant]
JP 2018511234A · 2018 [cited by applicant]
TW 202005389A · 2020 [cited by applicant]
WO 2017048345A1 · 2017 [cited by applicant]
WO 2017118411A1 · 2017 [cited by applicant]
WO 2017148345A1 · 2017 [cited by applicant]
WO 2017164441A1 · 2017 [cited by applicant]
WO 2018065397A2 · 2018 [cited by applicant]
WO 2018110180A1 · 2018 [cited by applicant]
Alshina et al., “Performance of JEM1.0 tools analysis by Samsung,” Joint Video Exploration Team (JVET) of ITU-T SG WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 2nd Meeting: San Diego, USA, Feb. 20-26, 2016, Document: JVET-B0022,… [cited by applicant]
ITU-T: “Series H: Audiovisual and Multimedia Systems Infrastructure of Audiovisual Services—Coding of moving Video,” Versatile Video Coding, Recommendation ITU-T H.266, Apr. 2022, 536 Pages, [Retrieved on Feb. 28, 2023]… [cited by applicant]
Notice of Reasons for Refusal for Japanese Patent Application No. 2023187705, mailed Oct. 1, 2024, 8 pages. [cited by applicant]
ITU-T H.265 “High efficiency video coding” Series H: Audiovisual and Multimedia Systems Infrastructure of audiovisual services—Coding of moving video, Telecommunication Standardization Sector of ITU, Available at addres… [cited by applicant]
Chen et al. “Algorithm Description of Joint Exploration Test Model 7 (JEM 7),” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7th Meeting: Torino, IT, Jul. 13-21, 2017, document J… [cited by applicant]
JEM-7.0: https://jvet.hhi.fraunhofer.de/svn/svn_HMJEMSoftware/tags/ HM-16.6-JEM-7.0, Sep. 29, 2021. [cited by applicant]
Zhang et al. “Merge Mode for Deformable Block Motion Information Derivation,” IEEE Transactions on Circuits and Systems for Video Technology, Nov. 2017, 27(11):2437-2449. [cited by applicant]
Alshina et al. “Performance of JEM1.0 Tools Analysis by Sam sung,” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP3 and ISO/IEC JTC 1/SC 29/WG 11, 2nd Meeting, San Diego, USA, Feb. 20-26, 2016, document JVET-B0022… [cited by applicant]
Hsu et al. “Description of Core Experiment 10: Combined and Multihypothesis Prediction,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 10th Meeting, San Diego, US, Apr. 10-20, 2018, … [cited by applicant]
Guo et al. “CE2: Overlapped Block Motion Compensation for Geometry Partition Block,” Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, 4th Meeting, Daegu, KR, Jan. 20-28, 20… [cited by applicant]
Hsu et al. “Description of Core Experiment 10: Combined and Multihypothesis Prediction,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 1oth Meeting, San Diego, US, Apr. 10-20, 2018, … [cited by applicant]
Zhang et al. “CE4-Related: Interweaved Prediction for Affine Motion Compensation,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IECJTC 1/SC 29/WG 11, 11th Meeting, Ljubljana, SI, Jul. 10-18, 2018. ocument… [cited by applicant]
Zhang et al. “CE10: Interweaved Prediction for Affine Motion Compensation (Test 10.5.1 and Test 10.5.2),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macao, CN, Oct. … [cited by applicant]
Zhang et al. “Non-CE2: Interweaved Prediction for Affine Motion Compensation,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting, Marrakech, MA, Jan. 9-18, 2019, document JV… [cited by applicant]
Li et al. “CE2-Related: Simplifications of Interweaved Affine Mode,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting, Geneva, CH, Mar. 19-27, 2019, document JVET-N0399, 20… [cited by applicant]
Lee et al. “OBMC for IVC+,” International Organisation for Standardisation, ISO/IEC JTC1/SC29/WG11, Coding of Moving Pictures and Audio, Oct. 2016, Chengdu, China. [cited by applicant]
Chen et al. “Algorithm Description of Joint Exploration Test Model 4,” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 4th Meeting: Chengdu, CN, Oct. 15-21, 2016, document JVET-D10… [cited by applicant]
Document: JVET-J0021-v5, Chen, Y., et al., “Description of SDR, HDR and 360 video coding technology proposal by Qualcomm and Technicolor—low and high complexity versions,” Joint Video Exploration Team (JVET) of ITU-T SG… [cited by applicant]
Non Final Office Action from U.S. Appl. No. 17/090,122 dated Jan. 25, 2021. [cited by applicant]
Final Office Action from U.S. Appl. No. 17/090,122 dated Jun. 3, 2021. [cited by applicant]
International Search Report and Written Opinion from PCT/182019/055566 dated Sep. 26, 2019 (22 pages). [cited by applicant]
International Search Report and Written Opinion from PCT/182019/055592 dated Nov. 21, 2019 (18 pages). [cited by applicant]
Non Final Office Action from U.S. Appl. No. 17/490,220 dated Sep. 2, 2022. [cited by applicant]
Final Office Action from U.S. Appl. No. 17/490,220 dated Mar. 16, 2023. [cited by applicant]
Non Final Office Action from U.S. Appl. No. 17/490,220 dated Sep. 14, 2023. [cited by applicant]
Office Action from European Application No. 19740074.0 dated Jan. 26, 2024. [cited by applicant]
Decision of Refusal for Japanese Application No. 2023-187705, mailed Apr. 1, 2025, 7 pages. [cited by applicant]
Second Examination Opinion Notice for Chinese Patent Application No. 202210066664.2, mailed on May 28, 2025, 22 pages. [cited by applicant]
First Office Action for Chinese Application No. 20221066664.2, mailed on Dec. 27, 2024, 23 pages. [cited by applicant]
Document: JVET-J0057, Chen, X., et al., “DMVR extension based on template matching,” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 10th Meeting: San Diego, US, Apr. 10-20, 2018, 3… [cited by applicant]
Chinese Notice of Allowance from Chinese Patent Application No. 202210066664.2 dated Sep. 1, 2025, 6 pages. [cited by applicant]