IP Library › Granted Patent US 11,546,601
Granted Patent B2
US 11,546,601 · App. 17/207,060 · Granted Jan 3, 2023

Utilization of non-sub block spatial-temporal motion vector prediction in inter mode

Inventors: Hongbin Liu (Beijing, CN); Li Zhang (San Diego, CA); Kai Zhang (San Diego, CA); Yue Wang (Beijing, CN)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/137H04N19/105H04N19/176H04N19/30H04N19/503H04N19/513
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,546,601
App. No.
17/207,060
Granted
Jan 3, 2023
Kind
B2
Abstract

A method of video processing includes determining, for a current block, at least one motion candidate list; and performing a conversion between the current block and a bitstream representation of the current block using the at least one motion candidate list, the at least one motion candidate list including at least one motion candidate derived from a set of neighboring blocks including one or more spatial and temporal neighboring blocks.

Claims (39)

1. A method of video processing, comprising:

determining, for a current block of a video, at least one motion candidate list, wherein the at least one motion candidate list comprises at least one motion candidate each of which is derived from a plurality of spatial and temporal neighboring blocks, wherein the at least one motion candidate is calculated based on a plurality of motion vectors (MV) associated with the plurality of spatial and temporal neighboring blocks and is added to the at least one motion candidate list, and wherein multiple motion candidates are derived and added to the at least one motion candidate list;

scaling the plurality of MVs associated with the plurality of spatial and temporal neighboring blocks to a target reference picture, wherein two or more MVs are scaled to a same reference picture or scaled to a reference picture closest to a current picture, and

wherein, for a spatial temporal motion vector prediction (STMVP) of the current block, a collocated picture as used in temporal motion vector prediction (TMVP) and/or in advanced temporal motion vector prediction (ATMVP), is used as the target reference picture, or

wherein the target reference picture is determined by a first available merge candidate;

deriving a motion vector predictor of the current block by using the scaled plurality of MVs; and

performing a conversion between the current block and a bitstream of the video using the at least one motion candidate list.

2. The method of claim 1 , wherein the motion candidate list is an advanced motion vector prediction (AMVP) list.

3. The method of claim 1 , further comprising:

applying a linear function to the plurality of MVs associated with the plurality of spatial and temporal neighboring blocks to derive a motion vector predictor of the current block.

4. The method of claim 3 , wherein an average of the plurality of MVs is defined as the motion vector predictor of the current block.

5. The method of claim 3 , wherein different weighs are applied to different ones of the plurality of MVs to derive the motion vector predictor of the current block.

6. The method of claim 1 , wherein for deriving each of the multiple motion candidates, different spatial or temporal blocks are utilized.

7. The method of claim 1 , wherein the two or more MVs used to derive the motion candidate refer to a same reference picture.

8. The method of claim 1 , wherein the two or more MVs used to derive the motion candidate refer to reference pictures in a same reference list.

9. The method of claim 1 , wherein the two or more MVs used to derive the motion candidate refer to a reference picture with a same reference index in a same reference list.

10. The method of claim 1 , wherein the two or more MVs used to derive the motion candidate refer to different reference pictures.

11. The method of claim 1 , wherein the collocated picture is signaled at a slice header.

12. The method of claim 1 , wherein if a reference picture of a neighboring block is different from the collocated picture, the MV of the neighboring block is scaled by using a HEVC temporal MV scaling scheme and is used in the STMVP of the current block.

13. The method of claim 1 , wherein the target reference picture is determined as the reference picture closest to the current picture.

14. The method of claim 1 , wherein the conversion comprises at least one of encoding the current block into the bitstream and decoding the current block from the bitstream.

15. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

determine, for a current block of a video, at least one motion candidate list, wherein the at least one motion candidate list comprises at least one motion candidate each of which is derived from a plurality of spatial and temporal neighboring blocks, wherein the at least one motion candidate is calculated based on a plurality of motion vectors (MV) associated with the plurality of spatial and temporal neighboring blocks and is added to the at least one motion candidate list, and wherein multiple motion candidates are derived and added to the at least one motion candidate list;

scale the plurality of MVs associated with the plurality of spatial and temporal neighboring blocks to a target reference picture, wherein two or more MVs are scaled to a same reference picture or scaled to a reference picture closest to a current picture, and

wherein, for a spatial temporal motion vector prediction (STMVP) of the current block, a collocated picture as used in temporal motion vector prediction (TMVP) and/or in advanced temporal motion vector prediction (ATMVP), is used as the target reference picture, or

wherein the target reference picture is determined by a first available merge candidate;

derive a motion vector predictor of the current block by using the scaled plurality of MVs; and

perform a conversion between the current block and a bitstream of the video using the at least one motion candidate list.

16. A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by a video processing apparatus, wherein the method comprises:

determining, for a current block of the video, at least one motion candidate list, wherein the at least one motion candidate list comprises at least one motion candidate each of which is derived from a plurality of spatial and temporal neighboring blocks, wherein the at least one motion candidate is calculated based on a plurality of motion vectors (MV) associated with the plurality of spatial and temporal neighboring blocks and is added to the at least one motion candidate list, and wherein multiple motion candidates are derived and added to the at least one motion candidate list;

scaling the plurality of MVs associated with the plurality of spatial and temporal neighboring blocks to a target reference picture, wherein two or more MVs are scaled to a same reference picture or scaled to a reference picture closest to a current picture, and

wherein, for a spatial temporal motion vector prediction (STMVP) of the current block, a collocated picture as used in temporal motion vector prediction (TMVP) and/or in advanced temporal motion vector prediction (ATMVP), is used as the target reference picture, or

wherein the target reference picture is determined by a first available merge candidate;

deriving a motion vector predictor of the current block by using the scaled plurality of MVs; and

generating the bitstream from the current block using the at least one motion candidate list.

17. The apparatus of claim 15 , wherein the instructions upon execution by the processor, cause the processor to apply a linear function to the plurality of MVs associated with the plurality of spatial and temporal neighboring blocks to derive a motion vector predictor of the current block.

18. The apparatus of claim 17 , wherein an average of the plurality of MVs is defined as the motion vector predictor of the current block.

19. The apparatus of claim 17 , wherein different weighs are applied to different ones of the plurality of MVs to derive the motion vector predictor of the current block.

20. The apparatus of claim 15 , wherein if a reference picture of a neighboring block is different from the collocated picture, the MV of the neighboring block is scaled by using a HEVC temporal MV scaling scheme and is used in the STMVP of the current block.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2021
From: ZHANG, LI; ZHANG, KAI
To: BYTEDANCE INC.
Reel/Frame 055655/0405 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2021
From: LIU, HONGBIN; WANG, YUE
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 055655/0485 →
Priority Claims (1)
WO PCT/CN2019/107166 · Sep 23, 2018 · international
Continuity (2)
Continuation PCTIB2019058021 · Sep 23, 2019
Related Publication 20210211647A1 · Jul 8, 2021