IP Library › Granted Patent US 12,015,772
Granted Patent B2
US 12,015,772 · App. 17/950,443 · Granted Jun 18, 2024

Prediction refinement for affine merge and affine motion vector prediction mode

Inventors: Zhipin Deng (Beijing, CN); Li Zhang (San Diego, CA); Ye-Kui Wang (San Diego, CA); Kai Zhang (San Diego, CA); Jizheng Xu (San Diego, CA)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/117H04N19/105H04N19/132H04N19/139H04N19/159H04N19/174H04N19/176H04N19/186H04N19/463H04N19/52H04N19/70H04N19/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,015,772
App. No.
17/950,443
Granted
Jun 18, 2024
Kind
B2
Abstract

A method includes determining, for a conversion between a video block of a video and a bitstream of the video, a size of prediction block corresponding to the video block according to a rule. The method also includes performing the conversion based on the determining. The rule specifies that a first size of the prediction block is determined responsive to whether a prediction refinement using optical flow technique is used for coding the video block. The video block has a second size and is coded using an affine merge mode or an affine advanced motion vector prediction mode.

Claims (80)

1. A method of processing video data, comprising:

determining, for a conversion between a current video block of a video and a bitstream of the video, a size of a prediction block corresponding to the current video block according to a rule; and

performing the conversion based on the determining,

wherein an affine merge mode is enabled for the current video block, and

wherein the rule specifies that a first size of the prediction block is determined responsive to whether a prediction refinement using optical flow technique is enabled for the current video block, and wherein the current video block has a second size;

wherein a second width and a second height of the second size of the current video block are indicated by M and N, respectively, and M and N are integers greater than or equal to 0;

wherein a prediction sample of the prediction block is present as predSamplesLX[xL][yL],

wherein xL is between 0 and M+1 inclusively, and yL is between 0 and N+1 inclusively,

wherein the prediction sample predSamplesLX[xL][yL] is derived by invoking a luma integer sample fetching process for one or more of conditions are true: xL is equal to 0, xL is equal to M+1, yL is equal to 0 and yL is equal to N+1, and the prediction sample predSamplesLX[xL][yL] is derived by invoking a luma sample 8-tap interpolation filtering process for all conditions are false.

2. The method of claim 1 , wherein a first width and a first height of the first size of the prediction block are indicated by (M+M0) and (N+N0), respectively, and

wherein M0 and N0 are integers greater than or equal to 0.

3. The method of claim 2 , wherein in a case that the prediction refinement using optical flow technique is enabled for the current video block, at least one of M0 and N0 is not equal to 0.

4. The method of claim 3 , wherein M0 and N0 are both equal to 2.

5. The method of claim 2 , wherein a prediction refinement utility flag controls values of M0 and N0, and wherein the prediction refinement utility flag indicates whether the prediction refinement using optical flow technique is utilized.

6. The method of claim 5 , wherein the values of M0 and N0 are determined independently from an affine flag.

7. The method of claim 6 , wherein the affine flag is inter_affine_flag which is used to indicate whether to apply an affine motion vector prediction mode.

8. The method of claim 7 , wherein for a video block applied with the affine motion vector prediction mode, a size of a prediction block of the video block is equal to the first size.

9. The method of claim 1 , wherein the affine merge mode includes generating control point motion vectors by using a merge index to select an affine merge candidate from a sub-block merge candidate list which is constructed based on motion information of spatial neighboring coding units.

10. The method of claim 1 , wherein performing the conversion comprises encoding the video into the bitstream.

11. The method of claim 1 , wherein performing the conversion comprises decoding the video from the bitstream.

12. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

determine, for a conversion between a current video block of a video and a bitstream of the video, a size of a prediction block corresponding to the current video block according to a rule; and

perform the conversion based on the determining,

wherein an affine merge mode is enabled for the current video block,

wherein the rule specifies that a first size of the prediction block is determined responsive to whether a prediction refinement using optical flow technique is enabled for the current video block, and

wherein the current video block has a second size;

wherein a second width and a second height of the second size of the current video block are indicated by M and N, respectively, and M and N are integers greater than or equal to 0,

wherein a prediction sample of the prediction block is present as predSamplesLX[xL][yL],

wherein xL is between 0 and M+1 inclusively, and yL is between 0 and N+1 inclusively, and

wherein the prediction sample predSamplesLX[xL][yL] is derived by invoking a luma integer sample fetching process for one or more of conditions are true: xL is equal to 0, xL is equal to M+1, yL is equal to 0 and yL is equal to N+1, and the prediction sample predSamplesLX[xL][yL] is derived by invoking a luma sample 8-tap interpolation filtering process for all conditions are false.

13. The apparatus of claim 12 , wherein a first width and a first height of the first size of the prediction block are indicated by (M+M0) and (N+N0), respectively,

wherein M0 and N0 are integers greater than or equal to 0,

wherein in a case that the prediction refinement using optical flow technique is enabled for the current video block, at least one of M0 and N0 is not equal to 0, and

wherein M0 and N0 are both equal to 2.

14. The apparatus of claim 13 , wherein a prediction refinement utility flag controls values of M0 and N0,

wherein the prediction refinement utility flag indicates whether the prediction refinement using optical flow technique is utilized,

wherein the values of M0 and N0 are determined independently from an affine flag,

wherein the affine flag is inter_affine_flag which is used to indicate whether to apply an affine motion vector prediction mode,

wherein for a video block applied with the affine motion vector prediction mode, a size of a prediction block of the video block is equal to the first size, and

wherein the affine merge mode includes generating control point motion vectors by using a merge index to select an affine merge candidate from a sub-block merge candidate list which is constructed based on motion information of spatial neighboring coding units.

15. A non-transitory computer-readable storage medium storing instructions that cause a processor to:

determine, for a conversion between a current video block of a video and a bitstream of the video, a size of a prediction block corresponding to the current video block according to a rule; and

perform the conversion based on the determining,

wherein an affine merge mode is enabled for the current video block, and

wherein the rule specifies that a first size of the prediction block is determined responsive to whether a prediction refinement using optical flow technique is enabled for the current video block, and

wherein the current video block has a second size;

wherein a second width and a second height of the second size of the current video block are indicated by M and N, respectively, and M and N are integers greater than or equal to 0,

wherein a prediction sample of the prediction block is present as predSamplesLX[xL][yL],

wherein xL is between 0 and M+1 inclusively, and yL is between 0 and N+1 inclusively,

wherein the prediction sample predSamplesLX[xL][yL] is derived by invoking a luma integer sample fetching process for one or more of conditions are true: xL is equal to 0, xL is equal to M+1, yL is equal to 0 and yL is equal to N+1, and the prediction sample predSamplesLX[xL][yL] is derived by invoking a luma sample 8-tap interpolation filtering process for all conditions are false.

16. The non-transitory computer-readable storage medium of claim 15 , wherein a first width and a first height of the first size of the prediction block are indicated by (M+M0) and (N+N0), respectively,

wherein M0 and N0 are integers greater than or equal to 0,

wherein in a case that the prediction refinement using optical flow technique is enabled for the current video block, at least one of M0 and N0 is not equal to 0, and

wherein M0 and N0 are both equal to 2.

17. The non-transitory computer-readable storage medium of claim 16 , wherein a prediction refinement utility flag controls values of M0 and N0,

wherein the prediction refinement utility flag indicates whether the prediction refinement using optical flow technique is utilized,

wherein the values of M0 and N0 are determined independently from an affine flag,

wherein the affine flag is inter_affine_flag which is used to indicate whether to apply an affine motion vector prediction mode,

wherein for a video block applied with the affine motion vector prediction mode, a size of a prediction block of the video block is equal to the first size, and

wherein the affine merge mode includes generating control point motion vectors by using a merge index to select an affine merge candidate from a sub-block merge candidate list which is constructed based on motion information of spatial neighboring coding units.

18. A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by a video processing apparatus, wherein the method comprises:

determining, for a current video block of the video, a size of a prediction block corresponding to the current video block according to a rule; and

generating the bitstream based on the determining,

wherein an affine merge mode is enabled for the current video block, and

wherein the rule specifies that a first size of the prediction block is determined responsive to whether a prediction refinement using optical flow technique is enabled for the current video block, and

wherein the current video block has a second size;

wherein a second width and a second height of the second size of the current video block are indicated by M and N, respectively, and M and N are integers greater than or equal to 0,

wherein a prediction sample of the prediction block is present as predSamplesLX[xL][yL],

wherein xL is between 0 and M+1 inclusively, and yL is between 0 and N+1 inclusively, and

wherein the prediction sample predSamplesLX[xL][yL] is derived by invoking a luma integer sample fetching process for one or more of conditions are true: xL is equal to 0, xL is equal to M+1, yL is equal to 0 and yL is equal to N+1, and the prediction sample predSamplesLX[xL][yL] is derived by invoking a luma sample 8-tap interpolation filtering process for all conditions are false.

19. The non-transitory computer-readable recording medium of claim 18 , wherein a first width and a first height of the first size of the prediction block are indicated by (M+M0) and (N+N0), respectively,

wherein M0 and N0 are integers greater than or equal to 0,

wherein in a case that the prediction refinement using optical flow technique is enabled for the current video block, at least one of M0 and N0 is not equal to 0,

wherein M0 and N0 are both equal to 2,

wherein a prediction refinement utility flag controls values of M0 and N0,

wherein the prediction refinement utility flag indicates whether the prediction refinement using optical flow technique is utilized,

wherein the values of M0 and N0 are determined independently from an affine flag,

wherein the affine flag is inter_affine_flag which is used to indicate whether to apply an affine motion vector prediction mode,

wherein for a video block applied with the affine motion vector prediction mode, a size of a prediction block of the video block is equal to the first size, and

wherein the affine merge mode includes generating control point motion vectors by using a merge index to select an affine merge candidate from a sub-block merge candidate list which is constructed based on motion information of spatial neighboring coding units.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2023
From: ZHANG, LI; WANG, YE-KUI; ZHANG, KAI; XU, JIZHENG
To: BYTEDANCE INC.
Reel/Frame 062658/0893 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2023
From: DENG, ZHIPIN
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 062658/0965 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2023
From: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 062659/0052 →
Priority Claims (1)
WO PCT/CN2020/080602 · Mar 23, 2020 · international
Continuity (2)
Continuation PCTCN2021082243 · Mar 23, 2021
Related Publication 20230042746A1 · Feb 9, 2023