IP Library Granted Patent US 12689726
Granted Patent B2
US 12689726 · App. 18/194,420 · Granted Jul 21, 2026

Multiple merge lists and orders for inter prediction with geometric partitioning

Inventors: Li Zhang (San Diego, CA); Kai Zhang (San Diego, CA); Hongbin Liu (Beijing, CN); Yue Wang (Beijing, CN); Na Zhang (Beijing, CN)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/119H04N19/105H04N19/132H04N19/137H04N19/139H04N19/159H04N19/176H04N19/46H04N19/52
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12689726
App. No.
18/194,420
Granted
Jul 21, 2026
Kind
B2
Abstract

Devices, systems and methods for digital video coding, which include geometric partitioning, are described. An exemplary method for video processing includes making a decision, based on a priority rule, regarding an order of insertion of motion candidates into a motion candidate list for a conversion between a current block of video and a bitstream representation of the video, wherein the current block is coded using a geometry partition mode; and performing, based on the decision and the motion candidate list, the conversion.

Claims (99)

1 . A method of processing video data, comprising:

determining, during a conversion between a current block of a video and a bitstream of the video, that the current block is coded with a geometric partitioning mode;

determining a first motion information MVInfo1 and a second motion information MVInfo2;

performing the conversion based on the MVInfo1 and the MVInfo2, wherein the conversion comprises applying a weighting process to generate a final prediction for the current block based on a weighted sum of prediction samples derived from the MVInfo1 and the MVInfo2;

storing, in response to both the MVInfo1 and the MVInfo2 are from a first reference picture list LX, a first single set of motion information with a m×n subblock within a weighted area of the current block, and wherein the storing is independent of a presence or an absence of a reference picture of the MVInfo1 or a reference picture of the MVInfo2 in a second reference picture list L(1−X), wherein the first single set of motion information comprises an uni-prediction motion information, wherein X=0 or X=1;

wherein the first single set of motion information is allowed to be used for encoding or decoding of a subsequent video block of the current block of the video,

wherein the bitstream includes a field in a sequence parameter set (SPS) that is indicative of a maximum number of allowed motion candidates for the geometric partitioning mode, and wherein the field is explicitly included in the bitstream and set to a difference between M and the maximum number of allowed motion candidates, wherein M is an integer.

2 . The method of claim 1 , further comprising:

storing, in response to the MVInfo1 is from the first reference picture list LX and the MVInfo2 is from the second reference picture list L(1−X), a second single set of motion information with the m×n subblock within the weighted area of the current block, wherein the second single set of motion information comprises bi-prediction motion information;

wherein the second single set of motion information is allowed to be used for encoding or decoding of a subsequent video block of the video.

3 . The method of claim 1 , further comprising:

storing a third single set of motion information with a m×n subblock within a non-weighted area, wherein the third single set of motion information comprises uni-prediction motion information which is based on the MVInfo1 or the MVinfo2;

wherein the third single set of motion information is allowed to be used for encoding or decoding of a subsequent video block of the video.

4 . The method of claim 1 , wherein the first single set of motion information comprising the uni-prediction motion information is based on the MVInfo1 or the MVInfo2; or wherein the first single set of motion information comprising the uni-prediction motion information is equal to the MVInfo1 or the MVInfo2.

5 . The method of claim 1 , wherein the stored first single set of motion information is used in at least one of temporal motion prediction for other blocks in different pictures, spatial motion prediction for other blocks in a current picture comprising the current block or an in-loop process.

6 . The method of claim 2 , wherein the bi-prediction motion information is based on combining the MVInfo1 and the MVInfo2.

7 . The method of claim 6 , wherein the bi-prediction motion information is derived by combining motion vectors and reference picture indices of the MVInfo1 and the MVInfo2.

8 . The method of claim 1 , wherein m=4 or 8 and n=4 or 8.

9 . The method of claim 1 , wherein the geometric partitioning mode includes multiple partition schemes and at least one partition scheme divides the current block in to two partitions with a split boundary, at least one of which is non-square and non-rectangular, wherein in each partition, multiple pixels around the split boundary are part of the weighted area.

10 . The method of claim 9 , wherein the MVInfo1 comprises a set of motion information associated with a partition covering a top-right corner sample and MVInfo2 comprises a set of motion information associated with a partition covering a bottom-left corner sample upon the top-right corner sample and the bottom-left corner sample are in two different partitions; and wherein a direction of the split boundary for the current block is from a top-left corner to a bottom-right corner, or

wherein the MVInfo1 comprises a set of motion information associated with a partition covering a top-left corner sample and MVInfo2 comprises a set of motion information associated with a partition covering a bottom-right corner sample upon the top-left corner sample and the bottom-right corner sample are in two different partitions; and wherein a direction of the split boundary for the current block is from a top-right corner to a bottom-left corner.

11 . The method of claim 1 , wherein neither of weights of prediction samples within the weighted area is equal to 0.

12 . The method of claim 1 ,

wherein the maximum number of allowed motion candidates for the current block is set to M minus the difference indicated by the field, and

wherein M is equal to or less than 6.

13 . The method of claim 12 , wherein M is equal to 5 or 6.

14 . The method of claim 1 , wherein the conversion comprises decoding the current block from the bitstream.

15 . The method of claim 1 , wherein the conversion comprises encoding the current block into the bitstream.

16 . An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

determine, during a conversion between a current block of a video and a bitstream of the video, that the current block is coded with a geometric partitioning mode;

determine a first motion information MVInfo1 and a second motion information MVInfo2;

perform the conversion based on the MVInfo1 and the MVInfo2, wherein the conversion comprises applying a weighting process to generate a final prediction for the current block based on a weighted sum of prediction samples derived from the MVInfo1 and the MVInfo2, store, in response to both the MVInfo1 and the MVInfo2 are from a first reference picture list LX, a first single set of motion information with a m×n subblock within a weighted area of the current block, and wherein the storing is independent of a presence or an absence of a reference picture of the MVInfo1 or a reference picture of the MVInfo2 in a second reference picture list L(1−X), wherein the first single set of motion information comprises an uni-prediction motion information, wherein X=0 or X=1;

wherein the first single set of motion information is allowed to be used for encoding or decoding of a subsequent video block of the current block of the video,

wherein the bitstream includes a field in a sequence parameter set (SPS) that is indicative of a maximum number of allowed motion candidates for the geometric partitioning mode, and wherein the field is explicitly included in the bitstream and set to a difference between M and the maximum number of allowed motion candidates, wherein M is an integer.

17 . The apparatus of claim 16 , further comprising:

storing, in response to the MVInfo1 is from the first reference picture list LX and the MVInfo2 is from the second reference picture list L(1−X), a second single set of motion information with the m×n subblock within the weighted area of the current block, wherein the second single set of motion information comprises bi-prediction motion information, wherein the second single set of motion information is allowed to be used for encoding or decoding of a subsequent video block of the video;

storing a third single set of motion information with a m×n subblock within a non-weighted area, wherein the third single set of motion information comprises uni-prediction motion information which is based on the MVInfo1 or the MVinfo2, wherein the third single set of motion information is allowed to be used for encoding or decoding of a subsequent video block of the video;

wherein the first single set of motion information comprising the uni-prediction motion information is based on the MVInfo1 or the MVInfo2; or

wherein the first single set of motion information comprising the uni-prediction motion information is equal to the MVInfo1 or the MVInfo2;

wherein the stored first single set of motion information is used in at least one of temporal motion prediction for other blocks in different pictures, spatial motion prediction for other blocks in a current picture comprising the current block or an in-loop process;

wherein the bi-prediction motion information is based on combining the MVInfo1 and the MVInfo2;

wherein the bi-prediction motion information is derived by combining motion vectors and reference picture indices of the MVInfo1 and the MVInfo2;

wherein m=4 or 8 and n=4 or 8;

wherein the geometric partitioning mode includes multiple partition schemes and at least one partition scheme divides the current block in to two partitions with a split boundary, at least one of which is non-square and non-rectangular, wherein in each partition, multiple pixels around the split boundary are part of the weighted area;

wherein the MVInfo1 comprises a set of motion information associated with a partition covering a top-right corner sample and MVInfo2 comprises a set of motion information associated with a partition covering a bottom-left corner sample upon the top-right corner sample and the bottom-left corner sample are in two different partitions; and

wherein a direction of the split boundary for the current block is from a top-left corner to a bottom-right corner, or

wherein the MVInfo1 comprises a set of motion information associated with a partition covering a top-left corner sample and MVInfo2 comprises a set of motion information associated with a partition covering a bottom-right corner sample upon the top-left corner sample and the bottom-right corner sample are in two different partitions; and wherein a direction of the split boundary for the current block is from a top-right corner to a bottom-left corner;

wherein neither of weights of prediction samples within the weighted area is equal to 0;

wherein the maximum number of allowed motion candidates for the current block is set to M minus the difference indicated by the field, and

wherein M is equal to or less than 6;

wherein M is equal to 5 or 6.

18 . A non-transitory computer-readable storage medium storing instructions that cause a processor to:

determine, during a conversion between a current block of a video and a bitstream of the video, that the current block is coded with a geometric partitioning mode;

determine a first motion information MVInfo1 and a second motion information MVInfo2;

perform the conversion based on the MVInfo1 and the MVInfo2, wherein the conversion comprises applying a weighting process to generate a final prediction for the current block based on a weighted sum of prediction samples derived from the MVInfo1 and the MVInfo2,

store, in response to both the MVInfo1 and the MVInfo2 are from a first reference picture list LX, a first single set of motion information with a m×n subblock within a weighted area of the current block, and wherein the storing is independent of a presence or an absence of a reference picture of the MVInfo1 or a reference picture of the MVInfo2 in a second reference picture list L(1−X), wherein the first single set of motion information comprises an uni-prediction motion information, wherein X=0 or X=1;

wherein the first single set of motion information is allowed to be used for encoding or decoding of a subsequent video block of the current block of the video,

wherein the bitstream includes a field in a sequence parameter set (SPS) that is indicative of a maximum number of allowed motion candidates for the geometric partitioning mode, and wherein the field is explicitly included in the bitstream and set to a difference between M and the maximum number of allowed motion candidates, wherein M is an integer.

19 . The storage medium of claim 18 , further comprising:

storing, in response to the MVInfo1 is from the first reference picture list LX and the MVInfo2 is from the second reference picture list L(1−X), a second single set of motion information with the m×n subblock within the weighted area of the current block, wherein the second single set of motion information comprises bi-prediction motion information, wherein the second single set of motion information is allowed to be used for encoding or decoding of a subsequent video block of the video;

storing a third single set of motion information with a m×n subblock within a non-weighted area, wherein the third single set of motion information comprises uni-prediction motion information which is based on the MVInfo1 or the MVinfo2, wherein the third single set of motion information is allowed to be used for encoding or decoding of a subsequent video block of the video;

wherein the first single set of motion information comprising the uni-prediction motion information is based on the MVInfo1 or the MVInfo2; or

wherein the first single set of motion information comprising the uni-prediction motion information is equal to the MVInfo1 or the MVInfo2;

wherein the stored first single set of motion information is used in at least one of temporal motion prediction for other blocks in different pictures, spatial motion prediction for other blocks in a current picture comprising the current block or an in-loop process;

wherein the bi-prediction motion information is based on combining the MVInfo1 and the MVInfo2;

wherein the bi-prediction motion information is derived by combining motion vectors and reference picture indices of the MVInfo1 and the MVInfo2;

wherein m=4 or 8 and n=4 or 8;

wherein the geometric partitioning mode includes multiple partition schemes and at least one partition scheme divides the current block in to two partitions with a split boundary, at least one of which is non-square and non-rectangular, wherein in each partition, multiple pixels around the split boundary are part of the weighted area;

wherein the MVInfo1 comprises a set of motion information associated with a partition covering a top-right corner sample and MVInfo2 comprises a set of motion information associated with a partition covering a bottom-left corner sample upon the top-right corner sample and the bottom-left corner sample are in two different partitions; and

wherein a direction of the split boundary for the current block is from a top-left corner to a bottom-right corner, or

wherein the MVInfo1 comprises a set of motion information associated with a partition covering a top-left corner sample and MVInfo2 comprises a set of motion information associated with a partition covering a bottom-right corner sample upon the top-left corner sample and the bottom-right corner sample are in two different partitions; and wherein a direction of the split boundary for the current block is from a top-right corner to a bottom-left corner;

wherein neither of weights of prediction samples within the weighted area is equal to 0;

wherein the maximum number of allowed motion candidates for the current block is set to M minus the difference indicated by the field, and

wherein M is equal to or less than 6;

wherein M is equal to 5 or 6.

20 . A method for storing a bitstream of a video, wherein the method comprises:

determining that a current block of the video is coded with a geometric partitioning mode;

determining a first motion information MVInfo1 and a second motion information MVInfo2;

generating the bitstream of the video from the current block based on the MVInfo1 and the MVInfo2 and storing the bitstream in a non-transitory computer-readable recording medium, wherein the generating comprises applying a weighting process to generate a final prediction for the current block based on a weighted sum of prediction samples derived from the MVInfo1 and the MVInfo2;

storing, in response to both the MVInfo1 and the MVInfo2 are from a first reference picture list LX, a first single set of motion information with a m×n subblock within a weighted area of the current block, and wherein the storing is independent of a presence or an absence of a reference picture of the MVInfo1 or a reference picture of the MVInfo2 in a second reference picture list L(1−X), wherein the first single set of motion information comprises an uni-prediction motion information, wherein X=0 or X=1;

wherein the first single set of motion information is allowed to be used for encoding or decoding of a subsequent video block of the current block of the video,

wherein the bitstream includes a field in a sequence parameter set (SPS) that is indicative of a maximum number of allowed motion candidates for the geometric partitioning mode, and wherein the field is explicitly included in the bitstream and set to a difference between M and the maximum number of allowed motion candidates, wherein M is an integer.

21 . The method of claim 20 , further comprising:

storing, in response to the MVInfo1 is from the first reference picture list LX and the MVInfo2 is from the second reference picture list L(1−X), a second single set of motion information with the m×n subblock within the weighted area of the current block, wherein the second single set of motion information comprises bi-prediction motion information, wherein the second single set of motion information is allowed to be used for encoding or decoding of a subsequent video block of the video;

storing a third single set of motion information with a m×n subblock within a non-weighted area, wherein the third single set of motion information comprises uni-prediction motion information which is based on the MVInfo1 or the MVinfo2, wherein the third single set of motion information is allowed to be used for encoding or decoding of a subsequent video block of the video;

wherein the first single set of motion information comprising the uni-prediction motion information is based on the MVInfo1 or the MVInfo2; or

wherein the first single set of motion information comprising the uni-prediction motion information is equal to the MVInfo1 or the MVInfo2;

wherein the stored first single set of motion information is used in at least one of temporal motion prediction for other blocks in different pictures, spatial motion prediction for other blocks in a current picture comprising the current block or an in-loop process;

wherein the bi-prediction motion information is based on combining the MVInfo1 and the MVInfo2;

wherein the bi-prediction motion information is derived by combining motion vectors and reference picture indices of the MVInfo1 and the MVInfo2;

wherein m=4 or 8 and n=4 or 8;

wherein the geometric partitioning mode includes multiple partition schemes and at least one partition scheme divides the current block in to two partitions with a split boundary, at least one of which is non-square and non-rectangular, wherein in each partition, multiple pixels around the split boundary are part of the weighted area;

wherein the MVInfo1 comprises a set of motion information associated with a partition covering a top-right corner sample and MVInfo2 comprises a set of motion information associated with a partition covering a bottom-left corner sample upon the top-right corner sample and the bottom-left corner sample are in two different partitions; and

wherein a direction of the split boundary for the current block is from a top-left corner to a bottom-right corner, or

wherein the MVInfo1 comprises a set of motion information associated with a partition covering a top-left corner sample and MVInfo2 comprises a set of motion information associated with a partition covering a bottom-right corner sample upon the top-left corner sample and the bottom-right corner sample are in two different partitions; and wherein a direction of the split boundary for the current block is from a top-right corner to a bottom-left corner;

wherein neither of weights of prediction samples within the weighted area is equal to 0;

wherein the maximum number of allowed motion candidates for the current block is set to M minus the difference indicated by the field, and

wherein M is equal to or less than 6;

wherein M is equal to 5 or 6.