IP Library › Granted Patent US 11,758,145
Granted Patent B2
US 11,758,145 · App. 17/173,601 · Granted Sep 12, 2023

Overlapped block motion compensation using temporal neighbors

Inventors: Hongbin Liu (Beijing, CN); Li Zhang (San Diego, CA); Kai Zhang (San Diego, CA); Yue Wang (Beijing, CN)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD; BYTEDANCE INC.
H04N19/139H04N19/105H04N19/107H04N19/137H04N19/159H04N19/176H04N19/196H04N19/31H04N19/51H04N19/521
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,758,145
App. No.
17/173,601
Granted
Sep 12, 2023
Kind
B2
Abstract

Devices, systems and methods for digital video coding, which includes an overlapped block motion compensation (OBMC) process based on temporal neighbors, are described. An exemplary method for video processing includes generating, based on a weighted sum of at least two temporary prediction blocks, a prediction block for a current video block, a first of the at least two temporary prediction blocks being based on a first motion information associated with the current video block, a second of the at least two temporary prediction blocks being based on a second motion information associated with at least one neighboring block of the current video block, and the at least one neighboring block including a temporally neighboring block, and performing, based on the prediction block, a conversion between the current video block and a bitstream representation of the current video block.

Claims (59)

1. A method for video processing, comprising:

generating, based on a weighted sum of at least two temporal prediction blocks, a prediction block for a current video block, wherein a first of the at least two temporal prediction blocks is based on a first motion information associated with the current video block, wherein a second of the at least two temporal prediction blocks is based on a second motion information associated with at least one neighboring block of the current video block, and wherein the at least one neighboring block comprises a temporally neighboring block; and

performing, based on the prediction block, a conversion between the current video block and a bitstream of the current video block;

wherein a weighting factor of the second of the at least two temporal prediction blocks is based on a location or coding mode of the at least one neighboring block;

the weighting factor is a first weighting factor upon a determination that the at least one neighboring block comprises a spatially neighboring block of the current video block, wherein the weighting factor is a second weighting factor upon a determination that the at least one neighboring block comprises a temporally neighboring block of the current video block; wherein the weighting factor is a third weighting factor upon a determination that the current video block is coded using an intra prediction mode;

wherein the first, second or third weighting factors are further based on a dimension of the current video block;

wherein performing the conversion is based on a coding mode of the current video block, a size or a shape of the current video block, or a size of a sub-block of the current video block;

wherein the coding mode of the current video block comprises a conventional translation motion with an affine mode being disabled;

wherein a product of a height of the current video block and a width of the current video block is greater than or equal to a threshold;

wherein the height of the current video block is greater than or equal to a first threshold, and wherein the width of the current video block is greater than or equal to a second threshold; and

wherein performing the conversion is further based on a slice type of a slice comprising the current video block, a low-delay check flag or a temporal layer, or based on signaling in a sequence parameter set (SPS), a picture parameter set (PPS), a video parameter set (VPS), a slice header, a coding tree unit (CTU), a coding unit (CU), a group of CTUs or a group of CUs.

2. The method of claim 1 , wherein the current video block is coded with a sub-block based coding tool, and wherein a final prediction of a current sub-block of the current video block is based on at least a motion information of temporally neighboring blocks of the current sub-block;

wherein a final prediction for each of a subset of sub-blocks of the current video block is based on the second motion information, and the subset excludes at least one sub-block of the current video block; and

wherein the first motion information and the second motion information are not derived from a same prediction process.

3. The method of claim 1 , wherein the current video block is coded without a sub-block based coding tool.

4. The method of claim 1 , wherein performing the conversion is further based, upon a determination of an availability of motion information associated with at least one spatially neighboring block of the current video block, on a third motion information associated with the at least one spatially neighboring block.

5. The method of claim 1 , wherein the temporally neighboring block of the at least one neighboring block is located in a collocated picture that is signaled in the sequence parameter set (SPS), the picture parameter set (PPS), the video parameter set (VPS) or the slice header; or alternatively, the temporally neighboring block of the at least one neighboring block is located in one of a plurality of reference pictures that are signaled in the sequence parameter set (SPS), the picture parameter set (PPS), the video parameter set (VPS), the slice header or a tile header.

6. The method of claim 1 , wherein the temporally neighboring block of the at least one neighboring block is located in a predetermined reference picture, and wherein the predetermined reference picture is in list 0 or list 1.

7. The method of claim 1 , wherein the temporally neighboring block of the at least one neighboring block is a collocated block in a selected reference picture; wherein a current prediction unit (PU) or coding unit (CU) comprises the current video block, and wherein a motion vector of the current PU or CU comprises an identification of the temporally neighboring block of the at least one neighboring block; wherein the motion vector is a scaled motion vector; and wherein a motion vector of the temporally neighboring block of the at least one neighboring block is scaled to one of a plurality of reference pictures that are signaled in the sequence parameter set (SPS), the picture parameter set (PPS), the video parameter set (VPS) or the slice header.

8. The method of claim 1 , wherein a current prediction unit (PU) or coding unit (CU) comprises the current video block, wherein a motion vector of the current PU or CU is scaled to a first reference picture of the current PU or CU, and wherein a motion vector of the temporally neighboring block of the at least one neighboring block is scaled to the first reference picture.

9. The method of claim 1 , wherein a motion vector of the at least one neighboring block is scaled to a predetermined reference picture, and wherein the predetermined reference picture is a first reference picture in list 0 or list 1.

10. The method of claim 1 , wherein the second and third weighting factors are signaled in the sequence parameter set (SPS), the picture parameter set (PPS), the video parameter set (VPS) or the slice header, wherein the first weighting factor is greater than the second weighting factor, and wherein the second weighting factor is greater than the third weighting factor, and wherein the second weighting factor is equal to the third weighting factor.

11. The method of claim 1 , wherein the performing the conversion comprises applying a motion compensation process on a luma component of the current video block, or alternatively, applying a motion compensation process on one or more of a plurality of chroma components of the current video block.

12. The method of claim 1 , wherein one or more weights of the weighted sum are based on a coordinate of a sample within the current video block, or, the one or more weights of the weighted sum are based on a distance of the sample within the current video block to a boundary of the current video block, and wherein generating the prediction block is part of an overlapped block motion compensation (OBMC) process.

13. The method of claim 1 , wherein performing the conversion comprises encoding the current video block into the bitstream.

14. The method of claim 1 , wherein performing the conversion comprises decoding the current video block from the bitstream.

15. An apparatus for video processing comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

generate, based on a weighted sum of at least two temporal prediction blocks, a prediction block for a current video block, wherein a first of the at least two temporal prediction blocks is based on a first motion information associated with the current video block, wherein a second of the at least two temporal prediction blocks is based on a second motion information associated with at least one neighboring block of the current video block, and wherein the at least one neighboring block comprises a temporally neighboring block; and

perform, based on the prediction block, a conversion between the current video block and a bitstream of the current video block;

wherein a weighting factor of the second of the at least two temporal prediction blocks is based on a location or coding mode of the at least one neighboring block;

the weighting factor is a first weighting factor upon a determination that the at least one neighboring block comprises a spatially neighboring block of the current video block, wherein the weighting factor is a second weighting factor upon a determination that the at least one neighboring block comprises a temporally neighboring block of the current video block; wherein the weighting factor is a third weighting factor upon a determination that the current video block is coded using an intra prediction mode;

wherein the first, second or third weighting factors are further based on a dimension of the current video block;

wherein performing the conversion is based on a coding mode of the current video block, a size or a shape of the current video block, or a size of a sub-block of the current video block;

wherein the coding mode of the current video block comprises a conventional translation motion with an affine mode being disabled;

wherein a product of a height of the current video block and a width of the current video block is greater than or equal to a threshold;

wherein the height of the current video block is greater than or equal to a first threshold, and wherein the width of the current video block is greater than or equal to a second threshold; and

wherein performing the conversion is further based on a slice type of a slice comprising the current video block, a low-delay check flag or a temporal layer, or based on signaling in a sequence parameter set (SPS), a picture parameter set (PPS), a video parameter set (VPS), a slice header, a coding tree unit (CTU), a coding unit (CU), a group of CTUs or a group of CUs.

16. A non-transitory computer-readable storage medium storing instructions that cause a processor to:

generate, based on a weighted sum of at least two temporal prediction blocks, a prediction block for a current video block, wherein a first of the at least two temporal prediction blocks is based on a first motion information associated with the current video block, wherein a second of the at least two temporal prediction blocks is based on a second motion information associated with at least one neighboring block of the current video block, and wherein the at least one neighboring block comprises a temporally neighboring block; and

perform, based on the prediction block, a conversion between the current video block and a bitstream of the current video block;

wherein a weighting factor of the second of the at least two temporal prediction blocks is based on a location or coding mode of the at least one neighboring block;

the weighting factor is a first weighting factor upon a determination that the at least one neighboring block comprises a spatially neighboring block of the current video block, wherein the weighting factor is a second weighting factor upon a determination that the at least one neighboring block comprises a temporally neighboring block of the current video block; wherein the weighting factor is a third weighting factor upon a determination that the current video block is coded using an intra prediction mode;

wherein the first, second or third weighting factors are further based on a dimension of the current video block;

wherein performing the conversion is based on a coding mode of the current video block, a size or a shape of the current video block, or a size of a sub-block of the current video block;

wherein the coding mode of the current video block comprises a conventional translation motion with an affine mode being disabled;

wherein a product of a height of the current video block and a width of the current video block is greater than or equal to a threshold;

wherein the height of the current video block is greater than or equal to a first threshold, and wherein the width of the current video block is greater than or equal to a second threshold; and

wherein performing the conversion is further based on a slice type of a slice comprising the current video block, a low-delay check flag or a temporal layer, or based on signaling in a sequence parameter set (SPS), a picture parameter set (PPS), a video parameter set (VPS), a slice header, a coding tree unit (CTU), a coding unit (CU), a group of CTUs or a group of CUs.

17. A non-transitory computer-readable recording medium storing a bitstream which is generated by a method performed by a video processing apparatus, wherein the method comprises:

generating, for a conversion between a current video block and the bitstream of the current video block, based on a weighted sum of at least two temporal prediction blocks, a prediction block for the current video block, wherein a first of the at least two temporal prediction blocks is based on a first motion information associated with the current video block, wherein a second of the at least two temporal prediction blocks is based on a second motion information associated with at least one neighboring block of the current video block, and wherein the at least one neighboring block comprises a temporally neighboring block; and

generating, based on the prediction block, the bitstream from the current video block,

wherein a weighting factor of the second of the at least two temporal prediction blocks is based on a location or coding mode of the at least one neighboring block;

the weighting factor is a first weighting factor upon a determination that the at least one neighboring block comprises a spatially neighboring block of the current video block, wherein the weighting factor is a second weighting factor upon a determination that the at least one neighboring block comprises a temporally neighboring block of the current video block; wherein the weighting factor is a third weighting factor upon a determination that the current video block is coded using an intra prediction mode;

wherein the first, second or third weighting factors are further based on a dimension of the current video block;

wherein performing the conversion is based on a coding mode of the current video block, a size or a shape of the current video block, or a size of a sub-block of the current video block;

wherein the coding mode of the current video block comprises a conventional translation motion with an affine mode being disabled;

wherein a product of a height of the current video block and a width of the current video block is greater than or equal to a threshold;

wherein the height of the current video block is greater than or equal to a first threshold, and wherein the width of the current video block is greater than or equal to a second threshold; and

wherein performing the conversion is further based on a slice type of a slice comprising the current video block, a low-delay check flag or a temporal layer, or based on signaling in a sequence parameter set (SPS), a picture parameter set (PPS), a video parameter set (VPS), a slice header, a coding tree unit (CTU), a coding unit (CU), a group of CTUs or a group of CUs.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2021
From: ZHANG, LI; ZHANG, KAI
To: BYTEDANCE INC.
Reel/Frame 055238/0075 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2021
From: LIU, HONGBIN; WANG, YUE
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 055238/0127 →
Priority Claims (1)
WO PCT/CN2018/102163 · Aug 24, 2018 · international
Continuity (2)
Continuation PCTIB2019057140 · Aug 26, 2019
Related Publication 20210195205A1 · Jun 24, 2021
Cited By (1)
US 12,689,763