IP Library Granted Patent US 11,632,566
Granted Patent B2
US 11,632,566 · App. 17/317,452 · Granted Apr 18, 2023

Inter prediction with refinement in video processing

Inventors: Hongbin Liu (Beijing, CN); Li Zhang (San Diego, CA); Kai Zhang (San Diego, CA); Jizheng Xu (San Diego, CA); Yue Wang (Beijing, CN)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/517H04N19/105H04N19/132H04N19/137H04N19/139H04N19/159H04N19/176H04N19/196H04N19/521H04N19/573H04N19/577
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,632,566
App. No.
17/317,452
Filed
May 11, 2021
Granted
Apr 18, 2023
Kind
B2
Art Unit
2482
USPC
375/240.02
Abstract

A method for processing a video includes performing a conversion between a current block of visual media data and a corresponding coded representation of the visual media data, wherein the conversion of the current block includes determining whether a use of one or both of a bi-directional optical flow (BIO) technique or a decoder-side motion vector refinement (DMVR) technique to the current block is enabled or disabled, and wherein the determining the use of the BIO technique or the DMVR technique is based on a cost criterion associated with the current block.

Claims (38)

1. A method of processing video data, comprising:

determining a first temporal gradient using reference pictures associated with a first video block or a sub-block thereof;

determining a second temporal gradient using reference pictures associated with a second video block or a sub-block thereof;

performing a modification of the first temporal gradient and a modification of the second temporal gradient conditionally based on a comparison of a mean difference between the reference pictures associated with the first video block and/or the second video block and at least one threshold value to generate a modified first temporal gradient and a modified second temporal gradient, wherein the modification of the first temporal gradient associated with the first video block is different from the modification of the second temporal gradient associated with the second video block; and

performing a conversion between the first video block and the second video block and their corresponding bitstream based on the modified first temporal gradient and the modified second temporal gradient.

2. The method of claim 1 , wherein the modification of the first temporal gradient and/or the modification of the second temporal gradient is conditionally based on at least one of an absolute mean difference between the reference pictures associated with the first video block and/or the second video block being greater than a first threshold value, and the absolute mean difference between the reference pictures associated with the first video block and/or the second video block being less than a second threshold value.

3. The method of claim 2 , wherein the first threshold value is 4, and the second threshold value is 20.

4. The method of claim 1 , further comprising:

disabling a use of a bi-directional optical flow (BIO) technique on the first video block and/or the second block based on an absolute mean difference between the reference pictures associated with the first video block and/or the second video block being greater than a threshold value.

5. The method of claim 2 , wherein the first threshold value or the second threshold is indicated in VPS, SPS, PPS, a picture, a slice, or a tile level associated with the first video block and/or the second video block, or the first threshold value or the second threshold are implicitly predefined parameters.

6. The method of claim 2 , wherein the first threshold value or the second threshold is different for different coding units (CUs), largest coding units (LCUs), slices, tiles, or pictures associated with the first video block and/or the second video block.

7. The method of claim 2 , wherein the first threshold value or the second threshold is based on a decoded or an encoded pixel value associated with the first video block and/or the second video block.

8. The method of claim 2 , wherein the first threshold value or the second threshold for a first set of reference pictures is different from the first threshold value or the second threshold for a second set of reference pictures.

9. The method of claim 1 , wherein the modification of the first temporal gradient and/or the modification of the second temporal gradient is conditionally based on at least one of an absolute mean of the reference pictures associated with the first video block and/or the second video block being greater than a third threshold value, and the absolute mean of the reference pictures associated with the first video block and/or the second video block being smaller than a fourth threshold value.

10. The method of claim 9 , wherein the third threshold value is 40, and the fourth threshold value is 100.

11. The method of claim 1 , wherein the modification of the first temporal gradient and/or the modification of the second temporal gradient is conditionally based on an absolute mean of the reference pictures associated with the first video block and/or the second video block being greater than an absolute mean difference of the reference pictures associated with the first video block and/or the second video block times a first multiplication factor, or the absolute mean of the reference pictures associated with the first video block and/or the second video block being less than an absolute mean difference of the reference pictures associated with the first video block and/or the second video block times a second multiplication factor.

12. The method of claim 11 , wherein the first multiplication factor or the second multiplication factor is 4.5.

13. The method of claim 1 , further comprising:

modifying a first reference block to generate a first modified reference block, and a second reference block to generate a second modified reference block, wherein both the first reference block and the second reference block are associated with a current block of video data;

determining differences between the first modified reference block and the second modified reference block, the differences including one or more of: a sum of absolute transformed differences (SATD), a mean removed sum of absolute transformed differences (MRSATD), a sum of squares error (SSE), a mean removed sum of squares error (MRSSE), a mean value difference, or gradient values; and

performing a conversion between the current block of video data and a corresponding bitstream of the video data, wherein the conversion includes a use of the differences between the first modified reference block and the second modified reference block generated from respectively modifying the first reference block and the second reference block.

14. The method of claim 13 , wherein the modifying the first reference block and the second reference block includes:

computing a first arithmetic mean based on sample values included in the first reference block and a second arithmetic mean based on sample values included in the second reference block;

subtracting the first arithmetic mean from samples included in the first reference block and the second arithmetic mean from samples included in the second reference block.

15. The method of claim 14 , wherein the first arithmetic mean and the second arithmetic mean are based on a subset of samples respectively included in the first reference block and the second reference block.

16. The method of claim 13 , wherein the first reference block and/or the second reference block are sub-blocks associated with the current block.

17. The method of claim 1 , wherein the conversion includes encoding the video block into the bitstream.

18. The method of claim 1 , wherein the conversion includes decoding the video block from the bitstream.

19. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

determine a first temporal gradient using reference pictures associated with a first video block or a sub-block thereof;

determine a second temporal gradient using reference pictures associated with a second video block or a sub-block thereof;

perform a modification of the first temporal gradient and a modification of the second temporal gradient conditionally based on a comparison of a mean difference between the reference pictures associated with the first video block and/or the second video block and at least one threshold value to generate a modified first temporal gradient and a modified second temporal gradient, wherein the modification of the first temporal gradient associated with the first video block is different from the modification of the second temporal gradient associated with the second video block; and

perform a conversion between the first video block and the second video block and their corresponding bitstream based on the modified first temporal gradient and the modified second temporal gradient.

20. A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by a video processing apparatus, wherein the method comprises:

determining a first temporal gradient using reference pictures associated with a first video block or a sub-block thereof;

determining a second temporal gradient using reference pictures associated with a second video block or a sub-block thereof;

performing a modification of the first temporal gradient and a modification of the second temporal gradient conditionally based on a comparison of a mean difference between the reference pictures associated with the first video block and/or the second video block and at least one threshold value to generate a modified first temporal gradient and a modified second temporal gradient, wherein the modification of the first temporal gradient associated with the first video block is different from the modification of the second temporal gradient associated with the second video block; and

generating corresponding bitstreams from the first video block and the second video block based on the modified first temporal gradient and the modified second temporal gradient.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2021
From: LIU, HONGBIN; WANG, YUE
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 056204/0521 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2021
From: ZHANG, LI; ZHANG, KAI; XU, JIZHENG
To: BYTEDANCE INC.
Reel/Frame 056204/0556 →
Priority Claims (3)
WO PCT/CN2018/116371 · Nov 20, 2018 · international
WO PCT/CN2019/081155 · Apr 2, 2019 · international
WO PCT/CN2019/085796 · May 7, 2019 · international
Continuity (2)
Continuation PCTCN2019119742 · Nov 20, 2019
Related Publication 20210266530A1 · Aug 26, 2021
Cited By (4)
US 12,348,760 US 12,363,337 US 12,432,355 US 12,477,106