IP Library › Granted Patent US 11,272,206
Granted Patent B2
US 11,272,206 · App. 17/405,179 · Granted Mar 8, 2022

Decoder side motion vector derivation

Inventors: Hongbin Liu (Beijing, CN); Li Zhang (San Diego, CA); Kai Zhang (San Diego, CA); Jizheng Xu (San Diego, CA); Yue Wang (Beijing, CN)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/513H04N19/105H04N19/132H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,272,206
App. No.
17/405,179
Granted
Mar 8, 2022
Kind
B2
Abstract

A method for processing a video includes performing a conversion between a current block of visual media data and a corresponding coded representation of the visual media data, wherein the conversion of the current block includes determining whether a use of one or both of a bi-directional optical flow (BIO) technique or a decoder-side motion vector refinement (DMVR) technique to the current block is enabled or disabled, and wherein the determining the use of the BIO technique or the DMVR technique is based on a cost criterion associated with the current block.

Claims (187)

1. A method of processing video data, comprising:

determining, for a current block of a video, an initial prediction sample;

refining the initial prediction sample, based on an optical flow refinement technology, with a prediction sample offset to acquire a final prediction sample; and

performing a conversion between the current block and a bitstream of the video based on the final prediction sample,

wherein the prediction sample offset is determined based on at least one spatial gradient of the initial prediction sample, wherein the spatial gradient is calculated based on at least a difference between two first prediction samples from a same reference picture list, and

wherein before calculating the difference between the two first prediction samples, values of the two first prediction sample are right shifted with a first value.

2. The method of claim 1 , wherein for a sample location (x,y) in the current block, the two first prediction samples have locations (hx+1, vy) and (hx−1, vy) corresponding to the same reference picture list X or locations (hx, vy+1) and (hx, vy−1) corresponding to the same reference picture list X, and

wherein X=0 or 1, hx=Clip3(1, nCbW, x) and vy=Clip3(1, nCbH, y), nCbW and nCbH is a width and a height of the current block, and wherein Clip3 is a clipping function which is defined as:

Clip

⁢

⁢

3

⁢

(

u

,

v

,

w

)

=

{

u

;

w

<

u

v

;

w

>

v

w

;

otherwise

.

3. The method of claim 1 , wherein the prediction sample offset is determined further based on at least one temporal gradient, wherein the temporal gradient is calculated based on at least a difference between two second prediction samples from different reference picture lists, and wherein before calculating the difference between the two second prediction samples, values of the two second prediction sample are right shifted with a second value.

4. The method of claim 3 , wherein for a sample location (x,y) in the current block, the two second prediction samples have locations (hx, vy) corresponding to a reference picture list 0 and a reference picture list 1, and

wherein hx=Clip3(1, nCbW, x) and vy=Clip3(1, nCbH, y), nCbW and nCbH is a width and a height of the current block, and wherein Clip3 is a clipping function which is defined as:

Clip

⁢

⁢

3

⁢

(

u

,

v

,

w

)

=

{

u

;

w

<

u

v

;

w

>

v

w

;

otherwise

.

5. The method of claim 3 , wherein the first value is different from the second value.

6. The method of claim 1 , wherein the prediction sample offset is determined further based on at least one temporal gradient, wherein the temporal gradient is calculated based on at least a difference between two second prediction samples from different reference picture list, and wherein

a shifting rule of the difference between the two second prediction samples is same as that of the difference between the two first prediction samples, and the shifting rule includes an order of a right-shifting operation and a subtraction operation.

7. The method of claim 1 , wherein whether the optical flow refinement procedure is enabled based on a condition related with a size of the current block.

8. The method of claim 7 , wherein whether a decoder-side motion vector refinement technique is enabled for the current block is based on the same condition, wherein the decoder-side motion vector refinement technique is used to derive a refined motion information of the current block based on a cost between at least one prediction sample acquired based on at least one reference sample of reference picture list 0 and at least one prediction sample acquired based on at least one reference sample of reference picture list 1.

9. The method of claim 8 , wherein the optical flow refinement technology and the decoder-side motion vector refinement technique are enabled at least based on a height of the current block is equal to or greater than T 1 .

10. The method of claim 9 , wherein T 1 =8.

11. The method of claim 1 , wherein performing the conversion includes decoding the current block from the bitstream.

12. The method of claim 1 , wherein performing the conversion includes encoding the current block into the bitstream.

13. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

determine, for a current block of a video, an initial prediction sample;

refine the initial prediction sample, based on an optical flow refinement technology, with a prediction sample offset to acquire a final prediction sample; and

perform a conversion between the current block and a bitstream of the video based on the final prediction sample,

wherein the prediction sample offset is determined based on at least one spatial gradient of the initial prediction sample, wherein the spatial gradient is calculated based on at least a difference between two first prediction samples from a same reference picture list, and

wherein before calculating the difference between the two first prediction samples, values of the two first prediction sample are right shifted with a first value.

14. The apparatus of claim 13 , wherein for a sample location (x,y) in the current block, the two first prediction samples have locations (hx+1, vy) and (hx−1, vy) corresponding to the same reference picture list X or locations (hx, vy+1) and (hx, vy−1) corresponding to the same reference picture list X, and

wherein X=0 or 1, hx=Clip3(1, nCbW, x) and vy=Clip3(1, nCbH, y), nCbW and nCbH is a width and a height of the current block, and wherein Clip3 is a clipping function which is defined as:

Clip

⁢

⁢

3

⁢

(

u

,

v

,

w

)

=

{

u

;

w

<

u

v

;

w

>

v

w

;

otherwise

.

15. The apparatus of claim 13 , wherein the prediction sample offset is determined further based on at least one temporal gradient, wherein the temporal gradient is calculated based on at least a difference between two second prediction samples from different reference picture lists, and

wherein before calculating the difference between the two second prediction samples, values of the two second prediction sample are right shifted with a second value.

16. The apparatus of claim 15 , wherein for a sample location (x,y) in the current block, the two second prediction samples have locations (hx, vy) corresponding to a reference list picture 0 and a reference picture list 1, and

wherein hx=Clip3(1, nCbW, x) and vy=Clip3(1, nCbH, y), nCbW and nCbH is a width and a height of the current block, and wherein Clip3 is a clipping function which is defined as:

Clip

⁢

⁢

3

⁢

(

u

,

v

,

w

)

=

{

u

;

w

<

u

v

;

w

>

v

w

;

otherwise

.

17. The apparatus of claim 15 , wherein the first value is different from the second value.

18. A non-transitory computer-readable storage medium storing instructions that cause a processor to:

determine, for a current block of a video, an initial prediction sample;

refine the initial prediction sample, based on an optical flow refinement technology, with a prediction sample offset to acquire a final prediction sample; and

perform a conversion between the current block and a bitstream of the video based on the final prediction sample,

wherein the prediction sample offset is determined based on at least one spatial gradient of the initial prediction sample, wherein the spatial gradient is calculated based on at least a difference between two first prediction samples from a same reference picture list, and

wherein before calculating the difference between the two first prediction samples, values of the two first prediction sample are right shifted with a first value.

19. The non-transitory computer-readable storage medium of claim 18 , wherein for a sample location (x,y) in the current block, the two first prediction samples have locations (hx+1, vy) and (hx−1, vy) corresponding to the same reference picture list X or locations (hx, vy+1) and (hx, vy−1) corresponding to the same reference picture list X, and

wherein X=0 or 1, hx=Clip3(1, nCbW, x) and vy=Clip3(1, nCbH, y), nCbW and nCbH is a width and a height of the current block, and wherein Clip3 is a clipping function which is defined as:

Clip

⁢

⁢

3

⁢

(

u

,

v

,

w

)

=

{

u

;

w

<

u

v

;

w

>

v

w

;

otherwise

.

20. A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by a video processing apparatus, wherein the method comprises:

determining, for a current block of a video, an initial prediction sample;

refining the initial prediction sample, based on an optical flow refinement technology, with a prediction sample offset to acquire a final prediction sample; and

generating the bitstream based on the final prediction sample,

wherein the prediction sample offset is determined based on at least one spatial gradient of the initial prediction sample, wherein the spatial gradient is calculated based on at least a difference between two first prediction samples from a same reference picture list, and

wherein before calculating the difference between the two first prediction samples, values of the two first prediction sample are right shifted with a first value.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2021
From: ZHANG, LI; ZHANG, KAI; XU, JIZHENG
To: BYTEDANCE INC.
Reel/Frame 057219/0344 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2021
From: LIU, HONGBIN; WANG, YUE
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 057219/0358 →
Priority Claims (2)
WO PCT/CN2019/081155 · Apr 2, 2019 · international
WO PCT/CN2019/085796 · May 7, 2019 · international
Continuity (2)
Continuation PCTCN2020082937 · Apr 2, 2020
Related Publication 20210385481A1 · Dec 9, 2021