IP Library › Granted Patent US 12,238,335
Granted Patent B2
US 12,238,335 · App. 18/302,625 · Granted Feb 25, 2025

Efficient sub-pixel motion vector search for high-performance video encoding

Inventors: Yongmao Tang (Shanghai, CN); Jianjun Chen (Shanghai, CN); Junan Chen (Jiangsu, CN); Yonghai Wu (Shanghai, CN); Zejun Hu (Jiangsu, CN); Wei Feng (Shanghai, CN)
Assignee: NVIDIA Corporation
H04N19/61H04N19/105H04N19/109H04N19/122H04N19/139H04N19/156H04N19/172H04N19/176H04N19/182H04N19/423
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,238,335
App. No.
18/302,625
Granted
Feb 25, 2025
Kind
B2
Abstract

Disclosed are systems and techniques for efficient real-time codec encoding of video files. In one embodiment, the techniques include obtaining a first plurality of motion vectors of a first resolution, generating a second plurality of motion vectors of a second resolution, and calculating a first cost of the motion vector using a first cost function of a first size. The techniques include selecting a subset of motion vectors of the second plurality of motion vectors, calculating a second cost using a second cost function of a second size, and generating a plurality of combined motion vectors based on the subset of motion vectors. The techniques include calculating a third cost using the second cost function of the second size, selecting a final motion vector, and generating, based on the selected final motion vector, a block of predicted pixels that approximates a block of source pixels of an image frame.

Claims (83)

1. A method comprising:

obtaining a first plurality of motion vectors of a first resolution;

generating a second plurality of motion vectors of a second resolution based on the first plurality of motion vectors of the first resolution;

calculating, for each motion vector of the second plurality of motion vectors, a first cost of the motion vector using a first cost function of a first size;

selecting a subset of motion vectors of the second plurality of motion vectors based on the first cost of each motion vector;

calculating, for each motion vector of the subset of motion vectors, a second cost using a second cost function of a second size;

generating a plurality of combined motion vectors based on the subset of motion vectors;

calculating, for each motion vector of the plurality of combined motion vectors, a third cost using the second cost function of the second size;

selecting, based on the cost of each motion vector of the subset of motion vectors and the cost of each motion vector of the plurality of combined motion vectors, a final motion vector; and

generating, based on the final motion vector, a block of predicted pixels that approximates a block of source pixels of an image frame.

2. The method of claim 1 , wherein calculating the first cost of the motion vector comprises, for at least a first motion vector of the second plurality of motion vectors:

obtaining pixel values of a corresponding reconstructed reference frame;

transforming the pixel values based on the first motion vector;

calculating a difference between the transformed pixel values and corresponding source pixel values; and

calculating the first cost using the first cost function of the first size based on the calculated difference.

3. The method of claim 1 , wherein calculating the first cost of the motion vector comprises, for at least a first motion vector of the second plurality of motion vectors, retrieving, from a buffer, the first cost of the first motion vector of the second plurality of motion vectors.

4. The method of claim 1 , wherein calculating the third cost comprises, for at least a first combined motion vector of the plurality of combined motion vectors:

retrieving, from a buffer, predicted pixel values corresponding to a first motion vector of the first combined motion vector;

obtaining pixel values of a reconstructed reference frame corresponding to a second motion vector of the first combined motion vector;

transforming the pixel values of the reconstructed reference frame based on the second motion vector;

calculating average pixel values based on the predicted pixel values and the transformed pixel values;

calculating a difference between the average pixel values and corresponding source pixel values; and

calculating the third cost using the second cost function of the second size based on the calculated difference.

5. The method of claim 1 , wherein the first resolution is an integer-pixel resolution and wherein the second resolution comprises at least one of a half-pixel resolution or a quarter-pixel resolution.

6. The method of claim 1 , wherein the first cost function of the first size is a 4×4 sum of absolute Hadamard-transformed difference function.

7. The method of claim 1 , wherein the second cost function of the second size is an 8×8 sum of absolute Hadamard-transformed difference function.

8. A system comprising:

a memory device; and

one or more circuit groups communicatively coupled to the memory device, the one or more circuit groups to:

obtain a first plurality of motion vectors of a first resolution;

generate a second plurality of motion vectors of a second resolution based on the first plurality of motion vectors of the first resolution;

calculate, for each motion vector of the second plurality of motion vectors, a first cost of the motion vector using a first cost function of a first size;

select a subset of motion vectors of the second plurality of motion vectors based on the first cost of each motion vector;

calculate, for each motion vector of the subset of motion vectors, a second cost using a second cost function of a second size;

generate a plurality of combined motion vectors based on the subset of motion vectors;

calculate, for each motion vector of the plurality of combined motion vectors, a third cost using the second cost function of the second size;

select, based on the cost of each motion vector of the subset of motion vectors and the cost of each motion vector of the plurality of combined motion vectors, a final motion vector; and

generate, based on the final motion vector, a block of predicted pixels that approximates a block of source pixels of an image frame.

9. The system of claim 8 , wherein to calculate the first cost of the motion vector, the one or more circuit groups are further to, for at least a first motion vector of the second plurality of motion vectors:

obtain pixel values of a corresponding reconstructed reference frame;

transform the pixel values based on the first motion vector;

calculate a difference between the transformed pixel values and corresponding source pixel values; and

calculate the first cost using the first cost function of the first size based on the calculated difference.

10. The system of claim 8 , wherein to calculate the first cost of the motion vector, the one or more circuit groups are further to, for at least a first motion vector of the second plurality of motion vectors, retrieve, from a buffer, the first cost of the first motion vector of the second plurality of motion vectors.

11. The system of claim 8 , wherein to calculate the third cost, the one or more circuit groups are further to, for at least a first combined motion vector of the plurality of combined motion vectors:

retrieve, from a buffer, predicted pixel values corresponding to a first motion vector of the first combined motion vector;

obtain pixel values of a reconstructed reference frame corresponding to a second motion vector of the first combined motion vector;

transform the pixel values of the reconstructed reference frame based on the second motion vector;

calculate average pixel values based on the predicted pixel values and the transformed pixel values;

calculate a difference between the average pixel values and corresponding source pixel values; and

calculate the third cost using the second cost function of the second size based on the calculated difference.

12. The system of claim 8 , wherein the first resolution is an integer-pixel resolution and wherein the second resolution comprises at least one of a half-pixel resolution or a quarter-pixel resolution.

13. The system of claim 8 , wherein the first cost function of the first size is a 4×4 sum of absolute Hadamard-transformed difference function.

14. The system of claim 8 , wherein the second cost function of the second size is an 8×8 sum of absolute Hadamard-transformed difference function.

15. A system comprising:

a memory device; and

one or more circuit groups communicatively coupled to the memory device, the one or more circuit groups comprising:

a first circuit group to:

obtain a first plurality of motion vectors of a first resolution;

generate a second plurality of motion vectors of a second resolution based on the first plurality of motion vectors of the first resolution;

calculate, for each motion vector of the second plurality of motion vectors, a first cost of the motion vector using a first cost function of a first size; and

select a subset of motion vectors of the second plurality of motion vectors based on the first cost of each motion vector; and

a second circuit group communicatively coupled to the first circuit group, the second circuit group to:

calculate, for each motion vector of the subset of motion vectors, a second cost using a second cost function of a second size;

generate a plurality of combined motion vectors based on the subset of motion vectors;

calculate, for each motion vector of the plurality of combined motion vectors, a third cost using the second cost function of the second size;

select, based on the cost of each motion vector of the subset of motion vectors and the cost of each motion vector of the plurality of combined motion vectors, a final motion vector; and

generate, based on the final motion vector, a block of predicted pixels that approximates a block of source pixels of an image frame.

16. The system of claim 15 , wherein to calculate the first cost of the motion vector, the first circuit group is further to, for at least a first motion vector of the second plurality of motion vectors:

obtain pixel values of a corresponding reconstructed reference frame;

transform the pixel values based on the first motion vector;

calculate a difference between the transformed pixel values and corresponding source pixel values; and

calculate the first cost using the first cost function of the first size based on the calculated difference.

17. The system of claim 15 , wherein to calculate the first cost of the motion vector, the first circuit group is further to, for at least a first motion vector of the second plurality of motion vectors, retrieve, from a buffer, the first cost of the first motion vector of the second plurality of motion vectors.

18. The system of claim 15 , wherein to calculate the third cost, the second circuit group is further to, for at least a first combined motion vector of the plurality of combined motion vectors:

retrieve, from a buffer, predicted pixel values corresponding to a first motion vector of the first combined motion vector;

obtain pixel values of a reconstructed reference frame corresponding to a second motion vector of the first combined motion vector;

transform the pixel values of the reconstructed reference frame based on the second motion vector;

calculate average pixel values based on the predicted pixel values and the transformed pixel values;

calculate a difference between the average pixel values and corresponding source pixel values; and

calculate the third cost using the second cost function of the second size based on the calculated difference.

19. The system of claim 15 , wherein the first resolution is an integer-pixel resolution and wherein the second resolution comprises at least one of a half-pixel resolution or a quarter-pixel resolution.

20. The system of claim 15 , wherein the second cost function of the second size is an 8×8 sum of absolute Hadamard-transformed difference function.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2023
From: TANG, YONGMAO; CHEN, JIANJUN; CHEN, JUNAN; WU, YONGHAI; HU, ZEJUN; FENG, WEI
To: NVIDIA CORPORATION
Reel/Frame 063549/0088 →
Continuity (1)
Related Publication 20240357175A1 · Oct 24, 2024
References Cited (85)
US 8311111B2 · Xu et al. · 2012 [cited by applicant]
US 9432668B1 · Bossen et al. · 2016 [cited by applicant]
US 9998726B2 · Rusanovskyy et al. · 2018 [cited by applicant]
US 10070128B2 · Ugur et al. · 2018 [cited by applicant]
US 10091514B1 · Bossen et al. · 2018 [cited by applicant]
US 10621731B1 · Duenas et al. · 2020 [cited by applicant]
US 10687054B2 · Mahdi et al. · 2020 [cited by applicant]
US 10887611B2 · Seregin et al. · 2021 [cited by applicant]
US 10999594B2 · Hsieh et al. · 2021 [cited by applicant]
US 11017566B1 · Tourapis et al. · 2021 [cited by applicant]
US 11057636B2 · Huang et al. · 2021 [cited by applicant]
US 11070813B2 · Socek et al. · 2021 [cited by applicant]
US 11172195B2 · Seregin · 2021 [cited by applicant]
US 11197009B2 · Zhang et al. · 2021 [cited by applicant]
US 11202070B2 · Zhang et al. · 2021 [cited by applicant]
US 11218694B2 · Seregin et al. · 2022 [cited by applicant]
US 11272201B2 · Seregin et al. · 2022 [cited by applicant]
US 11317111B2 · Rusanovskyy et al. · 2022 [cited by applicant]
US 11343504B2 · Zhao et al. · 2022 [cited by applicant]
US 11368684B2 · Seregin et al. · 2022 [cited by applicant]
US 11388394B2 · Seregin et al. · 2022 [cited by applicant]
US 11496746B2 · Siddaramanna et al. · 2022 [cited by applicant]
US 11563933B2 · Seregin et al. · 2023 [cited by applicant]
US 11582475B2 · Rusanovskyy et al. · 2023 [cited by applicant]
US 11638025B2 · Pourreza et al. · 2023 [cited by applicant]
US 11638062B2 · Stockhammer et al. · 2023 [cited by applicant]
US 11677987B2 · Said · 2023 [cited by applicant]
US 20100118945A1 · Wada et al. · 2010 [cited by applicant]
US 20130016783A1 · Kim et al. · 2013 [cited by applicant]
US 20130101035A1 · Wang et al. · 2013 [cited by applicant]
US 20130259142A1 · Ikeda et al. · 2013 [cited by applicant]
US 20140198844A1 · Hsu et al. · 2014 [cited by applicant]
US 20150229921A1 · Chen et al. · 2015 [cited by applicant]
US 20160225161A1 · Hepper · 2016 [cited by examiner]
US 20160330445A1 · Ugur et al. · 2016 [cited by applicant]
US 20170085886A1 · Jacobson et al. · 2017 [cited by applicant]
US 20170142438A1 · Hepper · 2017 [cited by examiner]
US 20170201769A1 · Chon et al. · 2017 [cited by applicant]
US 20170272758A1 · Lin et al. · 2017 [cited by applicant]
US 20200021847A1 · Kim et al. · 2020 [cited by applicant]
US 20200029096A1 · Rusanovskyy · 2020 [cited by applicant]
US 20200099926A1 · Tanner et al. · 2020 [cited by applicant]
US 20200104976A1 · Mammou et al. · 2020 [cited by applicant]
US 20200105024A1 · Mammou et al. · 2020 [cited by applicant]
US 20200111237A1 · Tourapis et al. · 2020 [cited by applicant]
US 20200137415A1 · Esenlik · 2020 [cited by examiner]
US 20200204829A1 · Stepin et al. · 2020 [cited by applicant]
US 20200288122A1 · Kim · 2020 [cited by applicant]
US 20200359022A1 · Abe et al. · 2020 [cited by applicant]
US 20200382777A1 · Zhang et al. · 2020 [cited by applicant]
US 20200382804A1 · Zhang et al. · 2020 [cited by applicant]
US 20210006833A1 · Tourapis et al. · 2021 [cited by applicant]
US 20210021809A1 · Kim · 2021 [cited by applicant]
US 20210099701A1 · Tourapis et al. · 2021 [cited by applicant]
US 20210211661A1 · Toma et al. · 2021 [cited by applicant]
US 20210211703A1 · Kim et al. · 2021 [cited by applicant]
US 20210211724A1 · Kim et al. · 2021 [cited by applicant]
US 20210217203A1 · Kim et al. · 2021 [cited by applicant]
US 20210321093A1 · Sundaram et al. · 2021 [cited by applicant]
US 20210377868A1 · Anand · 2021 [cited by applicant]
US 20210392334A1 · Esenlik et al. · 2021 [cited by applicant]
US 20220021891A1 · Chaudhari et al. · 2022 [cited by applicant]
US 20220256169A1 · Siddaramanna et al. · 2022 [cited by applicant]
US 20220277164A1 · Malayath · 2022 [cited by applicant]
US 20220279204A1 · Malayath et al. · 2022 [cited by applicant]
US 20230063062A1 · Srinivasan et al. · 2023 [cited by applicant]
US 20230071018A1 · Tang et al. · 2023 [cited by applicant]
CN 108449603A · 2018 [cited by applicant]
CN 111918058A · 2020 [cited by applicant]
CN 113301347A · 2021 [cited by applicant]
JP 2014127891A · 2014 [cited by applicant]
WO 2010030752A2 · 2010 [cited by applicant]
WO 2012030752A2 · 2012 [cited by applicant]
WO 2013067903A1 · 2013 [cited by applicant]
WO 2019163794A1 · 2019 [cited by applicant]
Chen Y., et al., “An Overview of Core Coding Tools in AV1 Video Codec,” Picture Coding Symposium (PCS), Jun. 24-27, 2018, 5 Pages, DOI: 10.1109/PCS.2018.8456249. [cited by applicant]
Goebel et al., “Hardware Design of DC/CFL Intra-Prediction Decoder for AV1 Codec,” 32nd Symposium of Integrated Circuits and Systems Design (SBCCI), Sao Paulo, Brazil, Aug. 26-30, 2019, pp. 1-6. [cited by applicant]
Han J., et al., “A Technical Overview of AV1,” arXiv:2008.06091v2 [eess.IV], Feb. 8, 2021, 25 Pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/CN2021/116311, mailed May 31, 2022, 6 Pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/CN2021/116312, mailed May 26, 2022, 9 Pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/CN2021/116711, mailed Apr. 24, 2022, 7 Pages. [cited by applicant]
Final Office Action for U.S. Appl. No. 17/451,974, mailed Jan. 25, 2024, 13 Pages. [cited by applicant]
ITU-T; H.265 (Year: 2016). [cited by applicant]
ITU-T; H.266 (Year: 2020). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/451,972, mailed Dec. 28, 2023, 51 Pages. [cited by applicant]