IP Library › Granted Patent US 10,567,770
Granted Patent B2
US 10,567,770 · App. 15/345,315 · Granted Feb 18, 2020

Video decoding implementations for a graphics processing unit

Inventors: Daniel Dinu (Redmond, WA); Juan Carlos Arevalo Baeza (Bellevue, WA); Barry Friemel (Redmond, WA); William Chen (Issaquah, WA)
Assignee: Microsoft Technology Licensing, LLC
H04N19/13G06T1/20H04N19/105H04N19/124H04N19/15H04N19/159H04N19/31H04N19/42H04N19/43H04N19/436H04N19/46H04N19/51H04N19/593H04N19/61H04N19/82H04N19/91H04N19/112H04N19/137H04N19/16H04N19/172H04N19/174H04N19/176H04N19/184H04N19/44H04N19/89
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,567,770
App. No.
15/345,315
Granted
Feb 18, 2020
Kind
B2
Abstract

Video decoding innovations for multithreading implementations and graphics processor unit (“GPU”) implementations are described. For example, for multithreaded decoding, a decoder uses innovations in the areas of layered data structures, picture extent discovery, a picture command queue, and/or task scheduling for multithreading. Or, for a GPU implementation, a decoder uses innovations in the areas of inverse transforms, inverse quantization, fractional interpolation, intra prediction using waves, loop filtering using waves, memory usage and/or performance-adaptive loop filtering. Innovations are also described in the areas of error handling and recovery, determination of neighbor availability for operations such as context modeling and intra prediction, CABAC decoding, computation of collocated information for direct mode macroblocks in B slices, reduction of memory consumption, implementation of trick play modes, and picture dropping for quality adjustment.

Claims (41)

1. A computer system comprising one or more processing units, memory, and storage, wherein the memory and/or the storage has stored therein computer-executable instructions for causing the computer system, when programmed thereby, to perform video processing comprising:

receiving encoded data for a picture that includes plural portions; and

decoding the encoded data to reconstruct the picture, including performing decoding operations for the plural portions of the picture, on a wave-by-wave basis, as plural waves, the plural portions of the picture having variable size, each of the plural waves including one or more of the plural portions of the picture, wherein, for at least one of the plural waves, at least some of the one or more portions within the wave are processed in parallel, wherein the plural waves ripple from a top-left corner of the picture toward a bottom-right corner of the picture, wherein each of the plural waves depends on results of the decoding operations for preceding waves among the plural waves, and wherein, for each of the plural waves, for each given portion of the wave the decoding operations have completed for (1) the portion, if any, left of the given portion, (2) the portion, if any, above-left of the given portion, (3) the portion, if any, above the given portion, and (4) the portion, if any, above-right of the given portion.

2. The computer system of claim 1 , wherein the plural waves roughly correspond to diagonal lines of portions of the picture, and wherein the portions are macroblocks.

3. The computer system of claim 1 , wherein the plural waves are based upon static assumptions of dependencies for the plural portions of the picture.

4. The computer system of claim 1 , wherein the plural waves are based upon actual dependencies for the plural portions of the picture.

5. The computer system of claim 1 , wherein the one or more processing units include a central processing unit (“CPU”) and a graphics processing unit (“GPU”), and wherein execution units of the GPU perform the decoding operations for the plural portions of the picture on a wave-by-wave basis.

6. The computer system of claim 1 , wherein each of the plural portions of the picture is an arrangement of sample values for luma with associated arrangements of sample values for chroma.

7. The computer system of claim 1 , wherein the plural portions of the picture are intra-coded, and wherein the decoding operations include intra prediction operations.

8. The computer system of 28 , wherein, for at least one of the plural waves, the one or more portions of that wave include a first intra portion having a first size and a second intra portion having a second size different than the first size.

9. The computer system of claim 7 , wherein the picture is a P picture or a B picture, and wherein one or more non-intra portions of the picture are omitted from the plural waves.

10. The computer system of claim 7 , wherein the picture is an I picture.

11. The computer system of claim 7 , wherein the performing the intra prediction operations include, for one of the plural waves, processing at least some luma blocks and at least some chroma blocks in parallel.

12. The computer system of claim 7 , wherein the intra prediction operations use plural intra prediction modes, and wherein the performing the intra prediction operations includes applying results of refactored operations for the plural intra prediction modes, the refactored operations reducing branches in implementations of the plural intra prediction modes.

13. The computer system of claim 1 , wherein the decoding operations include loop filtering operations.

14. The computer system of claim 13 , wherein the loop filtering operations include block-by-block processing along a row or column in a macroblock.

15. The computer system of claim 13 , wherein the plural waves are independent of edge strengths of the plural portions of the picture.

16. The computer system of claim 13 , wherein the loop filtering operations include:

in a first loop filtering pass for the picture, calculating boundary strength values; and

in a second loop filtering pass for the picture:

loop filtering plural luma blocks; and

loop filtering plural chroma blocks.

17. The computer system of claim 16 , wherein the second loop filtering pass includes a luma pass for the loop filtering the plural luma blocks and a chroma pass for the loop filtering the plural chroma blocks.

18. A method comprising:

receiving encoded data for a picture that includes plural portions; and

decoding the encoded data to reconstruct the picture, including performing decoding operations for the plural portions of the picture, on a wave-by-wave basis, as plural waves, the plural portions of the picture having variable size, each of the plural waves including one or more of the plural portions of the picture, wherein, for at least one of the plural waves, at least some of the one or more portions within the wave are processed in parallel, wherein the plural waves ripple from a top-left corner of the picture toward a bottom-right corner of the picture, wherein each of the plural waves depends on results of the decoding operations for preceding waves among the plural waves, and wherein, for each of the plural waves, for each given portion of the wave the decoding operations have completed for (1) the portion, if any, left of the given portion, (2) the portion, if any, above-left of the given portion, (3) the portion, if any, above the given portion, and (4) the portion, if any, above-right of the given portion.

19. The method of claim 18 , wherein execution units of a graphics processing unit (“GPU”) perform the decoding operations for the plural portions of the picture on a wave-by-wave basis.

20. The method of claim 18 , wherein the plural portions of the picture are intra-coded, and wherein the decoding operations include intra prediction operations.

21. The method of claim 20 , wherein:

the picture is a P picture or a B picture, and one or more non-intra portions of the picture are omitted from the plural waves; or

the picture is an I picture.

22. The method of claim 18 , wherein the decoding operations include loop filtering operations.

23. A non-volatile memory device having stored therein computer-executable instructions for causing a computer system, when programmed thereby, to perform video processing comprising:

receiving encoded data for a picture that includes plural portions of the picture; and

decoding the encoded data to reconstruct the picture, including performing decoding operations for the plural portions of the picture, on a wave-by-wave basis, as plural waves, the plural portions of the picture having variable size, each of the plural waves including one or more of the plural portions of the picture, wherein, for at least one of the plural waves, at least some of the one or more portions within the wave are processed in parallel, wherein the plural waves ripple from a top-left corner of the picture toward a bottom-right corner of the picture, wherein each of the plural waves depends on results of the decoding operations for preceding waves among the plural waves, and wherein, for each of the plural waves, for each given portion of the wave the decoding operations have completed for (1) the portion, if any, left of the given portion, (2) the portion, if any, above-left of the given portion, (3) the portion, if any, above the given portion, and (4) the portion, if any, above-right of the given portion.

24. The non-volatile memory device of claim 23 , wherein execution units of a graphics processing unit (“GPU”) perform the decoding operations for the plural portions of the picture on a wave-by-wave basis.

25. The non-volatile memory device of claim 23 , wherein the plural portions of the picture are intra-coded, and wherein the decoding operations include intra prediction operations.

26. The non-volatile memory device of claim 25 , wherein:

the picture is a P picture or a B picture, and one or more non-intra portions of the picture are omitted from the plural waves; or

the picture is an I picture.

27. The non-volatile memory device of claim 23 , wherein the decoding operations include loop filtering operations.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2018
From: BAEZA, JUAN CARLOS AREVALO; CHRISTOFFERSEN, ERIC S.; CALLAHAN, SEAN M.; DINU, DANIEL; FRIEMEL, BARRY; CHEN, WILLIAM; ZHAO, WEIDONG; WU, YONGJUN
To: MICROSOFT CORPORATION
Reel/Frame 046913/0917 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2018
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 046913/0943 →
Continuity (2)
Continuation 11824508 · Jun 30, 2007
Related Publication 20170155907A1 · Jun 1, 2017
Cited By (1)
US 12,355,968