IP Library › Granted Patent US 12,223,682
Granted Patent B2
US 12,223,682 · App. 17/357,038 · Granted Feb 11, 2025

Variable width interleaved coding for graphics processing

Inventors: Stephen Junkins (Bend, OR); Sreenivas Kothandaraman (Sammamish, WA); Prasoonkumar Surti (Folsom, CA); Srihari Pratapa (Seattle, WA); William Hux (Hillsboro, OR); John Feit (Folsom, CA)
Assignee: INTEL CORPORATION
G06T9/00G06T1/20G06T1/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,682
App. No.
17/357,038
Granted
Feb 11, 2025
Kind
B2
Abstract

Variable width interleaved coding for graphics processing is described. An example of an apparatus includes one or more processors including a graphic processor; and memory for storage of data including data for graphics processing, wherein the graphics processor includes an encoder pipeline to provide variable width interleaved coding and a decoder pipeline to decode the variable width interleaved coding, and wherein the encoder pipeline is to receive a plurality of bitstreams from workgroups; perform parallel entropy encoding on the bitstreams to generate a plurality of encoded bitstreams for each of the workgroups; perform variable interleaving of the bitstreams for each workgroup based at least in part on data requirements for decoding received from the decoder pipeline; and compact outputs for each of the workgroups into a contiguous stream of interleaved data.

Claims (53)

1. An apparatus comprising:

one or more processors including a graphic processor; and

memory for storage of data including data for graphics processing;

wherein the graphics processor includes an encoder pipeline to provide variable width interleaved coding and a decoder pipeline to decode the variable width interleaved coding, and wherein the encoder pipeline is to:

receive a plurality of bitstreams from a plurality of workgroups;

perform parallel entropy encoding on the bitstreams to generate a plurality of encoded bitstreams for each of the workgroups;

perform variable interleaving of the bitstreams for each workgroup of the plurality of workgroups, wherein a number of data elements that are interleaved for each workgroup in each of a plurality of iterations is based at least in part on current data requirements for decoding for each workgroup received from the decoder pipeline; and

compact interleaved bitstream outputs for each of the plurality of workgroups into a contiguous stream of interleaved data.

2. The apparatus of claim 1 , wherein the entropy encoding comprises Huffman coding.

3. The apparatus of claim 1 , wherein the received plurality of bitstreams include Single Instruction Multiple Data (SIMD) bitstreams.

4. The apparatus of claim 1 , wherein:

the encoder pipeline includes a plurality of parallel entropy encoders for each of the plurality of workgroups; and

the decoder pipeline includes a plurality of parallel entropy decoders for each of the plurality of workgroups.

5. The apparatus of claim 4 , wherein bitstream write patterns produced by the parallel entropy encoders for the plurality of workgroups are to be synchronized with bitstream read patterns of the parallel entropy decoders for the plurality of workgroups.

6. The apparatus of claim 1 , wherein the encoded bitstreams may be of variable lengths.

7. The apparatus of claim 6 , wherein a number of encoded bitstreams that are interleaved in the variable interleaving may vary with each iteration of the encoder pipeline.

8. The apparatus of claim 1 , wherein the decoder pipeline is to:

receive the contiguous data stream;

provide the current data requirements for decoding for each workgroup to the encoder pipeline;

separate the contiguous data stream into a plurality of sets of data based at least in part on the data requirements for decoding for each workgroup provided to the encoder pipeline; and

perform parallel entropy decoding of the plurality of sets of data for each workgroup to generate a plurality of bitstreams.

9. The apparatus of claim 8 , wherein the performance of parallel entropy decoding includes use of an entropy code lookup table.

10. The apparatus of claim 1 , wherein the encoder pipeline and the decoder pipeline utilize a kernel workgroup size that is equal to a target SIMD width.

11. One or more non-transitory computer-readable storage mediums having stored thereon executable computer program instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

receiving a plurality of bitstreams from a plurality of workgroups at an encoder pipeline;

performing parallel entropy encoding on the bitstreams to generate a plurality of encoded bitstreams for each of the workgroups;

performing variable interleaving of the bitstreams for each workgroup of the plurality of workgroups, wherein a number of data elements that are interleaved for each workgroup in each of a plurality of iterations is based at least in part on current data requirements for decoding received from a decoder pipeline; and

compacting interleaved bitstream outputs for each of the plurality of workgroups into a contiguous stream of interleaved data.

12. The one or more storage mediums of claim 11 , wherein:

the encoder pipeline includes a plurality of parallel entropy encoders for each of the plurality of workgroups; and

the decoder pipeline includes a plurality of parallel entropy decoders for each of the plurality of workgroups.

13. The one or more storage mediums of claim 11 , wherein the encoded bitstreams may be of variable lengths, and wherein a number of encoded bitstreams that are interleaved in the variable interleaving may vary with each iteration of the encoder pipeline.

14. The one or more storage mediums of claim 11 , wherein the instructions further include instructions for:

receiving the contiguous data stream at the decoder pipeline;

providing the current data requirements for decoding for each workgroup to the encoding pipeline;

separating the contiguous data stream into a plurality of sets of data based at least in part on the data requirements for decoding for each workgroup provided to the encoding pipeline; and

performing parallel entropy decoding of the plurality of sets of data for each workgroup to generate a plurality of bitstreams.

15. The one or more storage mediums of claim 14 , wherein performing parallel entropy decoding includes use of an entropy code lookup table.

16. A method comprising:

receiving a plurality of bitstreams from a plurality of workgroups at an encoder pipeline;

performing parallel entropy encoding on the bitstreams to generate a plurality of encoded bitstreams for each of the workgroups;

performing variable interleaving of the bitstreams for each workgroup of the plurality of workgroups, wherein a number of data elements that are interleaved for each workgroup in each of a plurality of iterations is based at least in part on current data requirements for decoding received from a decoder pipeline; and

compacting interleaved bitstream outputs for each of the workgroups into a contiguous stream of interleaved data.

17. The method of claim 16 , wherein:

the encoder pipeline includes a plurality of parallel entropy encoders for each of the plurality of workgroups; and

the decoder pipeline includes a plurality of parallel entropy decoders for each of the plurality of workgroups.

18. The method of claim 16 , wherein the encoded bitstreams may be of variable lengths, and wherein a number of encoded bitstreams that are interleaved in the variable interleaving may vary with each iteration of the encoder pipeline.

19. The method of claim 16 , further comprising:

receiving the contiguous data stream at the decoder pipeline;

providing the current data requirements for decoding for each workgroup to the encoding pipeline;

separating the contiguous data stream into a plurality of sets of data based at least in part on the data requirements for decoding for each workgroup provided to the encoding pipeline; and

performing parallel entropy decoding of the plurality of sets of data for each workgroup to generate a plurality of bitstreams.

20. The method of claim 19 , wherein performing parallel entropy decoding includes use of an entropy code lookup table.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2021
From: JUNKINS, STEPHEN; KOTHANDARAMAN, SREENIVAS; SURTI, PRASOONKUMAR; PRATAPA, SRIHARI; HUX, WILLIAM; FEIT, JOHN
To: INTEL CORPORATION
Reel/Frame 057427/0979 →
Continuity (2)
Provisional Application 63163685 · Mar 19, 2021
Related Publication 20220301228A1 · Sep 22, 2022
References Cited (29)
US 3978319A · Vinal · 1976 [cited by examiner]
US 20030128884A1 · Lee et al. · 2003 [cited by applicant]
US 20030147470A1 · Lee et al. · 2003 [cited by applicant]
US 20060238390A1 · Yi · 2006 [cited by applicant]
US 20070053600A1 · Lee et al. · 2007 [cited by applicant]
US 20070183674A1 · Lee et al. · 2007 [cited by applicant]
US 20090006510A1 · Laker · 2009 [cited by examiner]
US 20100191859A1 · Raveendran · 2010 [cited by examiner]
US 20100322308A1 · Lee et al. · 2010 [cited by applicant]
US 20120081242A1 · He et al. · 2012 [cited by applicant]
US 20180293692A1 · Koker · 2018 [cited by examiner]
US 20190005703A1 · Golas · 2019 [cited by examiner]
US 20220159255A1 · Xu · 2022 [cited by examiner]
US 20230136121A1 · Leleannec et al. · 2023 [cited by applicant]
DE 102022101975A1 · 2022 [cited by applicant]
WO 2009005758A3 · 2009 [cited by applicant]
Sitaridi Eva et al: “Parallel lossless compression using GPUs”, GPU Tech Cont, Dec. 31, 2014 (Dec. 31, 2014), pp. 1-42, XP093015123,. [cited by examiner]
Various: “can there be some performance gain from using lizard in highly parallel compute?”, encode.su, Jul. 23, 2019 (Jul. 23, 2019), pp. 1-8, XP093015141. [cited by examiner]
Andrew Yeung. (Sep. 2020). DirectStorage is coming to PC. Retrieved from devblogs.microsoft. com: https://devblogs.microsoft.com/directx/directstorage-is-coming-to-pc/. [cited by applicant]
Burnes, A. (Sep. 2020). Introducing NVIDIA RTX IO: GPU-Accelerated Storage Technology For The Next Generation of Games. Retrieved from Nvidia.com: https://www.nvidia.com/en-us/geforce/news/rtx-io-gpu-accelerated-storage… [cited by applicant]
Duda, J. (2014). Asymmetric numeral systems: entropy coding combining speed of Huffman coding with compression rate of arithmetic coding. Retrieved from https://arxiv.org/abs/1311.2540. [cited by applicant]
Giesen, F. (2014). Interleaved Entropy Coders. Retrieved from https://arxiv.org/abs/1402.3392. [cited by applicant]
Weissenberger, A., & Schmidt, B. (2018). Massively Parallel Huffman Decoding on GPUs. ACM ICPP 2018: 47th Internation Conference on Parallel Processing. Eugene, Oregon. Retrieved fromhttps://dl.acm.org/doi/10.1145/32250… [cited by applicant]
Yamamoto, N., Nakano, K., Ito, Y., Takafuji, D., Kasagi, A., & Tabaru, T. (2020). Huffman Coding with Gap Arrays for GPU Acceleration. ACM IJCP 2020; International Conference on Parallel Processing. ACM. Retrieved from … [cited by applicant]
Extended European Search Report for EP22183479.9 mailed Jan. 30, 2023, 14 pages. [cited by applicant]
Sitaridi Eva et al: “Parallel lossless compression using GPUs”, GPU Tech Conf, Dec. 31, 2014 (Dec. 31, 2014), pp. 1-42, XP093015123, Retrieved from the Internet: URL:https://on-demand.gputechconf.com/gtc/20 14/presentat… [cited by applicant]
Various: “can there be some performance gain from using lizard in highly parallel compute?”, encode.su, Jul. 23, 2019 (Jul. 23, 2019), pp. 1-8, XP093015141, Retrieved from the Internet: URL:https://encode.su/threads/315… [cited by applicant]
Notification of Publication for Chinese Patent Application No. 202210154266.6, Publication No. CN 115115719 A, mailed Oct. 11, 2022, 2 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/854,310 mailed Sep. 12, 2024, 13 pages. [cited by applicant]
Cited By (1)
US 12,561,752