IP Library › Granted Patent US 12,561,752
Granted Patent B2
US 12,561,752 · App. 18/085,367 · Granted Feb 24, 2026

Apparatus and method for a framework for GPU-driven data loading

Inventor: Minmin Gong (Kirkland, WA)
Assignee: TENCENT AMERICA LLC
G06T1/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,752
App. No.
18/085,367
Granted
Feb 24, 2026
Kind
B2
Abstract

In a method of data loading at a computing device, data to load is identified at a graphics processing unit (GPU) based on execution of an application program. Data chunks of the identified data in encoded form are loaded via the GPU from a data storage device to a video memory associated with the GPU. The data chunks are decoded in parallel by using plural GPU thread groups in parallel to decode the data chunks. Each of the data chunks is decoded independently of other data chunks. Apparatus, device, and non-transitory computer-readable storage medium counterparts are also contemplated.

Claims (54)

1 . A method of data loading at a computing device including a central processing unit (CPU) and a graphics processing unit (GPU), the method comprising:

identifying, by the GPU, data chunks to load from a plurality of candidate data chunks based on execution of an application program;

bypassing the CPU by directly loading, by the GPU, the identified data chunks of the data in encoded form directly from a data storage device to a video memory that is separate from the data storage device;

decoding, by the GPU, the data chunks in parallel by using plural GPU thread groups in parallel to decode the data chunks, wherein each of the data chunks is decoded independently of other data chunks; and

processing, by the GPU, the decoded data chunks to render the decoded data for visual representation.

2 . The method of claim 1 , wherein each of the plural GPU thread groups decodes a respective one of the data chunks independently of other data chunks by employing in-group decoding functions limited to using data from the respective thread group.

3 . The method of claim 2 , wherein the decoding the data chunks in parallel includes decoding a first data chunk using a first GPU thread group of the plural GPU thread groups in parallel with and independently from decoding a second data chunk using a second GPU thread group of the plural thread groups.

4 . The method of claim 2 , wherein the in-group decoding functions include a shuffle function that allows an inter-thread data exchange within one of the plural GPU thread groups.

5 . The method of claim 2 , wherein the in-group decoding functions include a radix sort of one or more threads within one of the plural GPU thread groups, the radix sort having an input bit length that is less than or equal to a number of threads in one of the plural GPU thread groups.

6 . The method of claim 2 , wherein the in-group decoding functions include a merge sort applicable to a maximum number of elements equal to a product of a number of elements in one thread and a number of threads in a GPU thread group.

7 . The method of claim 2 , wherein the in-group decoding functions include a match operation that includes iteratively (i) copying a non-overlap region split from an overlap region of the data, and (ii) dividing the overlap region into a subsequent non-overlap region and a subsequent overlap region until a length of the subsequent overlap region is less than a distance of the subsequent overlap region.

8 . The method of claim 1 , wherein the directly loading the data chunks to the video memory includes:

mapping a file of the data to a memory block; and

associating the memory block mapped to the file of the data with a buffer using an application program interface (API) extension group.

9 . The method of claim 1 , wherein the data chunks correspond to video data, texture data, mesh data, neural network data, or text data.

10 . The method of claim 1 , wherein the processing includes rendering the decoded data for the visual representation.

11 . An apparatus for data loading at a computing device including a central processing unit (CPU) and a graphics processing unit (GPU), the apparatus comprising:

the GPU configured to:

identify data chunks to load from a plurality of candidate data chunks based on execution of an application program;

directly load the identified data chunks of data in encoded form directly from a data storage device to a video memory that is separate from the data storage device;

decode the data chunks in parallel by using plural GPU thread groups in parallel to decode the data chunks, wherein each of the data chunks is decoded independently of other data chunks; and

process the decoded data chunks to render the decoded data for visual representation.

12 . The apparatus of claim 11 , wherein each of the plural GPU thread groups decodes a respective one of the data chunks independently of other data chunks by employing in-group decoding functions limited to using data from the respective thread group.

13 . The apparatus of claim 12 , wherein the GPU is configured to decode a first data chunk using a first GPU thread group of the plural GPU thread groups in parallel with and independently from a second data chunk using a second GPU thread group of the plural thread groups.

14 . The apparatus of claim 12 , wherein the in-group decoding functions include a shuffle function that allows an inter-thread data exchange within one of the plural GPU thread groups.

15 . The apparatus of claim 12 , wherein the in-group decoding functions include a radix sort of one or more threads within one of the plural GPU thread groups, the radix sort having an input bit length that is less than or equal to a number of threads in one of the plural GPU thread groups.

16 . The apparatus of claim 12 , wherein the in-group decoding functions include a merge sort applicable to a maximum number of elements equal to a product of a number of elements in one thread and a number of threads in a GPU thread group.

17 . The apparatus of claim 12 , wherein the in-group decoding functions include a match operation that includes iteratively (i) copying a non-overlap region split from an overlap region of the data, and (ii) dividing the overlap region into a subsequent non-overlap region and a subsequent overlap region until a length of the subsequent overlap region is less than a distance of the subsequent overlap region.

18 . The apparatus of claim 11 , wherein the GPU is configured to:

map a file of the data to a memory block; and

associate the memory block mapped to the file of the data with a buffer using an application program interface (API) extension group.

19 . The apparatus of claim 11 , wherein the data chunks correspond to video data, texture data, mesh data, neural network data, or text data.

20 . A non-transitory computer-readable storage medium storing instructions which when executed by a graphics processing unit (GPU) cause the GPU to perform:

identifying data chunks to load from a plurality of candidate data chunks based on execution of an application program;

bypassing a central processing unit (CPU) by directly loading the identified data chunks of data in encoded form directly from a data storage device to a video memory that is separate from the data storage device; and

decoding the data chunks in parallel by using plural GPU thread groups in parallel to decode the data chunks, wherein each of the data chunks is decoded independently of other data chunks; and

processing the decoded data chunks to render the decoded data for visual representation.

21 . A device, comprising:

a controller in a graphics processing unit (GPU), the controller being configured to:

identify data chunks to load from a plurality of candidate data chunks based on execution of an application program;

bypass a central processing unit (CPU) by directly load the identified data chunks of data in encoded form directly from a data storage device to a video memory that is separate from the data storage device;

decode the data chunks in parallel by using plural GPU thread groups in parallel to decode the data chunks, wherein each of the data chunks is decoded independently of other data chunks; and

process the decoded data chunks to render the decoded data for visual representation.

22 . The device of claim 21 , wherein each of the plural GPU thread groups decodes a respective one of the data chunks independently of other data chunks by employing in-group decoding functions limited to using data from the respective thread group.

23 . The device of claim 22 , wherein the controller is configured to decode a first data chunk using a first GPU thread group of the plural GPU thread groups in parallel with and independently from a second data chunk using a second GPU thread group of the plural thread groups.

24 . The device of claim 22 , wherein the in-group decoding functions include a shuffle function that allows an inter-thread data exchange within one of the plural GPU thread groups.

25 . The device of claim 22 , wherein the in-group decoding functions include a radix sort of one or more threads within one of the plural GPU thread groups, the radix sort having an input bit length that is less than or equal to a number of threads in one of the plural GPU thread groups.

26 . The device of claim 22 , wherein the in-group decoding functions include a merge sort applicable to a maximum number of elements equal to a product of a number of elements in one thread and a number of threads in a GPU thread group.

27 . The device of claim 22 , wherein the in-group decoding functions include a match operation that includes iteratively (i) copying a non-overlap region split from an overlap region of the data, and (ii) dividing the overlap region into a subsequent non-overlap region and a subsequent overlap region until a length of the subsequent overlap region is less than a distance of the subsequent overlap region.

28 . The device of claim 22 , wherein the controller is configured to:

map a file of the data to a memory block; and

associate the memory block mapped to the file of the data with a buffer using an application program interface (API) extension group.

29 . The device of claim 22 , wherein the data chunks correspond to video data, texture data, mesh data, neural network data, or text data.

30 . The device of claim 21 , wherein the controller is further configured to render the decoded data for the visual representation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2022
From: GONG, MINMIN
To: TENCENT AMERICA LLC
Reel/Frame 062163/0691 →
Continuity (1)
Related Publication 20240202859A1 · Jun 20, 2024
References Cited (35)
US 8396122B1 · Taylor · 2013 [cited by examiner]
US 8537899B1 · Taylor · 2013 [cited by examiner]
US 9311721B1 · Loughry · 2016 [cited by examiner]
US 12223682B2 · Junkins · 2025 [cited by examiner]
US 12243118B2 · Fontaine · 2025 [cited by examiner]
US 20070291846A1 · Hussain · 2007 [cited by applicant]
US 20080276262A1 · Munshi · 2008 [cited by examiner]
US 20110063296A1 · Bolz · 2011 [cited by examiner]
US 20110109639A1 · Sreenivas · 2011 [cited by examiner]
US 20120320070A1 · Arvo · 2012 [cited by examiner]
US 20140184606A1 · de Richebourg · 2014 [cited by examiner]
US 20150262385A1 · Satoh et al. · 2015 [cited by applicant]
US 20160055608A1 · Frascati · 2016 [cited by examiner]
US 20160366424A1 · Wu · 2016 [cited by examiner]
US 20170091029A1 · Cho · 2017 [cited by examiner]
US 20170124166A1 · Thomas · 2017 [cited by examiner]
US 20170318068A1 · Kikkeri Shivadatta · 2017 [cited by examiner]
US 20190129950A1 · Sarkar · 2019 [cited by examiner]
US 20210001220A1 · Cerny · 2021 [cited by examiner]
US 20210184795A1 · Ibars Casas · 2021 [cited by examiner]
US 20220036498A1 · Li · 2022 [cited by examiner]
US 20220124335A1 · Dinu et al. · 2022 [cited by applicant]
US 20220270203A1 · Tu et al. · 2022 [cited by applicant]
US 20230244470A1 · Shimon · 2023 [cited by examiner]
US 20230336731A1 · Dinu · 2023 [cited by examiner]
US 20240061943A1 · Durham · 2024 [cited by examiner]
JP 2015176492A · 2015 [cited by applicant]
International Search Report with Written Opinion issued in Application No. PCT/US2023/064531, mailed Jul. 7, 2023, 9 pages. [cited by applicant]
Memory-mapped files, .NET fundamentals documentation, https://learn.microsoft.com/en-US/dotnet/standard/io/ memory-mapped-files, Dec. 14, 2022, pp. 1-10. [cited by applicant]
VK_EXT_external_memory_host(3) Manual p. https://registry.khronos.org/vulkan/specs/1.3-extensions/man/html/ VK_EXT_external_memory_host.html, pp. 1-4 Nov. 10, 2017. [cited by applicant]
VK_KHR_external_memory_fd(3) Manual p. https://registry.khronos.org/vulkan/specs/1.3-extensions/man/html/ VK_KHR_external_memory_fd.html, pp. 1-3 Oct. 21, 2016. [cited by applicant]
DirectStorage Overview, Game Development Kit documentation, https://learn.microsoft.com/en-US/gaming/gdk/ _content/gc/system/overviews/directstorage/directstorage-overview, Mar. 15, 2023, pp. 1-17. [cited by applicant]
Wikipedia contributors. “LZ77 and LZ78.” Wikipedia, The Free Encyclopedia. Wikipedia, The Free Encyclopedia, Jan. 31, 2024. Web. Feb. 13, 2024, pp. 1-7. [cited by applicant]
Saher Odeh, et al., Merge Path - Parallel Merging Made Simple, Conference: Parallel and Distributed Processing Symposium Workshops & PhD Forum (IPDPSW), 2012 IEEE 26th International, May 2012, pp. 1-9. [cited by applicant]
Office Action received for Japanese Patent Application No. 2024-563945, mailed on Oct. 7, 2025, 12 pages (6 pages of English Translation and 6 pages of Original Document). [cited by applicant]