IP Library Granted Patent US 12,462,324
Granted Patent B2
US 12,462,324 · App. 18/127,554 · Granted Nov 4, 2025

Multicore state caching in graphics processing

Inventor: Ian King (Hertfordshire, GB)
Assignee: Imagination Technologies Limited
G06T1/20G06F9/4881G06F15/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,324
App. No.
18/127,554
Granted
Nov 4, 2025
Kind
B2
Abstract

A set of image rendering tasks and state information are distributed in a graphics processing unit (GPU) having a plurality of cores. A first master unit in one of the cores receives the set of image rendering tasks and the state information, and stores the state information in a memory. The first master unit splits the set of image rendering tasks into a first subset of tasks and a second subset of tasks, wherein the first subset of tasks is assigned to the first core, and the second subset of tasks is assigned to the second core. At least a first portion of the state information is transmitted to the first core, and at least a second portion of the state information is transmitted to the second core. The first subset of tasks is transmitted to the first core, and the second subset of tasks is transmitted to the second core.

Claims (99)

1 . A graphics processing unit comprising a plurality of cores, wherein one of the plurality of cores comprises a first master unit configured to:

receive a set of image rendering tasks and state information, wherein the state information comprises elements of state information required for processing the image rendering tasks;

store the state information in a memory;

split the set of image rendering tasks into at least a first subset of tasks and a second subset of tasks;

assign the first subset of tasks to a first core of the plurality of cores;

assign the second subset of tasks to a second core of the plurality of cores;

transmit an indication of at least a first portion of the state information to the first core;

transmit an indication of at least a second portion of the state information to the second core;

transmit an indication of the first subset of tasks to the first core;

transmit an indication of the second subset of tasks to the second core; and

wherein each core of the plurality of cores comprises a slave unit configured to perform image rendering tasks.

2 . The graphics processing unit of claim 1 , wherein the first master unit is further configured to:

identify first elements of the state information that are required for processing the first subset of tasks; and

identify second elements of the state information that are required for processing the second subset of tasks;

wherein the first portion of the of the state information consists of the first elements of the state information, and the second portion of the state information consists of the second elements of the state information.

3 . The graphics processing unit of claim 2 , wherein the first master unit is configured to:

not transmit to the first core any elements of state information other than the first elements of state information; and

not transmit to the second core any elements of state information other than the second elements of state information.

4 . The graphics processing unit of claim 1 , wherein the first master unit is configured to maintain a record of the cores to which each element of state information has been transmitted, and to update the record each time an element of state information is transmitted to one of the cores.

5 . The graphics processing unit of claim 1 , wherein one of the plurality of cores comprises a second master unit configured to:

receive a second set of image rendering tasks and second state information, wherein the second state information comprises elements of second state information required for processing the second set of image rendering tasks;

store the second state information in a second memory;

split the second set of image rendering tasks into at least a fifth subset of tasks and a sixth subset of tasks;

assign the fifth subset of tasks to the first core;

transmit an indication of at least a portion of the second state information to the first core; and

transmit the fifth subset of tasks to the first core.

6 . The graphics processing unit of claim 1 , wherein:

the plurality of cores are connected by a register bus configured to communicate register write commands between the cores;

the first master unit is configured to output at least a first register write command and a second register write command;

the first register write command is addressed to the first core and comprises an indication of the elements of state information required to process the first subset of tasks; and

the second register write command is addressed to the second core and comprises an indication of the elements of state information required to process the second subset of tasks.

7 . The graphics processing unit of claim 1 , wherein the first core comprises a plurality of processing units configured to process image rendering tasks, and wherein the slave unit of the first core is configured to:

receive the first subset of the image rendering tasks and the first portion of the state information;

split the first subset of image rendering tasks into a seventh subset of tasks and an eighth subset of tasks;

send the seventh subset to a first processing unit;

send the eighth subset to a second processing unit;

forward the first portion of the state information to the first processing unit; and

forward the first portion of the state information to the second processing unit.

8 . A method of distributing a set of image rendering tasks and state information in a graphics processing unit comprising a plurality of cores, the method comprising:

receiving, by a first master unit in one of the plurality of cores, the set of image rendering tasks and the state information, wherein the state information comprises elements of state information required for processing the image rendering tasks;

storing, by the first master unit, the state information in a memory;

splitting, by the first master unit, the set of image rendering tasks into at least a first subset of tasks and a second subset of tasks;

assigning, by the master unit, the first subset of tasks to the first core;

assigning, by the master unit, the second subset of tasks to the second core;

transmitting, by the master unit to the first core, at least a first portion of the state information;

transmitting, by the master unit to the second core; at least a second portion of the state information;

transmitting, by the master unit to the first core, the first subset of tasks; and

transmitting, by the master unit to the second core, the second subset of tasks.

9 . The method of claim 8 , further comprising:

identifying, by the first master unit, first elements of state information that are required for processing the first subset of tasks; and

identifying, by the first master unit, second elements of state information that are required for processing the second subset of tasks;

wherein the first portion of the state information consists of the first elements of the state information, and the second portion of the state information consists of the second elements of the state information.

10 . The method of claim 9 , wherein the first master unit:

does not transmit to the first core any elements of state information other than the first elements of state information; and

does not transmit to the second core, any elements of state information other than the second elements of state information.

11 . The method of claim 8 , further comprising:

maintaining, by the first master unit, a record of the cores to which each element of state information has been transmitted; and

updating, by the first master unit, the record each time an element of state information is transmitted to one of the cores.

12 . The method of claim 11 , further comprising:

receiving, by the first master unit, additional state information, wherein the additional state information comprises one or more of:

(A) a new element of state information, wherein the method further comprises:

storing, by the master unit, the new element of state information, and

updating, by the master unit, the record to include the new element of state information and indicate that the new element has not been transmitted to any of the cores, or

(B) an updated element of state information, wherein the method further comprises:

replacing, by the master unit, an element of state information stored in the memory with the updated element of state information, and

updating, by the master unit, the record to indicate that the updated element has not been transmitted to any of the cores;

receiving, by the first master unit, an additional set of image rendering tasks associated with the additional state information;

splitting, by the first master unit, the additional set of image rendering tasks into at least a third subset of tasks and a fourth subset of tasks;

assigning, by the first master unit, the third subset of tasks to the first core;

identifying, based on the record, third elements of state information, wherein the third elements of state information are elements of state information that are required to process the third subset of tasks and have not been transmitted to the first core;

transmitting, by the first master unit and to the first core, an indication of the third elements of state information; and

transmitting, by the first master unit, an indication of the third subset of tasks to the first core.

13 . The method of claim 8 , wherein:

the state information includes a cumulative element of state information; and

the method further comprises transmitting, by the first master unit, the cumulative element of state information to every core in the graphics processing unit.

14 . The method of claim 13 , further comprising:

receiving, by the first master unit, a new cumulative element of state information; and

transmitting, by the first master unit, the new cumulative element of state information to each core in the plurality of cores.

15 . The method of claim 8 , further comprising:

receiving, by a slave unit of the first core, the first subset of tasks and the first portion of the state information;

splitting, by the slave unit, the first subset of tasks into a fifth subset of tasks and a sixth subset of tasks;

sending, by the slave unit; the fifth subset of tasks to a first processing unit of the first core;

sending, by the slave unit; the sixth subset of tasks to a second processing unit of the first core;

forwarding, by the slave unit to the first processing unit, the first portion of the state information; and

forwarding, by the slave unit to the second processing unit, the first portion of the state information.

16 . The method of claim 8 , further comprising:

receiving, by a second master unit in one of the plurality of cores, a second set of image rendering tasks and second state information, wherein the second state information comprises elements of second state information required for processing the second set of image rendering tasks;

storing, by the second master unit, the second state information in a second memory;

splitting, by the second master unit, the second set of image rendering tasks into at least a seventh subset of tasks and an eighth subset of tasks;

assigning, by the second master unit, the seventh subset of tasks to the first core;

transmitting, by the second master unit to the first core, at least a portion of the second state information; and

transmitting, by the second master unit, the seventh subset of tasks to the first core.

17 . A method of manufacturing a graphics processing unit as set forth in claim 1 , the method comprising inputting to an integrated circuit manufacturing system an integrated circuit definition dataset that, when processed in said integrated circuit manufacturing system, configures the integrated circuit manufacturing system to manufacture said graphics processing unit.

18 . A non-transitory computer readable storage medium having stored thereon computer readable code configured to cause the method as set forth in claim 8 to be performed when the code is run.

19 . A non-transitory computer readable storage medium having stored thereon a computer readable dataset description of a graphics processing unit as set forth in claim 1 that, when processed in an integrated circuit manufacturing system, causes the integrated circuit manufacturing system to manufacture an integrated circuit embodying the graphics processing unit.

20 . An integrated circuit manufacturing system comprising:

a non-transitory computer readable storage medium having stored thereon a computer readable dataset description of a graphics processing unit as set forth in claim 1 ;

a layout processing system configured to process the computer readable dataset description so as to generate a circuit layout description of an integrated circuit embodying the graphics processing unit; and

an integrated circuit generation system configured to manufacture the graphics processing unit according to the circuit layout description.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2025
From: KING, IAN
To: IMAGINATION TECHNOLOGIES LIMITED
Reel/Frame 072171/0912 →
SECURITY INTEREST Recorded Jul 31, 2024
From: IMAGINATION TECHNOLOGIES LIMITED
To: FORTRESS INVESTMENT GROUP (UK) LTD
Reel/Frame 068221/0001 →
Priority Claims (2)
GB 2204508 · Mar 30, 2022 · national
GB 2204510 · Mar 30, 2022 · national
Continuity (1)
Related Publication 20230377088A1 · Nov 23, 2023
References Cited (49)
US 8074224B1 · Nordquist et al. · 2011 [cited by applicant]
US 8330766B1 · McAllister et al. · 2012 [cited by applicant]
US 10275851B1 · Zhao et al. · 2019 [cited by applicant]
US 10733695B2 · Andersson et al. · 2020 [cited by applicant]
US 20070091099A1 · Zhang et al. · 2007 [cited by applicant]
US 20090307464A1 · Steinberg · 2009 [cited by examiner]
US 20110109638A1 · Duluk, Jr. et al. · 2011 [cited by applicant]
US 20150254102A1 · Ueda et al. · 2015 [cited by applicant]
US 20160260249A1 · Persson et al. · 2016 [cited by applicant]
US 20170178401A1 · Agrawal et al. · 2017 [cited by applicant]
US 20170236244A1 · Price · 2017 [cited by examiner]
US 20180130253A1 · Hazel · 2018 [cited by applicant]
US 20180211435A1 · Nijasure et al. · 2018 [cited by applicant]
US 20180276876A1 · Yang et al. · 2018 [cited by applicant]
US 20180307490A1 · Hakura et al. · 2018 [cited by applicant]
US 20190355084A1 · Gierach et al. · 2019 [cited by applicant]
US 20200097293A1 · Havlir et al. · 2020 [cited by applicant]
US 20210097013A1 · Saleh et al. · 2021 [cited by applicant]
US 20210158598A1 · Bratt et al. · 2021 [cited by applicant]
US 20210241416A1 · Cerny · 2021 [cited by applicant]
US 20220083384A1 · Cerny · 2022 [cited by applicant]
US 20220319089A1 · Nemlekar et al. · 2022 [cited by applicant]
US 20240005444A1 · Stepuch · 2024 [cited by applicant]
US 20240070962A1 · Yang et al. · 2024 [cited by applicant]
US 20240127524A1 · Yang et al. · 2024 [cited by applicant]
CN 105261066A · 2016 [cited by applicant]
CN 109978751A · 2019 [cited by applicant]
CN 112862661A · 2021 [cited by applicant]
EP 1287494A1 · 2003 [cited by applicant]
EP 2548171A1 · 2013 [cited by applicant]
EP 3547248A1 · 2019 [cited by applicant]
EP 3796263A1 · 2021 [cited by applicant]
EP 3862975A1 · 2021 [cited by applicant]
GB 2547252A · 2017 [cited by applicant]
GB 2594764A · 2021 [cited by applicant]
WO 2009068895A1 · 2009 [cited by applicant]
WO 2018114957A1 · 2018 [cited by applicant]
Anonymous; “Graphics—SGX543MP4”; Retrieved from the Internet: URL:https://www.psdevwiki.com/vita/Graphics; Sep. 13, 2020; pp. 1-9. [cited by applicant]
Beets; “A look at the PowerVR graphics architecture: Tilebased rendering”; Retrieved from the Internet: URL:https://blog.imaginationtech.com/a-look-at-the-powervr-graphics-architecture-tile-based-rendering/; Apr. 2, 201… [cited by applicant]
Beets; “A look at the PowerVR graphics architecture”; Retrieved from the Internet: URL:https://blog.imaginationtech.com/the-dr-in-tbdr-deferred-rendering-in-rogue/; Jan. 28, 2016; pp. 1-13. [cited by applicant]
Beets; “Introducing Furian: the architectural changes”; Retrieved from the Internet: URL:https://blog.imaginationtech.com/introducing-furian-the-architectural-changes/; Mar. 13, 2017; pp. 1-13. [cited by applicant]
Yu et al; “A Credit-Based Load-Balance-Aware CTA Scheduling Optimization Scheme in GPGPU”; International Journal of Parallel Programming; vol. 44; No. 1; Aug. 22, 2014; 21 pages. [cited by applicant]
Fedorov, D.G.: “A new hierarchical parallelization scheme: generalized distributed data interface (GDDI), and an application to the fragment molecular orbital method (FMO)”, Journal of computational chemistry. Apr. 30, … [cited by applicant]
Ullman, S.: “Object recognition and segmentation by a fragment-based hierarchy”, Trends in cognitive sciences. Feb. 1, 2007; 11 (2):58-64. [cited by applicant]
Crisu et al; “Low-Power Techniques and 2D/3D Graphics Architectures”; Report Delft University of Technology; vol. Jan. 2001; Jun. 26, 2001; 139 pages. [cited by applicant]
Ma; “Concepts and metrics for measurement and prediction of the execution time of GPU rendering commands”; Retrieved from the Internet: URL:https://elib.uni-stuttgart.de/bitstream/11682/3467/1/MSTR_3635.pdf; Aug. 19, 20… [cited by applicant]
Nickolls et al; “Appendix C: Graphics and Computing GPU'S”; Computer Organization and Design: The Hardware/Software Interface; URL:http://booksite.elsevier.com/9780124077263/downloads/advance_contents_and_appendices/app… [cited by applicant]
Imagination Technologies: “Tiling positive or how Vulkan maps to PowerVR GPUs”; Retrieved from the Internet: URL: https%3A%2F%2Fblog.imaginationtech.com%2Ftiling-positive-or-how-vulkan-maps-to-powervr-gpus%2F; Mar. 9, 2… [cited by applicant]
Kayhan; “Chasing Triangles in a Tile-based Rasterizer”; Retrieved from the Internet: URL:https://tayfunkayhan.wordpress.com/2019/07/26/chasing-triangles-in-a-tile-based⋅-rasterizer/; Jul. 29, 2019; pp. 1-18. [cited by applicant]