IP Library › Granted Patent US 12,572,392
Granted Patent B2
US 12,572,392 · App. 17/827,373 · Granted Mar 10, 2026

Flexible partitioning of GPU resources

Inventors: David Cowperthwaite (Portland, OR); Kenneth Daxer (Sunnyvale, CA); Jeffery S. Boles (Folsom, CA); Hema Chand Nalluri (Bengaluru, IN); Aditya Navale (Folsom, CA); Prasoonkumar Surti (Folsom, CA); Arthur Hunter (Cameron Park, CA); Vasanth Ranganathan (El Dorado Hills, CA); Joydeep Ray (Folsom, CA); David Puffer (Tempe, AZ); Aravindh Anantaraman (Folsom, CA); Ankur Shah (Folsom, CA); Vidhya Krishnan (Folsom, CA); Kritika Bala (Folsom, CA)
Assignee: Intel Corporation
G06F9/5077G06F9/5016G06T1/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,392
App. No.
17/827,373
Granted
Mar 10, 2026
Kind
B2
Abstract

Described herein is a partitionable graphics processor having a plurality of flexibly partitioned processing resources. One embodiment provides a graphics processor comprising a plurality of processing resources configurable to be flexibly partitioned into a plurality of resource partitions and circuitry to compose multiple graphics processor device partitions from the plurality of resource partitions. The multiple graphics processor device partitions are configurable to be asymmetrically composed of different types of functional units.

Claims (29)

1 . A graphics processor comprising:

a plurality of processing resources configurable to be flexibly partitioned into a plurality of resource partitions, the plurality of processing resources respectively including a vector engine and a matrix engine; and

circuitry to compose multiple graphics processor device partitions from the plurality of resource partitions, wherein the multiple graphics processor device partitions are configurable to be asymmetrically composed of different types of functional units, the circuitry is configurable to compose a first partition and a second partition, the first partition includes a first processing resource, and the second partition includes a second processing resource.

2 . The graphics processor as in claim 1 , wherein the circuitry includes a network on chip interconnect that is configurable to compose a virtual graphics core including selected processing resources.

3 . The graphics processor as in claim 2 , wherein the circuitry is configured to compose a first virtual graphics core including the first processing resource for the first partition and a second virtual graphics core including the second processing resource for the second partition.

4 . The graphics processor as in claim 3 , wherein the circuitry is configured to map the matrix engine of the second processing resource to the first partition and map the vector engine of the second processing resource to the second partition.

5 . The graphics processor as in claim 4 , wherein the first partition is configured to execute a matrix instruction via the matrix engine of the first processing resource and the matrix engine of the second processing resource.

6 . The graphics processor as in claim 4 , wherein the first partition is configured to execute a first matrix instruction via the matrix engine of the first processing resource and a second matrix instruction via the matrix engine of the second processing resource.

7 . The graphics processor as in claim 4 , wherein the first processing resource and the second processing resource each include a media engine and a ray tracing unit, the circuitry is configured to map the media engine of the first processing resource and the media engine of the second processing resource to the second partition and map the ray tracing unit of the first processing resource and the second processing resource to the first partition.

8 . The graphics processor as in claim 4 , further comprising a system interface to present the graphics processor to a host system.

9 . The graphics processor as in claim 8 , wherein the system interface is configured to present the multiple graphics processor device partitions as separate sub-devices of the graphics processor.

10 . The graphics processor as in claim 9 , the separate sub-devices to be presented as single-root input/output virtualization virtual functions or scalable input/output virtualization virtual devices.

11 . A data processing system comprising:

a system interface; and

a graphics processor coupled with the system interface, the graphics processor including:

a plurality of processing resources configurable to be flexibly partitioned into a plurality of resource partitions, the plurality of processing resources respectively including a vector engine and a matrix engine; and

circuitry to compose multiple graphics processor device partitions from the plurality of resource partitions, wherein the multiple graphics processor device partitions are configurable to be asymmetrically composed of different types of functional units and presented via the system interface as multiple sub-devices of the graphics processor, the circuitry is configurable to compose a first partition and a second partition, the first partition includes a first processing resource, and the second partition includes a second processing resource.

12 . The data processing system as in claim 11 , wherein the circuitry includes a network on chip interconnect that is configurable to compose a virtual graphics core including selected processing resources.

13 . The data processing system as in claim 12 , wherein the circuitry is configured to compose a first virtual graphics core including the first processing resource for the first partition and a second virtual graphics core including the second processing resource for the second partition.

14 . The data processing system as in claim 13 , wherein the circuitry is configured to map the matrix engine of the second processing resource to the first partition and map the vector engine of the second processing resource to the second partition.

15 . The data processing system as in claim 14 , wherein the first partition is configured to execute a matrix instruction via the matrix engine of the first processing resource and the matrix engine of the second processing resource.

16 . The data processing system as in claim 14 , wherein the first partition is configured to execute a first matrix instruction via the matrix engine of the first processing resource and a second matrix instruction via the matrix engine of the second processing resource.

17 . A method comprising:

configuring a number of compute partitions for a graphics processor, the graphics processor including a processing resource having a plurality of vector engines and a plurality of matrix engines, wherein the processing resource is asymmetrically partitionable into a plurality of compute partitions;

composing a plurality of device partitions of the graphics processor via selection of one or more of the plurality of compute partitions for each device partition, the plurality of device partitions including a first device partition and a second device partition, wherein the first device partition includes the plurality of matrix engines and a first portion of the plurality of vector engines, and the second device partition includes a second portion of the plurality of vector engines; and

presenting the first device partition as a first sub-device and the second device partition as a second sub-device.

18 . The method as in claim 17 , wherein the graphics processor is included in a multi-client server device and the method further comprises presenting the first sub-device to a first client of the server device and presenting the second sub-device to a second client of the server device.

19 . The method as in claim 17 , further comprising configuring a number of cache and memory partitions for the graphics processor, wherein composing the plurality of device partitions of the graphics processor includes selecting one or more of a plurality of cache and memory partitions for each device partition.

20 . The method as in claim 17 , further comprising configuring a memory bandwidth allocation for each device partition, wherein configuring a memory bandwidth allocation for each device partition includes automatically determining an increased memory bandwidth allocation for the first device partition based on inclusion of the plurality of matrix engines in the first device partition.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2022
From: COWPERTHWAITE, DAVID; DAXER, KENNETH; BOLES, JEFFERY S.; NALLURI, HEMA CHAND; NAVALE, ADITYA; SURTI, PRASOONKUMAR; HUNTER, ARTHUR; RANGANATHAN, VASANTH; RAY, JOYDEEP; PUFFER, DAVID; ANANTARAMAN, ARAVINDH; SHAH, ANKUR; KRISHNAN, VIDHYA; BALA, KRITIKA
To: INTEL CORPORATION
Reel/Frame 062129/0757 →
Continuity (3)
Provisional Application 63321665 · Mar 19, 2022
Provisional Application 63321604 · Mar 18, 2022
Related Publication 20230297440A1 · Sep 21, 2023
References Cited (55)
US 5764999A · Wilcox · 1998 [cited by applicant]
US 7583268B2 · Huang · 2009 [cited by applicant]
US 7701461B2 · Fouladi · 2010 [cited by applicant]
US 9158569B2 · Mitra et al. · 2015 [cited by applicant]
US 10176550B1 · Baggerman · 2019 [cited by examiner]
US 10754649B2 · Bainville · 2020 [cited by examiner]
US 10891773B2 · Ray et al. · 2021 [cited by applicant]
US 11599490B1 · Machulsky · 2023 [cited by applicant]
US 20030160818A1 · Tschiegg et al. · 2003 [cited by applicant]
US 20090055157A1 · Soffer · 2009 [cited by applicant]
US 20100066762A1 · Yeh et al. · 2010 [cited by applicant]
US 20120081355A1 · Post et al. · 2012 [cited by applicant]
US 20130080567A1 · Pope · 2013 [cited by applicant]
US 20140132611A1 · Chen et al. · 2014 [cited by applicant]
US 20150294494A1 · Stone · 2015 [cited by applicant]
US 20160328823A1 · Rao et al. · 2016 [cited by applicant]
US 20170329729A1 · Chew · 2017 [cited by applicant]
US 20180293692A1 · Koker et al. · 2018 [cited by applicant]
US 20180307533A1 · Tian et al. · 2018 [cited by applicant]
US 20180308198A1 · Appu et al. · 2018 [cited by applicant]
US 20180341503A1 · Nair · 2018 [cited by applicant]
US 20190270005A1 · Gary · 2019 [cited by applicant]
US 20190332425A1 · Narayana et al. · 2019 [cited by applicant]
US 20200201758A1 · Asaro et al. · 2020 [cited by applicant]
US 20200210359A1 · Cornett et al. · 2020 [cited by applicant]
US 20200219223A1 · Vembu et al. · 2020 [cited by applicant]
US 20200264910A1 · Kraemer et al. · 2020 [cited by applicant]
US 20200264994A1 · Raisch et al. · 2020 [cited by applicant]
US 20200278938A1 · Vembu et al. · 2020 [cited by applicant]
US 20200379920A1 · Banerjee et al. · 2020 [cited by applicant]
US 20200409733A1 · Sankaran et al. · 2020 [cited by applicant]
US 20200410628A1 · Shah et al. · 2020 [cited by applicant]
US 20210073125A1 · Duluk, Jr. et al. · 2021 [cited by applicant]
US 20210165745A1 · Karve et al. · 2021 [cited by applicant]
US 20210263755A1 · Tian et al. · 2021 [cited by applicant]
US 20210272347A1 · Mccrary · 2021 [cited by applicant]
US 20220058047A1 · Epstein · 2022 [cited by applicant]
US 20220206833A1 · Bhandari · 2022 [cited by examiner]
US 20220276966A1 · Shcherbina et al. · 2022 [cited by applicant]
US 20220405128A1 · Aristarkhov · 2022 [cited by applicant]
US 20230288471A1 · Duluk et al. · 2023 [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/459,311, mailed Sep. 17, 24, 8 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/832,305, mailed Apr. 26, 23, 8 pages. [cited by applicant]
Intel® Architecture Instruction Set Extensions and Future Features, Programming Reference, May 2021, Ref. # 319433-044, 214 pages. [cited by applicant]
Intel® Scalable I/O Virtualization, Technical Specification, Ref. # 337679-002, Rev. 1.1, Sep. 2020, 29 pages. [cited by applicant]
Intel®, White Paper: “PCI-SIG Single Root I/O Virtualization (SR-IOV) Support in Intel Virtualization Technology for Connetivity”, 2008, Rev. 06/08-001US, 4 pages. [cited by applicant]
Nvidia, “Multi-Process Service”, vR495, Oct. 2021, 34 pages. [cited by applicant]
Nvidia, “Nvidia Multi-Instance GPU User Guide”, RN-08625-v.1.0_v01, Aug. 2021, 46 pages. [cited by applicant]
Office Action for U.S. Appl. No. 17/827,305 mailed May 8, 2025, 18 pages. [cited by applicant]
Office action for U.S. Appl. No. 17/827,346 mailed May 19, 2025, 17 pages. [cited by applicant]
Office Action for U.S. Appl. No. 17/827,444 dated Apr. 28, 2025, 30 pages. [cited by applicant]
Office Action for U.S. Appl. No. 17/849,106 mailed Jun. 3, 2025, 27 pages. [cited by applicant]
Office Action for U.S. Appl. No. 17/849,165 mailed Apr. 23, 2025, 8 pages. [cited by applicant]
Final Office Action for U.S. Appl. No. 17/827,346 mailed Aug. 26, 2025, 18 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/827,444 mailed Sep. 2, 2025, 16 pages. [cited by applicant]