IP Library › Granted Patent US 12,182,062
Granted Patent B1
US 12,182,062 · App. 17/961,833 · Granted Dec 31, 2024

Multi-tile memory management

Inventors: Abhishek R. Appu (El Dorado Hills, CA); Altug Koker (El Dorado Hills, CA); Aravindh Anantaraman (Folsom, CA); Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Valentin Andrei (San Jose, CA); Nicolas Galoppo Von Borries (Portland, OR); Varghese George (Folsom, CA); Mike Macpherson (Portland, OR); Subramaniam Maiyuran (Gold River, CA); Joydeep Ray (Folsom, CA); Lakshminarayanan Striramassarma (Folsom, CA); Scott Janus (Loomis, CA); Brent Insko (Portland, OR); Vasanth Ranganathan (El Dorado Hills, CA); Kamal Sinha (Rancho Cordova, CA); Arthur Hunter (Cameron Park, CA); Prasoonkumar Surti (Folsom, CA); David Puffer (Tempe, AZ); James Valerio (North Plains, OR); Ankur N. Shah (Folsom, CA)
Assignee: Intel Corporation
G06F15/7839G06F7/5443G06F7/575G06F7/588G06F9/3001G06F9/30014G06F9/30036G06F9/3004G06F9/30043G06F9/30047G06F9/30065G06F9/30079G06F9/3888G06F9/5011G06F9/5077G06F12/0215G06F12/0238G06F12/0246G06F12/0607G06F12/0802G06F12/0804G06F12/0811G06F12/0862G06F12/0866G06F12/0871G06F12/0875G06F12/0882G06F12/0891G06F12/0893G06F12/0895G06F12/0897G06F12/1009G06F12/128G06F15/8046G06F17/16G06F17/18G06T1/20G06T1/60H03M7/46G06F9/3802G06F9/3818G06F9/3867G06F2212/1021G06F2212/1044G06F2212/302G06F2212/401G06F2212/455G06F2212/60G06N3/08G06T15/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,182,062
App. No.
17/961,833
Granted
Dec 31, 2024
Kind
B1
Abstract

Methods and apparatus relating to techniques for multi-tile memory management. In an example, a graphics processor includes an interposer, a first chiplet coupled with the interposer, the first chiplet including a graphics processing resource and an interconnect network coupled with the graphics processing resource, cache circuitry coupled with the graphics processing resource via the interconnect network, and a second chiplet coupled with the first chiplet via the interposer, the second chiplet including a memory-side cache and a memory controller coupled with the memory-side cache. The memory controller is configured to enable access to a high-bandwidth memory (HBM) device, the memory-side cache is configured to cache data associated with a memory access performed via the memory controller, and the cache circuitry is logically positioned between the graphics processing resource and a chiplet interface.

Claims (33)

1. A graphics processor comprising:

an interposer;

a first chiplet coupled with the interposer, the first chiplet including a graphics processing resource and a switched interconnect network coupled with the graphics processing resource;

cache circuitry coupled with the graphics processing resource via the switched interconnect network; and

a second chiplet coupled with the first chiplet via the interposer, the second chiplet including a memory-side cache and a memory controller coupled with the memory-side cache,

wherein the memory controller is configured to enable access to a high-bandwidth memory (HBM) device,

the memory-side cache is configured to cache data associated with a memory access performed via the memory controller, and

the cache circuitry is logically positioned between the graphics processing resource and a chiplet interface, wherein the first chiplet and the second chiplet are distinct and at least partially packaged integrated circuits.

2. The graphics processor of claim 1 , wherein the cache circuitry includes a level-2 (L2) cache.

3. The graphics processor of claim 2 , wherein the L2 cache is a distributed cache having nodes interconnected via the switched interconnect network.

4. The graphics processor of claim 1 , wherein the memory-side cache includes a level-3 (L3) cache.

5. The graphics processor of claim 1 , wherein the interposer, first chiplet, and second chiplet have a 2.5 dimension (2.5D) arrangement and the interposer includes a first standardized slot configured to accept the first chiplet and a second standardized slot configured to accept the second chiplet.

6. The graphics processor of claim 5 , wherein the interposer is an active interposer.

7. The graphics processor of claim 1 , wherein the graphics processing resource includes a plurality of functional units having a single instruction multiple thread (SIMT) architecture.

8. The graphics processor of claim 7 , wherein the plurality of functional units includes a general-purpose graphics processor core and a tensor core.

9. The graphics processor of claim 7 , wherein the plurality of functional units includes a ray-tracing core.

10. The graphics processor of claim 1 , wherein the graphics processing resource includes a plurality of functional units having a single instruction multiple data (SIMD) architecture.

11. A system comprising:

an interconnect to a system interface; and

a multi-die graphics processor coupled with the interconnect, the multi-die graphics processor comprising:

an interposer;

a first chiplet coupled with the interposer, the first chiplet including a graphics processing resource and a switched interconnect network coupled with the graphics processing resource;

cache circuitry coupled with the graphics processing resource via the switched interconnect network; and

a second chiplet coupled with the first chiplet via the interposer, the second chiplet including a memory-side cache and a memory controller coupled with the memory-side cache, wherein the memory controller is configured to enable access to a high-bandwidth memory (HBM) device, the memory-side cache is configured to cache data associated with a memory access performed via the memory controller, and the cache circuitry is logically positioned between the graphics processing resource and a chiplet interface, wherein the first chiplet and the second chiplet are distinct and at least partially packaged integrated circuits.

12. The system of claim 11 , wherein the cache circuitry includes a level-2 (L2) cache.

13. The system of claim 12 , wherein the L2 cache is a distributed cache having nodes interconnected via the switched interconnect network.

14. The system of claim 11 , wherein the memory-side cache includes a level-3 (L3) cache.

15. The system of claim 11 , wherein the interposer, first chiplet, and second chiplet have a 2.5 dimension (2.5D) arrangement and the interposer includes a first standardized slot configured to accept the first chiplet and a second standardized slot configured to accept the second chiplet.

16. The system of claim 15 , wherein the interposer is an active interposer.

17. The system of claim 11 , wherein the graphics processing resource includes a plurality of functional units having a single instruction multiple thread (SIMT) architecture.

18. The system of claim 17 , wherein the plurality of functional units includes a general-purpose system core and a tensor core.

19. The system of claim 17 , wherein the plurality of functional units includes a ray-tracing core.

20. The system of claim 11 , wherein the graphics processing resource includes a plurality of functional units having a single instruction multiple data (SIMD) architecture.

Continuity (4)
Continuation 17431034
Provisional Application 62819337 · Mar 15, 2019
Provisional Application 62819435 · Mar 15, 2019
Provisional Application 62819361 · Mar 15, 2019
Cited By (3)
US 12,554,674 US 12,561,277 US 12,737,317