IP Library › Granted Patent US 12,462,186
Granted Patent B2
US 12,462,186 · App. 17/129,739 · Granted Nov 4, 2025

Stacked dies for machine learning accelerator

Inventors: Maxim V. Kazakov (San Diego, CA); Swapnil P. Sakharshete (San Diego, CA); Milind N. Nemlekar (San Diego, CA); Vineet Goel (San Diego, CA)
Assignee: Advanced Micro Devices, Inc.
G06N20/00G06F12/0893G06F13/1668G06F13/28G06F13/4027G06T15/005H01L25/0657
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,186
App. No.
17/129,739
Granted
Nov 4, 2025
Kind
B2
Abstract

A device is disclosed. The device includes a machine learning die including a memory and one or more machine learning accelerators; and a processing core die stacked with the machine learning die, the processing core die being configured to execute shader programs for controlling operations on the machine learning die, wherein the memory is configurable as either or both of a cache and a directly accessible memory.

Claims (38)

1 . A device, comprising:

a machine learning die including a memory and one or more machine learning accelerators; and

a processing core die stacked with the machine learning die, the processing core die being configured to execute shader programs for controlling operations on the machine learning die,

wherein each machine learning accelerator is coupled to a local memory portion, to each other machine learning accelerator via a memory interconnect, and to the processing core die via an inter-die interconnect,

wherein the memory is configurable as including both a cache and a directly accessible memory,

wherein the memory includes a portion configured to switch between being used as a cache and being used as a directly accessible memory, wherein the switching comprises altering how much of the memory is used as the cache based on a specified amount; and

wherein the cache is configured to perform caching operations for the processing core die and the directly accessible memory is configured to perform memory operations for the machine learning die.

2 . The device of claim 1 , wherein the machine learning accelerators are configured to perform matrix multiplication using data in the memory.

3 . The device of claim 1 , wherein the machine learning die and the processing core die are coupled via one or more inter-die interconnects.

4 . The device of claim 1 , wherein the processor core die is configured to modify a portion of the memory from being used for the one or more machine learning accelerators to being used as a cache for the processor core die.

5 . The device of claim 1 , wherein the processor core die is configured to modify a portion of the memory from being used as a cache for the APD core die to being used for the one or more machine learning accelerators.

6 . The device of claim 1 , wherein the processor core die is configured to execute shader instructions for storing data into the memory.

7 . The device of claim 1 , wherein the machine learning die further includes one or more controllers that control operations of the memory and the machine learning accelerators.

8 . The device of claim 7 , wherein the machine learning die further includes memory interconnects that couple the one or more controllers together.

9 . The device of claim 8 , wherein the memory interconnects are configured to provide data from one portion of the memory to a controller local to a different portion of the memory.

10 . A method, comprising:

executing a shader program on an accelerated processing device (“APD”) core die;

in accordance with instructions of the shader program, directing a set of machine learning (“ML”) arithmetic logic units (“ALUs”) of an ML accelerator die to perform a set of machine learning tasks, via one or more inter-die interconnects, wherein the ML accelerator die includes a memory, wherein the ML accelerator die is stacked with the APD core die, wherein the ML accelerator die includes one or more machine learning accelerators, and wherein each machine learning accelerator is coupled to a local memory portion, to each other machine learning accelerator via a memory interconnect, and to the APD core die via an inter-die interconnect,;

performing the set of machine learning tasks with ML ALUs; and

configuring the memory as both a cache and a directly accessible memory, wherein the cache is configured to perform caching operations for the APD core die and the directly accessible memory is configured to perform memory operations for the machine learning die, wherein the memory includes a portion configured to switch between being used as a cache and being used as a directly accessible memory, wherein the switching comprises altering how much of the memory is used as the cache based on a specified amount.

11 . The method of claim 10 , wherein the machine learning accelerators are configured to perform matrix multiplication using data in the memory.

12 . The method of claim 10 , wherein the machine learning die and the APD core die are coupled via one or more inter-die interconnects.

13 . The method of claim 10 , wherein the APD core die is configured to modify a portion of the memory from being used for the one or more machine learning accelerators to being used as a cache for the APD core die.

14 . The method of claim 10 , wherein the APD core die is configured to modify a portion of the memory from being used as a cache for the APD core die to being used for the one or more machine learning accelerators.

15 . The method of claim 10 , wherein the APD core die is configured to execute shader instructions for storing data into the memory.

16 . The method of claim 10 , wherein the machine learning die further includes one or more controllers that control operations of the memory and the machine learning accelerators.

17 . The method of claim 16 , wherein the machine learning die further includes memory interconnects that couple the one or more controllers together.

18 . The method of claim 17 , wherein the memory interconnects are configured to provide data from one portion of the memory to a controller local to a different portion of the memory.

19 . A device, comprising:

a processor; and

an accelerated processing device (“APD”) including:

a machine learning die including a memory and one or more machine learning accelerators; and

a processing core die stacked with the machine learning die, the processing core die being configured to execute shader programs for controlling operations on the machine learning die, one or more of the shader programs being specified by the processor,

wherein each machine learning accelerator is coupled to a local memory portion, to each other machine learning accelerator via a memory interconnect, and to the processing core die via an inter-die interconnect,

wherein the memory is configurable as including both a cache and a directly accessible memory,

wherein the memory includes a portion configured to switch between being used as a cache and being used as a directly accessible memory, wherein the switching comprises altering how much of the memory is used as the cache based on a specified amount; and

wherein the cache is configured to perform caching operations for the processing core die and the directly accessible memory is configured to perform memory operations for the machine learning die.

20 . The device of claim 19 , wherein the machine learning accelerators are configured to perform matrix multiplication using data in the memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2021
From: KAZAKOV, MAXIM V.; SAKHARSHETE, SWAPNIL P.; NEMLEKAR, MILIND N.; GOEL, VINEET
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 055039/0607 →
Continuity (2)
Provisional Application 63031954 · May 29, 2020
Related Publication 20210374607A1 · Dec 2, 2021
References Cited (21)
US 10268395B1 · Joshua · 2019 [cited by examiner]
US 20050102657A1 · Lewis · 2005 [cited by applicant]
US 20100228941A1 · Koob et al. · 2010 [cited by applicant]
US 20150379670A1 · Koker et al. · 2015 [cited by applicant]
US 20180157970A1 · Henry et al. · 2018 [cited by applicant]
US 20190042477A1 · Chhabra · 2019 [cited by examiner]
US 20190042923A1 · Janedula · 2019 [cited by examiner]
US 20190050040A1 · Baskaran et al. · 2019 [cited by applicant]
US 20190057300A1 · Mathuriya et al. · 2019 [cited by applicant]
US 20190073312A1 · Hu et al. · 2019 [cited by applicant]
US 20190156187A1 · Dasari et al. · 2019 [cited by applicant]
US 20190287208A1 · Yerli · 2019 [cited by applicant]
US 20200042477A1 · Malladi · 2020 [cited by applicant]
US 20200050476A1 · Xu et al. · 2020 [cited by applicant]
US 20200065113A1 · Gutierrez · 2020 [cited by applicant]
US 20200379911A1 · Wanner · 2020 [cited by examiner]
US 20210074059A1 · Ando · 2021 [cited by applicant]
CN 1391671A · 2003 [cited by applicant]
JP 2017517810A · 2017 [cited by applicant]
JP 2020027613A · 2020 [cited by applicant]
WO 2019225734A · 2021 [cited by applicant]