IP Library Granted Patent US 12,412,119
Granted Patent B2
US 12,412,119 · App. 17/135,383 · Granted Sep 9, 2025

Clustering of machine learning (ML) functional components

Inventors: Maxim V. Kazakov (San Diego, CA); Milind N. Nemlekar (San Diego, CA); Swapnil Sakharshete (San Diego, CA); Vineet Goel (San Diego, CA)
Assignee: ADVANCED MICRO DEVICES, INC.
G06N20/00G06F7/57G06F12/1081G06T1/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,119
App. No.
17/135,383
Granted
Sep 9, 2025
Kind
B2
Abstract

A graphics processing unit (GPU) for clustering of machine learning (ML) functional components, including: a plurality of compute units; a plurality of ML clusters, wherein each of the ML clusters comprises at least one arithmetic logic unit (ALU), and wherein each of the ML clusters is associated with a respective subset of the compute units; and a plurality of memory modules each positioned on the GPU adjacent to a respective ML cluster of the plurality of ML clusters, wherein each ML cluster is configured to directly access one or more adjacent memory modules.

Claims (32)

1. A graphics processing unit (GPU) for clustering of machine learning (ML) functional components, comprising:

a plurality of compute units;

a plurality of ML clusters, wherein each of the ML clusters comprises at least one arithmetic logic unit (ALU), and wherein each of the ML clusters is associated with a respective subset of the compute units; and

a plurality of memory modules each positioned on the GPU adjacent to a respective ML cluster of the plurality of ML clusters, wherein each ML cluster is configured to directly access one or more adjacent memory modules.

2. The GPU of claim 1 , wherein the plurality of ML clusters are associated with a first voltage domain distinct from at least one second voltage domain of the GPU.

3. The GPU of claim 1 , wherein a first portion of the memory modules comprise cache memory and a second portion of the memory modules comprise scratchpad memory.

4. The GPU of claim 1 , wherein the plurality of memory modules comprise static random access memory (SRAM) modules.

5. The GPU of claim 1 , wherein each of the ML clusters comprise at least one direct memory access (DMA) engine.

6. The GPU of claim 5 , wherein each of the ML clusters comprise a controller configured to issue commands to the at least one ALU and the at least one DMA engine.

7. The GPU of claim 1 , further comprising at least one control processor configured to issue commands to the plurality of ML clusters.

8. An apparatus for clustering of machine learning (ML) functional components, comprising:

a component;

a graphics processing unit (GPU) operatively coupled to the component, the GPU comprising:

a plurality of compute units;

a plurality of ML clusters, wherein each of the ML clusters comprises at least one arithmetic logic unit (ALU), and wherein each of the ML clusters is associated with a respective subset of the compute units; and

a plurality of memory modules each positioned on the GPU adjacent to a respective ML cluster of the plurality of ML clusters, wherein each ML cluster is configured to directly access one or more adjacent memory modules.

9. The apparatus of claim 8 , wherein the plurality of ML clusters are associated with a first voltage domain distinct from at least one second voltage domain of the GPU.

10. The apparatus of claim 8 , wherein a first portion of the memory modules comprise cache memory and a second portion of the memory modules comprise scratchpad memory.

11. The apparatus of claim 8 , wherein the plurality of memory modules comprise static random access memory (SRAM) modules.

12. The apparatus of claim 8 , wherein each of the ML clusters comprise at least one direct memory access (DMA) engine.

13. The apparatus of claim 12 , wherein each of the ML clusters comprise a controller configured to issue commands to the at least one ALU and the at least one DMA engine.

14. The apparatus of claim 8 , further comprising at least one control processor configured to issue commands to the plurality of ML clusters.

15. A graphics processing unit (GPU) for clustering of machine learning (ML) functional components, comprising:

a plurality of compute units;

a plurality of ML clusters, wherein each of the ML clusters is associated with a respective subset of the compute units; and

a plurality of memory modules each positioned on the GPU adjacent to a respective ML cluster of the plurality of ML clusters, wherein a ML cluster of the plurality of ML clusters is configured to directly access one or more adjacent memory modules and perform at least a portion of a general matrix multiply (GEMM) operation using the directly accessed one or more adjacent memory modules.

16. The GPU of claim 15 , wherein directly accessing the one or more adjacent memory modules comprises storing, by a DMA engine of the ML cluster, data into a scratchpad portion of the one or more adjacent memory modules; and

wherein performing the at least a portion of the GEMM operation comprises performing, by an arithmetic logic unit (ALU) of the ML cluster, the at least one operation on the data stored in the scratchpad portion of the one or more adjacent memory modules.

17. The GPU of claim 16 , wherein a controller of the ML cluster is to receive a first command, and wherein at least one second command is issued, based on the first command, to the ALU and the DMA engine of the ML cluster.

18. The GPU of claim 17 , wherein the first command is received from a control processor of the GPU.

19. The GPU of claim 17 , wherein the first command is received from a compute unit of the plurality of compute units of the GPU.

20. The GPU of claim 15 , wherein the plurality of ML clusters are associated with a first voltage domain separate from at least one second voltage domain of the GPU.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2021
From: KAZAKOV, MAXIM V.; NEMLEKAR, MILIND N.; SAKHARSHETE, SWAPNIL; GOEL, VINEET
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 054989/0100 →
Continuity (1)
Related Publication 20220207411A1 · Jun 30, 2022
References Cited (3)
US 20190370072A1 · Walter · 2019 [cited by examiner]
US 20210034136A1 · Hovis · 2021 [cited by examiner]
US 20210374607A1 · Kazakov · 2021 [cited by examiner]