IP Library Granted Patent US 12,327,124
Granted Patent B2
US 12,327,124 · App. 18/128,642 · Granted Jun 10, 2025

Vertical and horizontal broadcast of shared operands

Inventors: Sateesh Lagudu (Hyderabad, IN); Allen H. Rush (Santa Clara, CA); Michael Mantor (Orlando, FL); Arun Vaidyanathan Ananthanarayan (Hyderabad, IN); Prasad Nagabhushanamgari (Hyderabad, IN); Maxim V. Kazakov (San Diego, CA)
Assignee: Advanced Micro Devices, Inc.
G06F9/3887G06F9/3888G06F13/28G06F13/4027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,327,124
App. No.
18/128,642
Granted
Jun 10, 2025
Kind
B2
Abstract

An array processor includes processor element arrays distributed in rows and columns. The processor element arrays perform operations on parameter values. The array processor also includes memory interfaces that broadcast sets of the parameter values to mutually exclusive subsets of the rows and columns of the processor element arrays. In some cases, the array processor includes single-instruction-multiple-data (SIMD) units including subsets of the processor element arrays in corresponding rows, workgroup processors (WGPs) including subsets of the SIMD units, and a memory fabric configured to interconnect with an external memory that stores the parameter values. The memory interfaces broadcast the parameter values to the SIMD units that include the processor element arrays in rows associated with the memory interfaces and columns of processor element arrays that are implemented across the SIMD units in the WGPs. The memory interfaces access the parameter values from the external memory via the memory fabric.

Claims (26)

1. A device comprising:

a plurality of workgroup processors, each workgroup processor comprising a plurality of single-instruction-multiple-data (SIMD) units, each SIMD unit comprising a set of processor element arrays distributed in rows and columns and configured to perform operations on parameter values; and

memory interfaces configured to broadcast sets of parameter values to the processor element arrays, wherein each of the memory interfaces is to exclusively broadcast parameter values to one or more of one row and one column of the processor element arrays.

2. The device of claim 1 , wherein the processor element arrays comprise vector arithmetic logic unit (ALU) processors, and wherein the memory interfaces comprise direct memory access (DMA) engines.

3. The device of claim 1 , wherein a first memory interface of the memory interfaces broadcasts first parameter values to the processor element arrays and wherein a second memory interface of the memory interfaces broadcasts second parameter values to the processor element arrays.

4. The device of claim 1 , wherein the memory interfaces are connected to the processor element arrays via separate physical connections.

5. The device of claim 1 , wherein the memory interfaces are configured to concurrently populate registers associated with the processor element arrays with the parameter values.

6. The device of claim 1 , further comprising:

a memory fabric configured to interconnect with an external memory that stores the parameter values, and

wherein the memory interfaces are configured to access the parameter values from the external memory via the memory fabric.

7. A method comprising:

broadcasting, from memory interfaces, parameter values to processor element arrays distributed in rows and columns,

wherein subsets of the processor element arrays are implemented in corresponding single-instruction-multiple-data (SIMD) units, and

wherein each of the memory interfaces exclusively broadcasts parameter values to one or more of one row and one column of the processor element arrays.

8. The method of claim 7 , wherein the processor element arrays comprise vector arithmetic logic unit (ALU) processors, and wherein the memory interfaces comprise direct memory access (DMA) engines.

9. The method of claim 7 , wherein broadcasting the parameter values comprises broadcasting first parameter values from a first memory interface of the memory interfaces to the processor element arrays, and wherein broadcasting the parameter values comprises broadcasting second parameter values from a second memory interface of the memory interfaces to the processor element arrays.

10. The method of claim 7 , wherein broadcasting the parameter values comprises broadcasting the parameter values via separate physical connections between the memory interfaces.

11. The method of claim 7 , wherein broadcasting the parameter values comprises concurrently populating registers associated with the processor element arrays with the parameter values.

12. The method of claim 7 , wherein fetching the parameter values comprises accessing the parameter values via a memory fabric configured to interconnect with the memory that stores the parameter values.

13. A non-transitory computer readable medium embodying a set of executable instructions, the set of executable instructions to manipulate a computer system to perform a portion of a process to fabricate at least part of a processor, the processor comprising:

a plurality of workgroup processors, each workgroup processor comprising a plurality of single-instruction-multiple-data (SIMD) units, each SIMD unit comprising a set of processor element arrays distributed in rows and columns and configured to perform operations on parameter values; and

memory interfaces configured to broadcast sets of parameter values to the processor element arrays, wherein each of the memory interfaces is to exclusively broadcast parameter values to one or more of one row and one column of the processor element arrays.

14. The non-transitory computer readable medium of claim 13 , wherein the processor element arrays comprise vector arithmetic logic unit (ALU) processors, and wherein the memory interfaces comprise direct memory access (DMA) engines.

15. The non-transitory computer readable medium of claim 13 , wherein a first memory interface of the memory interfaces broadcasts first parameter values to the processor element arrays, and wherein a second memory interface of the memory interfaces broadcasts second parameter values to the processor element arrays.

16. The non-transitory computer readable medium of claim 13 , wherein the memory interfaces are connected to the processor element arrays via separate physical connections.

17. The non-transitory computer readable medium of claim 13 , wherein the memory interfaces are configured to concurrently populate registers associated with the processor element arrays with the parameter values.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2023
From: KAZAKOV, MAXIM V.
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 063175/0009 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2023
From: LAGUDU, SATEESH; NAGABHUSHANAMGARI, PRASAD; ANANTHANARAYAN, ARUN VAIDYANATHAN; MANTOR, MICHAEL; RUSH, ALLEN
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 063175/0192 →
Continuity (2)
Continuation 17032307 · Sep 25, 2020
Related Publication 20230289191A1 · Sep 14, 2023
References Cited (19)
US 5710935A · Barker · 1998 [cited by examiner]
US 12164923B2 · Binder · 2024 [cited by examiner]
US 20020198911A1 · Blomgren · 2002 [cited by examiner]
US 20090164752A1 · McConnell · 2009 [cited by examiner]
US 20130318511A1 · Tian · 2013 [cited by examiner]
US 20140189321A1 · Uliel · 2014 [cited by examiner]
US 20150310311A1 · Shi · 2015 [cited by examiner]
US 20160188336A1 · Valentine · 2016 [cited by examiner]
US 20190196814A1 · Han · 2019 [cited by examiner]
US 20220067483A1 · Yudanov · 2022 [cited by examiner]
US 20220100528A1 · Lagudu · 2022 [cited by examiner]
US 20220197973A1 · Lagudu · 2022 [cited by examiner]
US 20240220273A1 · Vivekraja · 2024 [cited by examiner]
US 20240427596A1 · Hildebrand · 2024 [cited by examiner]
US 20250061074A1 · Park · 2025 [cited by examiner]
International Preliminary Report on Patentability issued in Application No. PCT/US2021/051892, mailed Apr. 6, 2023, 7 pages. [cited by applicant]
Extended European Search Report issued in Application No. 21873489.5, mailed Oct. 15, 2024, 10 pages. [cited by applicant]
Anonymous, “Viseral ACAP AI Engine”, 2020, <https://0x04.net/˜mwk/xidocs/am/v-ai.pdf>, Accessed Sep. 26, 2024, 62 pages. [cited by applicant]
Anonymous, “RDNA 1.0 Instruction Set Architecture”, 2020, <https://web.archive.org/web/2020031195133if_/https://developer.amd.com/wp-content/resources/RDNA_Shader_ISA.pdf>, 308 pages. [cited by applicant]