IP Library Granted Patent US 12,299,413
Granted Patent B2
US 12,299,413 · App. 18/414,164 · Granted May 13, 2025

Dual vector arithmetic logic unit

Inventors: Bin He (Orlando, FL); Brian Emberling (Santa Clara, CA); Mark Leather (Santa Clara, CA); Michael Mantor (Orlando, FL)
Assignee: ADVANCED MICRO DEVICES, INC.
G06F7/57G06F9/3851G06F9/3867G06F9/3887G06F17/16G06T1/20G06F15/8015
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,413
App. No.
18/414,164
Granted
May 13, 2025
Kind
B2
Abstract

A processing system executes wavefronts at multiple arithmetic logic unit (ALU) pipelines of a single instruction multiple data (SIMD) unit in a single execution cycle. The ALU pipelines each include a number of ALUs that execute instructions on wavefront operands that are collected from vector general process register (VGPR) banks at a cache and output results of the instructions executed on the wavefronts at a buffer. By storing wavefronts supplied by the VGPR banks at the cache, a greater number of wavefronts can be made available to the SIMD unit without increasing the VGPR bandwidth, enabling multiple ALU pipelines to execute instructions during a single execution cycle.

Claims (39)

1. A system comprising:

a cache to store a set of operands transferred from a set of vector general purpose register (VGPR) banks; and

an execution unit comprising a first arithmetic logic unit (ALU) pipeline and a second ALU pipeline to selectively execute either a single instruction at each of the first ALU pipeline and the second ALU pipeline using the first set of operands in an execution cycle or a pair of independent instructions at respective ALU pipelines of the first ALU pipeline and the second ALU pipeline using the first set of operands in the execution cycle.

2. The system of claim 1 , further comprising:

the set of VGPR banks to transfer operands to the cache.

3. The system of claim 2 , wherein the set of VGPR banks receives the operands from local data share return data, texture return data, VGPR initialization inputs, or any combination thereof.

4. The system of claim 1 , further comprising:

a buffer to store results of the single instruction or the pair of independent instructions; and

a controller to cause some or all of the results to be sent to the cache, the set of VGPR banks, or both.

5. The system of claim 1 , wherein the cache is configured to store received results as operands for a future instruction.

6. The system of claim 1 , wherein a number of operands in the set of operands equals a number of ALUs in the first ALU pipeline plus a number of ALUs in the second ALU pipeline.

7. The system of claim 6 , wherein the set of operands further includes an indication that the number of operands in the set of operands equals the number of ALUs in the first ALU pipeline plus the number of ALUs in the second ALU pipeline.

8. The system of claim 1 , wherein the set of operands further includes indications of ALUs that are to receive respective operands of the set of operands.

9. An apparatus, comprising:

a set of vector general purpose register (VGPR) banks to store a set of vectors;

a cache to store the set of vectors transferred from the set of VGPR banks; and

an execution unit comprising a first arithmetic logic unit (ALU) pipeline and a second ALU pipeline to selectively execute either a single instruction or a dual instruction on the set of vectors in an execution cycle both at the first ALU pipeline and at the second ALU pipeline.

10. The apparatus of claim 9 , further comprising:

an instruction buffer to send, to the execution unit, a vector instruction to perform operations on vectors from the cache, wherein the vector instruction is either the single instruction or the dual instruction.

11. The apparatus of claim 9 , further comprising:

a buffer to store results of the single instruction or the dual instruction; and

a controller to cause the results of the single instruction or the dual instruction to be sent to the cache, the set of VGPR banks, or both.

12. The apparatus of claim 9 , wherein a number of work items in the set of vectors equals a number of ALUs in the first ALU pipeline.

13. The apparatus of claim 12 , wherein a number of ALUs in the second ALU pipeline equals the number of ALUs in the first ALU pipeline.

14. The apparatus of claim 12 , wherein the set of vectors further includes an indication that the number of vectors in the set of operands equals the number of ALUs in the first ALU pipeline.

15. The apparatus of claim 9 , wherein the set of vectors further includes indications of ALUs that are to receive respective vectors of the set of vectors.

16. A method, comprising:

transferring a set of warps from a set of vector general purpose register (VGPR) banks to an execution unit via a cache; and

selectively executing, at the execution unit, either a single instruction or a dual instruction on the set of warps in an execution cycle both at a first ALU pipeline and at a second ALU pipeline.

17. The method of claim 16 , wherein:

a number of work items in the set of warps equals a number of ALUs of the first ALU pipeline plus a number of ALUs of the second ALU pipeline; and

selectively executing comprises executing the single instruction both at the first ALU pipeline and at the second ALU pipeline in the execution cycle.

18. The method of claim 16 , wherein:

a number of work items in the set of warps equals a number of ALUs of the first ALU pipeline; and

selectively executing comprises executing the dual instruction at respective ALU pipelines of the first ALU pipeline and the second ALU pipeline.

19. The method of claim 16 , further comprising:

storing results of the single instruction or the dual instruction at a buffer.

20. The method of claim 19 , further comprising:

transferring the results from the buffer to the cache or the set of VGPR banks in response to a determination that the results are to be used in execution of a future instruction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2024
From: HE, BIN; EMBERLING, BRIAN; LEATHER, MARK; MANTOR, MICHAEL
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 067980/0772 →
Continuity (3)
Continuation 18201839 · May 25, 2023
Continuation 17121354 · Dec 14, 2020
Related Publication 20240168719A1 · May 23, 2024
References Cited (9)
US 10817302B2 · Chen et al. · 2020 [cited by applicant]
US 20180121386A1 · Chen · 2018 [cited by examiner]
US 20180239606A1 · Mantor · 2018 [cited by examiner]
US 20190278605A1 · Emberling · 2019 [cited by examiner]
Ubal, R. et al., (ACM paper entitiled Multi2Sim: A Simulation Framework for CPU-GPU Computing) 2012. ACM, 10 pages. (Year: 2012). [cited by examiner]
Gong, Xun et. al., HAWS: Accelerating GPU Wavefront Execution through Selective Out-of-order Execution, 2019, ACM, pp. 15:1-15:22 (Year: 2019). [cited by examiner]
Fung, Wilson Wai Lunl., Dynamic Warp Formation: Exploiting Thread Scheduling for Efficient MIMD Control Flow on SIMD Graphics Hardware,2006, Univ. of British Columbia (Vancouver), 99 pages. (Year: 2006). [cited by examiner]
Teodoro, George et al., Efficient irregular wavefront propagation algorithms on hybrid CPU-GPU machines. 2013, Elsevier, Parallel Computing 39 (2013), pp. 189-211). (Year: 2013). [cited by examiner]
Extended European Search Report issued in Application No. 21907572.8, mailed Oct. 25, 2024, 8 pages. [cited by applicant]