IP Library Granted Patent US 12,554,497
Granted Patent B2
US 12,554,497 · App. 18/606,865 · Granted Feb 17, 2026

Fused comparison add instructions

Inventor: Steven Isaac Reeves (Truckee, CA)
Assignee: Advanced Micro Devices, Inc.
G06F9/30098G06F9/3001G06F9/30021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,554,497
App. No.
18/606,865
Granted
Feb 17, 2026
Kind
B2
Abstract

An apparatus, system, and method for efficiently processing pairs of operations repeatedly used in applications. In various implementations, a computing system includes a parallel data processing circuit with multiple compute circuits. Each of the compute circuits includes multiple lanes of execution, each with a corresponding arithmetic logic unit (ALU). The ALU supports executing a single fused conditional ternary instruction that replaces two separate instructions that provide two operations (comparison and add). When executing the fused conditional ternary instruction, the ALU does not retrieve the intermediate result from the scalar register file, the vector register file, or bypass circuitry located externally from the ALU. Rather, the ALU generates the intermediate result and uses the intermediate result without routing the intermediate result externally from ALU.

Claims (52)

1 . An integrated circuit comprising:

a register file configured to store data; and

circuitry configured to:

read, from the register file, three source operands indicated by an instruction;

responsive to the instruction being a fused instruction comprising two different operations with one being a comparison operation:

generate an indication that a first operation of the two different operations comprises a comparison operation;

generate a single result of the instruction as an output of a second operation of the two different operations; and

send the single result to the register file.

2 . The integrated circuit as recited in claim 1 , wherein the circuitry is further configured to send the indication to a plurality of arithmetic logic units (ALUs) in a same pipeline stage.

3 . The integrated circuit as recited in claim 2 , wherein each of the plurality of ALUs is configured to generate a corresponding single result of the instruction.

4 . The integrated circuit as recited in claim 3 , wherein the instruction is used as an activation function of a machine learning data model.

5 . The integrated circuit as recited in claim 1 , wherein the circuitry is further configured to generate an indication specifying the second operation of the two different operations comprises an addition operation.

6 . The integrated circuit as recited in claim 5 , wherein the circuitry is further configured to:

send two of the three source operands to the comparison operation; and

send one of the three source operands and a result of the comparison operation to the addition operation, bypassing intermediate pipeline registers and the register file.

7 . The integrated circuit as recited in claim 5 , wherein the circuitry is further configured to:

send two of the three source operands to the addition operation; and

send one of the three source operands and a result of the addition operation to the comparison operation, bypassing intermediate pipeline registers and the register file.

8 . A method comprising:

reading, by circuitry from a register file, three source operands indicated by an instruction;

responsive to the instruction being a fused instruction comprising two different operations with one being a comparison operation:

generating, by the circuitry, an indication that a first operation of the two different operations comprises a comparison operation;

generating, by the circuitry, a single result of the instruction as an output of a second operation of the two different operations; and

sending, by the circuitry, the single result to the register file.

9 . The method as recited in claim 8 , further comprising sending, by the circuitry, the indication to a plurality of arithmetic logic units (ALUs) in a same pipeline stage.

10 . The method as recited in claim 9 , further comprising generating, by each of the plurality of ALUs, a corresponding single result of the instruction.

11 . The method as recited in claim 10 , wherein the instruction is used as an activation function of a machine learning data model.

12 . The method as recited in claim 11 , further comprising generating, by the circuitry, an indication specifying the second operation of the two different operations comprises an addition operation.

13 . The method as recited in claim 12 , further comprising:

sending, by the circuitry, two of the three source operands to the comparison operation; and

sending, by the circuitry, one of the three source operands and a result of the comparison operation to the addition operation bypassing intermediate pipeline registers and the register file.

14 . The method as recited in claim 12 , further comprising:

sending, by the circuitry, two of the three source operands to the addition operation; and

sending, by the circuitry, one of the three source operands and a result of the addition operation to the comparison operation bypassing intermediate pipeline registers and the register file.

15 . A computing system comprising:

a memory comprising a plurality of instructions; and

a plurality of compute circuits, each comprising:

a register file configured to store data; and

circuitry configured to:

receive the plurality of instructions;

read, from the register file, three source operands indicated by a given instruction of the plurality of instructions;

responsive to the given instruction being a fused instruction comprising two different operations with one being a comparison operation:

generate an indication that a first operation of the two different operations comprises a comparison operation;

generate a single result of the given instruction as an output of a second operation of the two different operations; and

send the single result to the register file.

16 . The computing system as recited in claim 15 , wherein the circuitry is further configured to send the indication to a plurality of arithmetic logic units (ALUs) in a same pipeline stage.

17 . The computing system as recited in claim 16 , wherein each of the plurality of ALUs is configured to generate a corresponding single result of the given instruction of the plurality of instructions.

18 . The computing system as recited in claim 17 , wherein the given instruction is used as an activation function of a machine learning data model.

19 . The computing system as recited in claim 15 , wherein the circuitry is further configured to generate an indication specifying the second operation of the two different operations comprises an addition operation.

20 . The computing system as recited in claim 19 , wherein the circuitry is further configured to:

send two of the three source operands to the comparison operation; and

send one of the three source operands and a result of the comparison operation to the addition operation, bypassing intermediate pipeline registers and the register file.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2024
From: REEVES, STEVEN ISAAC
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 066826/0094 →
Continuity (1)
Related Publication 20250291594A1 · Sep 18, 2025
References Cited (7)
US 11487541B2 · Corbal et al. · 2022 [cited by applicant]
US 11693657B2 · Yudanov et al. · 2023 [cited by applicant]
US 11934834B2 · Mirkes · 2024 [cited by examiner]
US 20090019262A1 · Tashiro · 2009 [cited by examiner]
US 20220066760A1 · Chang · 2022 [cited by examiner]
US 20220129752A1 · Lagudu et al. · 2022 [cited by applicant]
US 20230305844A1 · Tyrlik · 2023 [cited by examiner]