IP Library › Granted Patent US 12,554,503
Granted Patent B2
US 12,554,503 · App. 18/646,992 · Granted Feb 17, 2026

Processor pipeline for interlocked data transfer operations with variable latency

Inventors: Ricardo Ramirez (Sunnyvale, CA); Albert Anthony Martin (Sunnyvale, CA); Abhijit Sil (Dublin, CA); Rabin Sugumar (Sunnyvale, CA)
Assignee: Akeana, Inc.
G06F9/3836G06F9/30014G06F9/3016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,554,503
App. No.
18/646,992
Granted
Feb 17, 2026
Kind
B2
Abstract

Disclosed embodiments provide techniques for instruction execution with a processor pipeline for data transfer operations. A processor core is accessed. The processor core executes one or more instructions out of order. The processor core supports integer operations and floating-point operations. An instruction in the processor core is decoded. The instruction is a data transfer operation. The data transfer operation necessitates a floating-point operation and an integer operation. The floating-point operation and the integer operation are dispatched to one or more issue queues. The floating-point operation and the integer operation are interlocked. The interlocking is accomplished using at least one entry in the one or more issue queues. A first operation of the floating-point operation and the integer operation is executed. A second operation of the floating-point operation and the integer operation is executed. The execution of the second operation is based on the interlocking.

Claims (44)

1 . A processor-implemented method for instruction execution comprising:

accessing a processor core, wherein the processor core executes one or more instructions out of order, and wherein the processor core supports integer operations and floating-point operations, wherein the integer operations and the floating-point operations include various latency operations, wherein the various latency operations include variable latency operations;

decoding an instruction in the processor core, wherein the instruction comprises a data transfer operation, and wherein the data transfer operation necessitates a floating-point operation and an integer operation;

dispatching the floating-point operation and the integer operation to one or more issue queues;

interlocking the floating-point operation and the integer operation, wherein the interlocking is accomplished using at least one entry in the one or more issue queues, wherein a request/grant protocol is employed for interlocking the variable latency operations;

executing a first operation of the floating-point operation and the integer operation; and

executing a second operation of the floating-point operation and the integer operation, based on the interlocking.

2 . The method of claim 1 wherein the one or more issue queues comprise a floating-point issue queue and an integer issue queue.

3 . The method of claim 1 wherein a latency of the various latency operations determines one or more fields in corresponding entries of a floating-point issue queue and an integer issue queue.

4 . The method of claim 1 wherein the at least one entry in the one or more issue queues includes one or more of a companion instruction queue identification, a dependency bit, and an operation type.

5 . The method of claim 3 wherein the companion instruction queue identification, the dependency bit, and/or the operation type enable the interlocking at an instruction level.

6 . The method of claim 1 wherein the at least one entry in the one or more issue queues includes a reorder buffer identification and/or an operand validity indication.

7 . The method of claim 6 wherein the reorder buffer identification and/or the operand validity indication enable interlocking at an out-of-order instruction level.

8 . The method of claim 1 wherein the interlocking enables marking the second operation eligible for issuing.

9 . The method of claim 8 wherein the interlocking enables issuing an operation from an instruction queue to an execution unit.

10 . The method of claim 9 wherein the interlocking enables a source operand for the second operation to be obtained from the first operation.

11 . The method of claim 1 wherein the first operation comprises a floating-point operation and the second operation comprises an integer operation.

12 . The method of claim 11 wherein the first operation is executed in a floating-point unit and the second operation is executed in an arithmetic logic unit.

13 . The method of claim 12 wherein data resulting from the first operation is transferred directly from the floating-point unit to the arithmetic logic unit.

14 . The method of claim 12 wherein data resulting from the first operation is transferred from the floating-point unit to the arithmetic logic unit using temporary register file storage.

15 . The method of claim 1 wherein the first operation comprises an integer operation and the second operation comprises a floating-point element operation.

16 . The method of claim 15 wherein the first operation is executed in an arithmetic logic unit and the second operation is executed in a floating-point unit.

17 . The method of claim 16 wherein data resulting from the first operation is transferred directly from the arithmetic logic unit to the floating-point unit.

18 . The method of claim 16 wherein data resulting from the first operation is transferred from the arithmetic logic unit to the floating-point unit using temporary register file storage.

19 . The method of claim 1 wherein the data transfer operations include floating-point convert operations, floating-point move operations, floating-point compare operations, or floating-point class operations.

20 . The method of claim 19 wherein the data transfer operations include single-precision floating-point operations and double-precision floating-point operations.

21 . The method of claim 1 wherein the one or more instructions include vector floating-point instructions.

22 . The method of claim 1 wherein the interlocking is managed by a decode unit.

23 . A computer program product embodied in a non-transitory computer readable medium for instruction execution, the computer program product comprising code which causes one or more processors to generate semiconductor logic for:

accessing a processor core, wherein the processor core executes one or more instructions out of order, and wherein the processor core supports integer operations and floating-point operations, wherein the integer operations and the floating-point operations include various latency operations, wherein the various latency operations include variable latency operations;

decoding an instruction in the processor core, wherein the instruction comprises a data transfer operation, and wherein the data transfer operation necessitates a floating-point operation and an integer operation;

dispatching the floating-point operation and the integer operation to one or more issue queues;

interlocking the floating-point operation and the integer operation, wherein the interlocking is accomplished using at least one entry in the one or more issue queues, wherein a request/grant protocol is employed for interlocking the variable latency operations;

executing a first operation of the floating-point operation and the integer operation; and

executing a second operation of the floating-point operation and the integer operation, based on the interlocking.

24 . A computer system for instruction execution comprising:

a memory which stores instructions;

one or more processors coupled to the memory wherein the one or more processors, when executing the instructions which are stored, are configured to:

access a processor core, wherein the processor core executes one or more instructions out of order, and wherein the processor core supports integer operations and floating-point operations, wherein the integer operations and the floating-point operations include various latency operations, wherein the various latency operations include variable latency operations;

decode an instruction in the processor core, wherein the instruction comprises a data transfer operation, and wherein the data transfer operation necessitates a floating-point operation and an integer operation;

dispatch the floating-point operation and the integer operation to one or more issue queues;

interlock the floating-point operation and the integer operation, wherein the interlocking is accomplished using at least one entry in the one or more issue queues, wherein a request/grant protocol is employed for interlocking the variable latency operations;

execute a first operation of the floating-point operation and the integer operation; and

execute a second operation of the floating-point operation and the integer operation, based on the interlocking.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2025
From: RAMIREZ, RICARDO; MARTIN, ALBERT ANTHONY; SIL, ABHIJIT; SUGUMAR, RABIN
To: AKEANA, INC.
Reel/Frame 071248/0392 →
Continuity (20)
Provisional Application 63570281 · Mar 27, 2024
Provisional Application 63564529 · Mar 13, 2024
Provisional Application 63563492 · Mar 11, 2024
Provisional Application 63563102 · Mar 8, 2024
Provisional Application 63556944 · Feb 23, 2024
Provisional Application 63556951 · Feb 23, 2024
Provisional Application 63605620 · Dec 4, 2023
Provisional Application 63602514 · Nov 24, 2023
Provisional Application 63547574 · Nov 7, 2023
Provisional Application 63547404 · Nov 6, 2023
Provisional Application 63546769 · Nov 1, 2023
Provisional Application 63545961 · Oct 27, 2023
Provisional Application 63542797 · Oct 6, 2023
Provisional Application 63526009 · Jul 11, 2023
Provisional Application 63521365 · Jun 16, 2023
Provisional Application 63471283 · Jun 6, 2023
Provisional Application 63467335 · May 18, 2023
Provisional Application 63463371 · May 2, 2023
Provisional Application 63462542 · Apr 28, 2023
Related Publication 20250217151A1 · Jul 3, 2025
References Cited (26)
US 5991863A · Dao · 1999 [cited by examiner]
US 6061781A · Jain · 2000 [cited by examiner]
US 6330657B1 · Col · 2001 [cited by examiner]
US 6519696B1 · Henry · 2003 [cited by examiner]
US 6934809B2 · Tremblay et al. · 2005 [cited by applicant]
US 7506105B2 · Al-Sukhni et al. · 2009 [cited by applicant]
US 8918625B1 · O'Bleness · 2014 [cited by examiner]
US 10013356B2 · Chou · 2018 [cited by applicant]
US 10671394B2 · Britto et al. · 2020 [cited by applicant]
US 10929948B2 · Benthin et al. · 2021 [cited by applicant]
US 11288405B2 · Belgarric et al. · 2022 [cited by applicant]
US 11403099B2 · Cerny et al. · 2022 [cited by applicant]
US 11403225B2 · Zheng et al. · 2022 [cited by applicant]
US 11429529B2 · Hornung et al. · 2022 [cited by applicant]
US 11442863B2 · Shulyak et al. · 2022 [cited by applicant]
US 11474130B2 · Lentz et al. · 2022 [cited by applicant]
US 11486911B2 · Tuncer et al. · 2022 [cited by applicant]
US 20090172349A1 · Sprangle · 2009 [cited by examiner]
US 20090198974A1 · Barowski · 2009 [cited by examiner]
US 20100306510A1 · Olson · 2010 [cited by examiner]
US 20180267798A1 · Grisenthwaite · 2018 [cited by examiner]
US 20220004639A1 · Yardi et al. · 2022 [cited by applicant]
US 20220019436A1 · Lloyd · 2022 [cited by examiner]
US 20220029780A1 · Dafali · 2022 [cited by applicant]
US 20220197657A1 · Soundararajan et al. · 2022 [cited by applicant]
WO 2022117687A1 · 2022 [cited by applicant]