IP Library › Granted Patent US 12,748,598
Granted Patent B2
US 12,748,598 · App. 18/100,190 · Granted Sep 29, 2026

Arithmetic processing apparatus which executes plurality of instructions in parallel and sequentially from executable instructions and method for arithmetic processing

Inventor: Gen Oshiyama (Kawasaki, JP)
Assignee: Fujitsu Limited
G06F9/3826G06F9/3838G06F9/3867
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,598
App. No.
18/100,190
Granted
Sep 29, 2026
Kind
B2
Abstract

An arithmetic processing apparatus includes a queue and control circuitry, and executes a plurality of instructions in parallel and sequentially from executable instructions. In the arithmetic processing apparatus, the queue stores the plurality of instructions, and the control circuitry holds an indicator indicating a pipeline that executes a producer instruction included in the plurality of instructions and an execution stage of the producer instruction in the pipeline, executes data dependency resolution between the producer instruction and a consumer instruction that uses an execution result of the producer instruction and that is included in the plurality of instructions, and controls issuing timings of the plurality of instructions.

Claims (39)

1 . An arithmetic processing apparatus comprising:

a queue; and

control circuitry, wherein

the arithmetic processing apparatus executes a plurality of instructions in parallel and sequentially from executable instructions,

the queue stores the plurality of instructions, and

the control circuitry holds an indicator indicating a pipeline that executes a producer instruction included in the plurality of instructions and uniquely identifying an execution stage on the pipeline of the producer instruction from which a result of an operation is to be forwarded, the pipeline has multiple types of latency, and the control circuitry executes data dependency resolution between the producer instruction and a consumer instruction that uses an execution result of the producer instruction and that is included in the plurality of instructions and controls issuing timings of the plurality of instructions,

wherein the queue stores a plurality of consumer instructions, and the control circuitry holds a plurality of indicators each indicating the pipeline of the producer instruction for each of the plurality of consumer instructions,

wherein the indicator comprises:

pipeline identification information of the pipeline to which the producer instruction is issued; and

stage identification information uniquely allocated to each stage, of stages from a stage corresponding to a first issuing timing at which the result of the operation is forwarded from the producer instruction to the consumer instruction at a shortest forwarding timing to a stage corresponding to a second issuing timing at which the result of the operation is stored in a register and comes to be ready to be read by the consumer instruction, and

wherein:

the stage identification information is set at a first cycle earlier by a first given number of cycles than a last cycle at which the producer instruction is executed, and is reset at a second cycle later by a second given number of cycles than the last cycle; and

unique identifiers are allocated one to each of one or more cycles from the first cycle to the second cycle.

2 . The arithmetic processing apparatus according to claim 1 , wherein the control circuitry generates a plurality of indicators one for each entry of the queue and each operand of the plurality of instructions in relation to each of a plurality of producer instructions.

3 . The arithmetic processing apparatus according to claim 2 , wherein the control circuitry executes the data dependency resolution, using the indicator retained in each operand of the plurality of instructions.

4 . The arithmetic processing apparatus according to claim 1 , further comprising forwarding control circuitry that forwards, based on the indicator, the result of the operation of the producer instruction to an operand of the consumer instruction, the consumer instruction being issued to a pipeline that executes the consumer instruction at a timing according to the control by the control circuitry.

5 . The arithmetic processing apparatus according to claim 1 , further comprising canceling control circuitry that cancels, when a preceding load instruction exists in a dependency chain of the producer instruction and the preceding load instruction results in a cache miss, all instructions in a subsequent dependency chain to the preceding load instruction, the canceling being based on the indicator.

6 . The arithmetic processing apparatus according to claim 5 , wherein the canceling control circuitry sets information representing whether the data dependency resolution succeeds or not to a value indicating that, due to the cache miss causing data dependency resolution failure, each producer instruction is awaiting re-issue by the control circuitry, on all instructions in the subsequent dependency chain to the preceding load instruction, in the canceling.

7 . The arithmetic processing apparatus according to claim 1 , further comprising updating control circuitry that controls, based on the indicator, updating of an Inflight Condition Flag.

8 . A method for arithmetic processing in an arithmetic processing apparatus that executes a plurality of instructions in parallel and sequentially from executable instructions, the method comprising:

at a scheduler provided in the arithmetic processing apparatus,

storing, in a queue, the plurality of instructions being accepted; and

holding an indicator indicating a pipeline that executes a producer instruction included in the plurality of instructions stored in the queue and uniquely identifying an execution stage on the pipeline of the producer instruction from which a result of an operation is to be forwarded, the pipeline having multiple types of latency;

executing data dependency resolution between the producer instruction and a consumer instruction that uses an execution result of the producer instruction and that is included in the plurality of instructions; and

controlling issuing timings of the plurality of instructions,

wherein the storing comprises storing, in the queue, a plurality of consumer instructions, and

wherein the holding comprises holding a plurality of indicators each indicating the pipeline of the producer instruction for each of the plurality of consumer instructions,

wherein the indicator comprises:

pipeline identification information of the pipeline to which the producer instruction is issued; and

stage identification information uniquely allocated to each stage, from a stage corresponding to a first issuing timing at which the result of the operation is forwarded from the producer instruction to the consumer instruction at a shortest forwarding timing to a stage corresponding to a second issuing timing at which the result of the arithmetic operation is stored in a register and comes to be ready to be read by the consumer instruction, and

wherein:

the stage identification information is set at a first cycle earlier by a first given number of cycles than a last cycle at which the producer instruction is executed, and is reset at a second cycle later by a second given number of cycles than the last cycle; and

unique identifiers are allocated one to each of one or more cycles from the first cycle to the second cycle.

9 . The method according to claim 8 , further comprising at the scheduler, generating a plurality of indicators one for each entry of the queue and each operand of the plurality of instructions in relation to each of a plurality of producer instructions.

10 . The method according to claim 9 , further comprising at the scheduler, executing the data dependency resolution, using the indicator retained in each operand of the plurality of instructions.

11 . The method according to claim 8 , further comprising at the scheduler, forwarding, based on the indicator, the result of the operation of the producer instruction to an operand of the consumer instruction, the consumer instruction being issued to a pipeline that executes the consumer instruction at a timing according to the controlling.

12 . The method according to claim 8 , further comprising at the scheduler, detecting a cache miss of a preceding load instruction in a dependency chain of the producer instruction, and canceling all instructions in a subsequent dependency chain to the preceding load instruction, the canceling being based on the indicator.

13 . The method according to claim 12 , further comprising at the scheduler, setting information representing whether the data dependency resolution succeeds or not to a value indicating that, due to the cache miss causing data dependency resolution failure, each producer instruction is awaiting re-issue by the scheduler, on all instructions in the subsequent dependency chain to the preceding load instruction, in the canceling.

14 . The method according to claim 8 , further comprising at the scheduler, controlling, based on the indicator, updating of an Inflight Condition Flag.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2023
From: OSHIYAMA, GEN
To: FUJITSU LIMITED
Reel/Frame 062453/0640 →
Priority Claims (2)
JP 2022-055143 · Mar 30, 2022 · national
JP 2022-196251 · Dec 8, 2022 · national
Continuity (1)
Related Publication 20230315446A1 · Oct 5, 2023
References Cited (18)
US 5944811A · Motomura · 1999 [cited by applicant]
US 10514925B1 · Reynolds · 2019 [cited by examiner]
US 20030182536A1 · Teruyama · 2003 [cited by examiner]
US 20050060518A1 · Augsburg · 2005 [cited by examiner]
US 20060095732A1 · Tran · 2006 [cited by examiner]
US 20070204135A1 · Jiang · 2007 [cited by examiner]
US 20090276608A1 · Shimada · 2009 [cited by examiner]
US 20130298127A1 · Meier et al. · 2013 [cited by applicant]
US 20130326197A1 · Brown · 2013 [cited by examiner]
US 20140013089A1 · Henry et al. · 2014 [cited by applicant]
US 20140129805A1 · Husby · 2014 [cited by examiner]
US 20140281431A1 · Iyengar · 2014 [cited by examiner]
US 20200183684A1 · Oshiyama · 2020 [cited by examiner]
US 20210089312A1 · Kothinti Naresh · 2021 [cited by examiner]
US 20230118428A1 · Mirkes · 2023 [cited by examiner]
JP H10078872A · 1998 [cited by applicant]
JP 2015232902A · 2015 [cited by applicant]
Office Action dated Jun. 9, 2026, issued in counterpart JP Application No. 2022-196251, with English translation. (6 pages). [cited by applicant]