IP Library › Granted Patent US 11,907,712
Granted Patent B2
US 11,907,712 · App. 17/033,649 · Granted Feb 20, 2024

Methods, systems, and apparatuses for out-of-order access to a shared microcode sequencer by a clustered decode pipeline

Inventors: Thomas Madaelil (Austin, TX); Jonathan Combs (Austin, TX); Vikash Agarwal (Austin, TX)
Assignee: Intel Corporation
G06F9/223G06F9/382G06F9/3802G06F9/3822G06F9/3844
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,907,712
App. No.
17/033,649
Granted
Feb 20, 2024
Kind
B2
Abstract

Systems, methods, and apparatuses relating to circuitry to implement out-of-order access to a shared microcode sequencer by a clustered decode pipeline are described. In one embodiment, a hardware processor core includes a first decode cluster comprising a plurality of decoder circuits, a second decode cluster comprising a plurality of decoder circuits, a fetch circuit to fetch a first block of instructions and send the first block of instructions to the first decode cluster for decoding, and fetch a second block of instructions younger in program order than the first block of instructions and send the second block of instructions to the second decode cluster for decoding, a microcode sequencer comprising a memory that stores a plurality of micro-operations, and an arbitration circuit to arbitrate access by the first decode cluster and the second decode cluster to a shared read port of the memory, wherein the arbitration circuit is to allow the second decode cluster decoding the second block of instructions access to the shared read port of the memory instead of the first decode cluster decoding the first block of instructions when an instruction of the second block of instructions has a number of corresponding micro-operations in the microcode sequencer below an arbitration threshold.

Claims (41)

1. A hardware processor core comprising:

a first decode cluster comprising a plurality of decoder circuits;

a second decode cluster comprising a plurality of decoder circuits;

a fetch circuit to fetch a first block of instructions and send the first block of instructions to the first decode cluster for decoding, and fetch a second block of instructions younger in program order than the first block of instructions and send the second block of instructions to the second decode cluster for decoding;

a microcode sequencer comprising a memory that stores a plurality of micro-operations; and

an arbitration circuit to arbitrate access by the first decode cluster and the second decode cluster to a shared read port of the memory, wherein the arbitration circuit is to allow the second decode cluster decoding the second block of instructions access to the shared read port of the memory instead of the first decode cluster decoding the first block of instructions when an instruction of the second block of instructions has a number of corresponding micro-operations in the microcode sequencer below an arbitration threshold.

2. The hardware processor core of claim 1 , wherein the arbitration circuit it to, when the instruction of the second block of instructions has the number of corresponding micro-operations in the microcode sequencer above or equal to the arbitration threshold, cause a stall of access to the shared read port of the memory of the microcode sequencer by the second decode cluster until the second block of instructions is an oldest block of instructions in program order being decoded by the hardware processor core.

3. The hardware processor core of claim 2 , wherein the stall is a stall of decoding by the second decode cluster.

4. The hardware processor core of claim 1 , wherein the arbitration circuit is to allow the second decode cluster decoding the second block of instructions access to the shared read port of the memory instead of the first decode cluster decoding the first block of instructions when the instruction of the second block of instructions has the number of corresponding micro-operations in the microcode sequencer below the arbitration threshold and an instruction decode queue of the second decode cluster has available storage space for the number of corresponding micro-operations.

5. The hardware processor core of claim 1 , wherein the second decode cluster comprises a data structure to store one or more bits that indicate the number of corresponding micro-operations in the microcode sequencer for the instruction of the second block of instructions, and is to send the one or more bits to the arbitration circuit in response to a request to decode the instruction of the second block of instructions.

6. The hardware processor core of claim 5 , wherein the data structure of the second decode cluster is to store an entry point value that indicates an entry point in the memory for the corresponding micro-operations of the instruction of the second block of instructions, and the second decode cluster is to send the entry point value and the one or more bits to the arbitration circuit in response to the request to decode the instruction of the second block of instructions.

7. The hardware processor core of claim 1 , wherein the arbitration circuit is to allow the first decode cluster decoding the first block of instructions access to the shared read port of the memory instead of the second decode cluster decoding the second block of instructions when an instruction of the first block of instructions has one or more corresponding micro-operations in the microcode sequencer.

8. The hardware processor core of claim 1 , wherein the shared read port of the memory is the only read port into the memory of the microcode sequencer.

9. A method comprising:

sending a first block of instructions to a first decode cluster comprising a plurality of decoder circuits of a processor for decoding;

sending a second block of instructions younger in program order than the first block of instructions to a second decode cluster comprising a plurality of decoder circuits of the processor for decoding; and

arbitrating access, via an arbitration circuit of the processor, by the first decode cluster and the second decode cluster to a shared read port of a microcode sequencer comprising a memory that stores a plurality of micro-operations to allow the second decode cluster decoding the second block of instructions access to the shared read port instead of the first decode cluster decoding the first block of instructions when an instruction of the second block of instructions has a number of corresponding micro-operations in the microcode sequencer below an arbitration threshold.

10. The method of claim 9 , wherein the arbitrating comprises, when the instruction of the second block of instructions has the number of corresponding micro-operations in the microcode sequencer above or equal to the arbitration threshold, cause a stall of access to the shared read port of the memory of the microcode sequencer by the second decode cluster until the second block of instructions is an oldest block of instructions in program order being decoded by the first decode cluster and the second decode cluster.

11. The method of claim 10 , wherein the stall is a stall of decoding by the second decode cluster.

12. The method of claim 9 , wherein the arbitrating comprises allowing the second decode cluster decoding the second block of instructions access to the shared read port of the memory instead of the first decode cluster decoding the first block of instructions when the instruction of the second block of instructions has the number of corresponding micro-operations in the microcode sequencer below the arbitration threshold and an instruction decode queue of the second decode cluster has available storage space for the number of corresponding micro-operations.

13. The method of claim 9 , further comprising:

reading one or more bits from a data structure of the second decode cluster that indicate the number of corresponding micro-operations in the microcode sequencer for the instruction of the second block of instructions; and

sending the one or more bits to the arbitration circuit in response to a request to decode the instruction of the second block of instructions.

14. The method of claim 13 , further comprising:

reading an entry point value from the data structure of the second decode cluster that indicates an entry point in the memory for the corresponding micro-operations of the instruction of the second block of instructions; and

sending the entry point value and the one or more bits to the arbitration circuit in response to the request to decode the instruction of the second block of instructions.

15. The method of claim 9 , wherein the arbitrating comprises allowing the first decode cluster decoding the first block of instructions access to the shared read port of the memory instead of the second decode cluster decoding the second block of instructions when an instruction of the first block of instructions has one or more corresponding micro-operations in the microcode sequencer.

16. The method of claim 9 , wherein the shared read port of the memory is the only read port into the memory of the microcode sequencer.

17. A hardware processor core comprising:

a first decode cluster comprising a plurality of decoder circuits;

a second decode cluster comprising a plurality of decoder circuits;

a branch predictor to identify a first block of instructions and a second block of instructions younger in program order than the first block of instructions, cause the first block of instructions to be sent to the first decode cluster for decoding, and cause the second block of instructions to be sent to the second decode cluster for decoding;

a microcode sequencer comprising a memory that stores a plurality of micro-operations; and

an arbitration circuit to arbitrate access by the first decode cluster and the second decode cluster to a shared read port of the memory, wherein the arbitration circuit is to allow the second decode cluster decoding the second block of instructions access to the shared read port of the memory instead of the first decode cluster decoding the first block of instructions when an instruction of the second block of instructions has a number of corresponding micro-operations in the microcode sequencer below an arbitration threshold.

18. The hardware processor core of claim 17 , wherein the arbitration circuit it to, when the instruction of the second block of instructions has the number of corresponding micro-operations in the microcode sequencer above or equal to the arbitration threshold, cause a stall of access to the shared read port of the memory of the microcode sequencer by the second decode cluster until the second block of instructions is an oldest block of instructions in program order being decoded by the hardware processor core.

19. The hardware processor core of claim 18 , wherein the stall is a stall of decoding by the second decode cluster.

20. The hardware processor core of claim 17 , wherein the arbitration circuit is to allow the second decode cluster decoding the second block of instructions access to the shared read port of the memory instead of the first decode cluster decoding the first block of instructions when the instruction of the second block of instructions has the number of corresponding micro-operations in the microcode sequencer below the arbitration threshold and an instruction decode queue of the second decode cluster has available storage space for the number of corresponding micro-operations.

21. The hardware processor core of claim 17 , wherein the second decode cluster comprises a data structure to store one or more bits that indicate the number of corresponding micro-operations in the microcode sequencer for the instruction of the second block of instructions, and is to send the one or more bits to the arbitration circuit in response to a request to decode the instruction of the second block of instructions.

22. The hardware processor core of claim 21 , wherein the data structure of the second decode cluster is to store an entry point value that indicates an entry point in the memory for the corresponding micro-operations of the instruction of the second block of instructions, and the second decode cluster is to send the entry point value and the one or more bits to the arbitration circuit in response to the request to decode the instruction of the second block of instructions.

23. The hardware processor core of claim 17 , wherein the arbitration circuit is to allow the first decode cluster decoding the first block of instructions access to the shared read port of the memory instead of the second decode cluster decoding the second block of instructions when an instruction of the first block of instructions has one or more corresponding micro-operations in the microcode sequencer.

24. The hardware processor core of claim 17 , wherein the shared read port of the memory is the only read port into the memory of the microcode sequencer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2021
From: MADAELIL, THOMAS; COMBS, JONATHAN; AGARWAL, VIKASH
To: INTEL CORPORATION
Reel/Frame 055203/0560 →
Continuity (1)
Related Publication 20220100500A1 · Mar 31, 2022