IP Library Granted Patent US 7,793,080
Granted Patent B2
US 7,793,080 · App. 11/967,924 · Granted Sep 7, 2010

Processing pipeline having parallel dispatch and method thereof

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,793,080
App. No.
11/967,924
Granted
Sep 7, 2010
Kind
B2
Abstract

One or more processor cores of a multiple-core processing device each can utilize a processing pipeline having a plurality of execution units (e.g., integer execution units or floating point units) that together share a pre-execution front-end having instruction fetch, decode and dispatch resources. Further, one or more of the processor cores each can implement dispatch resources configured to dispatch multiple instructions in parallel to multiple corresponding execution units via separate dispatch buses. The dispatch resources further can opportunistically decode and dispatch instruction operations from multiple threads in parallel so as to increase the dispatch bandwidth. Moreover, some or all of the stages of the processing pipelines of one or more of the processor cores can be configured to implement independent thread selection for the corresponding stage.

Claims (47)

1. A processing device comprising:

a first execution unit;

a second execution unit; and

a front-end unit coupled to the first execution unit via a first dispatch bus and coupled to the second execution unit via a second dispatch bus separate from the first dispatch bus, the first dispatch bus configured to concurrently transmit a first dispatch group of up to N microcode operations from the front-end unit to the first execution unit for a dispatch cycle and the second dispatch bus configured to concurrently transmit a second dispatch group of up to N microcode operations from the front-end unit to the second execution unit for the dispatch cycle;

wherein the front-end unit comprises:

a dispatch module coupled to the first dispatch bus and the second dispatch bus, the dispatch module comprising:

a dispatch buffer configured to buffer microcode operations for dispatch; and

a dispatch controller configured to select microcode operations from the dispatch buffer for inclusion in the first and second dispatch groups; and

a decode module comprising a plurality of parallel decode paths, each decode path comprising:

a microcode decoder comprising a microcode table, the microcode decoder configured to decode an instruction into a set of one or more microcode operations based on the microcode table;

a first format decoder configured to format each microcode operation output by the microcode decoder according to a dispatch format and provide the resulting formatted microcode operation for storage in the dispatch buffer;

a fastpath hardware decoder configured to decode an instruction into a set of one or more microcode operations; and

a second format decoder configured to format each microcode operation output by the fastpath hardware decoder according to the dispatch format and provide the resulting formatted microcode operation for storage in the dispatch buffer.

2. The processing device of claim 1 , wherein N is four.

3. The processing device of claim 1 , wherein the first execution unit comprises a first integer execution unit and the second execution unit comprises a second integer execution unit.

4. The processing device of claim 1 , wherein the first execution unit comprises an integer execution unit and the second execution unit comprises a floating point unit.

5. The processing device of claim 1 , wherein the dispatch controller is configured to select microcode operations from the dispatch buffer based on dispatch criteria comprising at least one of: a maximum number of load microcode operations per dispatch group; a maximum number of store nmicrocode operations per dispatch group; a number of available entries in a retirement queue; and a number of available entries in a scheduler queue.

6. The processing device of claim 1 , further comprising:

a decode controller configured to:

direct instructions of a first thread to the microcode decoders of the decode paths for decoding to generate a first set of microcode operations for the first thread; and

direct instructions of a second thread to the fastpath hardware decoders of the decode paths for decoding to generate a second set of microcode operations for the second thread concurrent with the generation of the first set of microcode operations; and

wherein the first dispatch group includes microcode operations from the first set and the second dispatch group includes microcode operations from the second set.

7. The processing device of claim 1 , further comprising:

a decode controller configured to:

direct instructions of a thread to the microcode decoders of the decode paths for decoding to generate a first set of microcode operations for the thread; and

direct instructions of the thread to the fastpath hardware decoders of the decode paths for decoding to generate a second set of microcode operations for the thread; and

wherein the first dispatch group includes microcode operations from the first set and the second dispatch group includes microcode operations from the second set.

8. The processing device of claim 1 , wherein the dispatch controller is configured to:

in response to a thread switch at the processing device from a first thread to a second thread, select microcode operations associated with the first thread from the dispatch buffer for concurrent dispatch with microcode operations associated with the second thread until the dispatch buffer is devoid of microcode operations associated with the first thread.

9. The processing device of claim 1 , wherein the dispatch controller is disposed between the first execution unit and the second execution unit.

10. A method comprising:

providing a processing device comprising a first execution unit, a second execution unit, and a front-end unit coupled to the first execution unit via a first dispatch bus and coupled to the second execution unit via a second dispatch bus separate from the first dispatch bus;

concurrently dispatching a first dispatch group of up to N microcode operations from the front-end unit to the first execution unit via the first dispatch bus for a dispatch cycle and a second dispatch group of up to N microcode operations from the front-end unit to the second execution unit via the second dispatch bus for the dispatch cycle;

decoding instructions of a first thread using a plurality of microcode decoders of the processing device to generate a first set of one or more microcode operations for the first thread;

decoding instructions of a second thread using a plurality of fastpath hardware decoders of the processing device generate a second set of one or more microcode operations for the second thread concurrent with the generation of the first set of microcode operations; and

wherein the first dispatch group includes microcode operations from the first set and the second dispatch group includes microcode operations from the second set.

11. The method of claim 10 , wherein N is four.

12. The method of claim 10 , wherein the first execution unit comprises a first integer execution unit and the second execution unit comprises a second integer execution unit.

13. The method of claim 12 , wherein the first execution unit comprises an integer execution unit and the second execution unit comprises a floating point unit.

14. The method of claim 10 ,

wherein the first dispatch group and the second dispatch group each includes microcode operations from at least one of the first set and the second set.

15. The method of claim 10 , further comprising:

in response to a thread switch at the processing device from a first thread to a second thread, concurrently dispatching microcode operations associated with the first thread from and microcode operations associated with the second thread from a dispatch buffer until the dispatch buffer is devoid of microcode operations associated with the first thread.

16. A method comprising:

providing a processing device comprising a first execution unit, a second execution unit, and a front-end unit coupled to the first execution unit via a first dispatch bus and coupled to the second execution unit via a second dispatch bus separate from the first dispatch bus;

concurrently dispatching a first dispatch group of up to N microcode operations from the front-end unit to the first execution unit via the first dispatch bus for a dispatch cycle and a second dispatch group of up to N microcode operations from the front-end unit to the second execution unit via the second dispatch bus for the dispatch cycle; and

wherein concurrently dispatching instruction operations comprises selecting microcode operations for dispatch based on dispatch criteria comprising at least one of: a maximum number of load microcode operations per dispatch group; a maximum number of store microcode operations per dispatch group; a number of available entries in a retirement queue; and a number of available entries in a scheduler queue.

Assignments (7)
RELEASE OF SECURITY INTEREST Recorded May 12, 2021
From: WILMINGTON TRUST, NATIONAL ASSOCIATION
To: GLOBALFOUNDRIES U.S. INC.
Reel/Frame 056987/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2021
From: GLOBALFOUNDRIES US INC.
To: MEDIATEK INC.
Reel/Frame 055173/0781 →
RELEASE OF SECURITY INTEREST Recorded Nov 20, 2020
From: WILMINGTON TRUST, NATIONAL ASSOCIATION
To: GLOBALFOUNDRIES INC.
Reel/Frame 054636/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2020
From: GLOBALFOUNDRIES INC.
To: GLOBALFOUNDRIES U.S. INC.
Reel/Frame 054633/0001 →
SECURITY AGREEMENT Recorded Nov 29, 2018
From: GLOBALFOUNDRIES INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION
Reel/Frame 049490/0001 →
AFFIRMATION OF PATENT ASSIGNMENT Recorded Aug 18, 2009
From: ADVANCED MICRO DEVICES, INC.
To: GLOBALFOUNDRIES INC.
Reel/Frame 023120/0426 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 8, 2008
From: SHEN, GENE; LIE, SEAN
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 021492/0458 →