IP Library Granted Patent US 11,983,538
Granted Patent B2
US 11,983,538 · App. 17/659,569 · Granted May 14, 2024

Load-store unit dual tags and replays

Inventors: Robert T. Golla (Austin, TX); Ajay A. Ingle (Austin, TX)
Assignee: Cadence Design Systems, Inc.
G06F9/3834G06F12/0855
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,983,538
App. No.
17/659,569
Granted
May 14, 2024
Kind
B2
Abstract

Techniques are disclosed relating to a processor load-store unit. In some embodiments, the load-store unit is configured to execute load/store instructions in parallel using first and second pipelines and first and second tag memory arrays. In tag write conflict situations, the load-store unit may arbitrate between the first and second pipelines to ensure the first and second tag memory array contents remain identical. In some embodiments, a data cache tag replay scheme is utilized. In some embodiments, executing load/store instructions in parallel with fills, probes, and store-updates, using separate but identical tag memory arrays, may advantageously improve performance.

Claims (41)

1. An apparatus, comprising:

a processor configured to execute program instructions;

cache circuitry; and

a load-store unit configured to:

perform multiple types of memory access instructions executed by the processor, using first and second pipelines in parallel;

determine whether memory access instructions hit in the cache circuitry, including to:

use a first tag memory array for the first pipeline; and

use a second tag memory array for the second pipeline; and

control the first and second tag memory arrays such that they store matching tag information, wherein the load-store unit further comprises control circuitry,

configured to replay one or more instructions of a pipeline that loses an arbitration, and wherein the control circuitry is further configured to determine an oldest dependent instruction and to replay the corresponding instruction as well as all instructions younger than and including the oldest dependent instruction of the first pipeline.

2. The apparatus of claim 1 , wherein the second pipeline takes priority over the first pipeline and the control circuitry is further configured to replay a corresponding instruction as well as all younger instructions of the first pipeline.

3. The apparatus of claim 1 , wherein the second pipeline takes priority over the first pipeline.

4. The apparatus of claim 1 , wherein the load-store unit is configured to allow at most one of the first and second pipelines to write to the first and second tag memory arrays in a given cycle.

5. The apparatus of claim 1 , wherein to control the first and second tag memory arrays, the load-store unit is configured to write a same value to both the first and second tag memory arrays in response to either one of the first and second pipelines writing a tag.

6. The apparatus of claim 1 , wherein the multiple types of memory access instructions include:

a first subset of memory access types that the first pipeline is configured to perform; and

a second subset of memory access types that the second pipeline is configured to perform.

7. The apparatus of claim 6 , wherein the first subset of memory access types includes the following types that are not included in the second subset: load instructions, store instructions, and atomic operations.

8. The apparatus of claim 6 , wherein the second subset of memory access types includes the following types that are not included in the first subset: fills, probes, and store-updates.

9. A method, comprising:

performing, by a load-store unit of a processor, multiple types of memory access instructions executed by a processor, using first and second pipelines in parallel;

determining, by the load-store unit, whether memory access instructions hit in cache circuitry, including:

using a first tag memory array for the first pipeline; and

using a second tag memory array for the second pipeline;

controlling the first and second tag memory arrays such that they store matching tag information;

arbitrating, by control circuitry, between the first and second pipelines in response to an attempt for the first and second pipeline to write to the first and second tag memory arrays in a given cycle; and

replaying one or more instructions of a pipeline that loses an arbitration, wherein the control circuitry is further configured to determine an oldest dependent instruction and to replay the corresponding instruction as well as all instructions younger than and including the oldest dependent instruction of the first pipeline.

10. The method of claim 9 , further comprising:

wherein the first pipeline always loses arbitration if both the first and second pipelines attempt to write the first and second tag memory arrays in a given cycle.

11. The method of claim 9 , further comprising:

allowing at most one of the first and second pipelines to write to the first and second tag memory arrays in a given cycle.

12. A non-transitory computer readable storage medium having stored thereon design information that specifies a design of at least a portion of a hardware integrated circuit in a format recognized by a semiconductor fabrication system that is configured to use the design information to produce the circuit according to the design, wherein the design information specifies that the circuit includes

a processor configured to execute program instructions;

cache circuitry; and

a load-store unit configured to:

perform multiple types of memory access instructions executed by the processor, using first and second pipelines in parallel;

determine whether memory access instructions hit in the cache circuitry, including to:

use a first tag memory array for the first pipeline; and

use a second tag memory array for the second pipeline;

control the first and second tag memory arrays such that they store matching tag information, wherein the load-store unit further comprises control circuitry configured to arbitrate between the first and second pipelines in response to an attempt for the first and second pipelines to both write to the tag memory arrays in a given cycle; and

replay one or more instructions of a pipeline that loses an arbitration, and wherein the control circuitry is further configured to determine an oldest dependent instruction and to replay the corresponding instruction as well as all instructions younger than and including the oldest dependent instruction of the first pipeline.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2022
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: CADENCE DESIGN SYSTEMS, INC.
Reel/Frame 060334/0340 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2022
From: GOLLA, ROBERT T.; INGLE, AJAY A.
To: WESTERN DIGITAL TECHNOLOGIES, INC.
Reel/Frame 059625/0246 →
Continuity (1)
Related Publication 20230333856A1 · Oct 19, 2023