IP Library › Granted Patent US 12,210,874
Granted Patent B2
US 12,210,874 · App. 18/335,412 · Granted Jan 28, 2025

Processing for vector load or store micro-operation with inactive mask elements

Inventor: Yueh Chi Wu (Taichung, TW)
Assignee: SiFive, Inc.
G06F9/30036
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,210,874
App. No.
18/335,412
Granted
Jan 28, 2025
Kind
B2
Abstract

Apparatus and methods for processing of a vector load or store micro-operation with mask information as a no-operation (no-op) when a mask vector for the vector load or store micro-operation has all inactive mask elements or processing vector load or store sub-micro-operation(s) with active mask element(s) are described. An integrated circuit includes a load store unit configured to receive load or store micro-operations cracked from a vector load or store operation, determine that a mask vector for the vector load or store micro-operation is fully inactive, and process the vector load or store micro-operation as a no-operation. If the mask vector is not fully inactive, the vector load or store micro-operation is unrolled into vector load or store sub-micro-operation(s) which have active mask element(s). Vector load or store sub-micro-operation(s) which have inactive mask element(s) are ignored.

Claims (48)

1. An integrated circuit comprising:

a load store unit configured to:

receive load or store micro-operations cracked from a vector load or store operation;

determine that a mask vector for the vector load or store micro-operation is fully inactive; and

process the vector load or store micro-operation as a no-operation for a fully inactive mask vector.

2. The integrated circuit of claim 1 , the load store unit further configured to:

modify the vector load or store micro-operation to execute as the no-operation.

3. The integrated circuit of claim 2 , wherein modification of the vector load or store micro-operation includes modification of at least one of an operand, source, or destination of the vector load or store micro-operation.

4. The integrated circuit of claim 1 , the load store unit further configured to:

unroll each load or store micro-operation into multiple sub-micro-operations which have mask elements in the mask vector that are active.

5. The integrated circuit of claim 4 , the load store unit further configured to:

process an active vector load or store sub-micro-operation as normal.

6. The integrated circuit of claim 5 , the load store unit further configured to:

ignore an inactive vector load or store sub-micro-operation.

7. The integrated circuit of claim 6 , further comprising:

a Baler configured to receive load or store micro-operations cracked from the vector load or store operation and load at least the mask vector for access by the load store unit.

8. A method comprising:

receiving, at a load store unit, a vector load or store micro-operation;

determining, by the load store unit, that a mask vector for the vector load or store micro-operation is fully inactive; and

processing, by the load store unit, the vector load or store micro-operation as a no-operation for a fully inactive mask vector.

9. The method of claim 8 , further comprising:

modifying, by the load store unit, the vector load or store micro-operation to execute as the no-operation.

10. The method of claim 9 , the modifying further comprising:

modifying, by the load store unit, at least one of an operand, source, or destination of the vector load or store micro-operation.

11. The method of claim 9 , further comprising:

unrolling, by the load store unit, each load or store micro-operation into multiple sub-micro-operations which have mask elements in the mask vector that are active.

12. The method of claim 11 , further comprising:

processing, by the load store unit, an active vector load or store sub-micro-operation as normal.

13. The method of claim 12 , further comprising:

ignoring, by the load store unit, an inactive vector load or store sub-micro-operation.

14. The method of claim 13 , further comprising:

receiving, at a Baler unit, the vector load or store micro-operation; and

loading, by the Baler unit, at least the mask vector for access by the load store unit.

15. A non-transitory computer readable medium comprising a circuit representation that, when processed by a computer, is used to program or manufacture an integrated circuit comprising:

a load store unit configured to:

receive load or store micro-operations cracked from a vector load or store operation;

determine that a mask vector for the vector load or store micro-operation is fully inactive; and

process the vector load or store micro-operation as a no-operation for a fully inactive mask vector.

16. The non-transitory computer readable medium of claim 15 , wherein the circuit representation when processed by the computer is further used to program or manufacture the load store unit such that the load store unit is further configured to:

modify the vector load or store micro-operation to execute as the no-operation.

17. The non-transitory computer readable medium of claim 16 , wherein modification of the vector load or store micro-operation includes modification of at least one of an operand, source, or destination of the vector load or store micro-operation.

18. The non-transitory computer readable medium of claim 15 , wherein the circuit representation when processed by the computer is further used to program or manufacture the load store unit such that the load store unit is further configured to:

unroll each load or store micro-operation into multiple sub-micro-operations which have mask elements in the mask vector that are active; and

process an active vector load or store sub-micro-operation as normal.

19. The non-transitory computer readable medium of claim 15 , wherein the circuit representation when processed by the computer is further used to program or manufacture the load store unit such that the load store unit is further configured to:

ignore an inactive vector load or store sub-micro-operation.

20. The non-transitory computer readable medium of claim 15 , wherein the circuit representation when processed by the computer is further used to program or manufacture the integrate circuit such that the integrated circuit further comprises:

a Baler configured to receive load or store micro-operations cracked from the vector load or store operation and load at least the mask vector for access by the load store unit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2023
From: WU, YUEH CHI
To: SIFIVE, INC.
Reel/Frame 063961/0563 →
Continuity (2)
Provisional Application 63435588 · Dec 28, 2022
Related Publication 20240220250A1 · Jul 4, 2024
References Cited (1)
US 20190278577A1 · Plotnikov · 2019 [cited by examiner]