IP Library Granted Patent US 9,665,368
Granted Patent B2
US 9,665,368 · App. 13/631,666 · Granted May 30, 2017

Systems, apparatuses, and methods for performing conflict detection and broadcasting contents of a register to data element positions of another register

Inventors: Christopher J. Hughes (Santa Clara, CA); Mark J. Charney (Lexington, MA); Jesus Corbal (Barcelona, ES); Milind B. Girkar (Sunnyvale, CA); Elmoustapha Ould-Ahmed_Vall (Chandler, AZ); Bret L. Toll (Hillsboro, OR); Robert Valentine (Kiryat Tivon, IL)
Assignee: Intel Corporation
G06F9/3001G06F9/30018G06F9/30021G06F9/30036G06F9/30043
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,665,368
App. No.
13/631,666
Granted
May 30, 2017
Kind
B2
Abstract

Systems, apparatuses, and methods of performing in a computer processor broadcasting data in response to a single vector packed broadcasting instruction that includes a source writemask register operand, a destination vector register operand, and an opcode. In some embodiments, the data of the source writemask register is zero extended prior to broadcasting.

Claims (43)

1. A method of comprising:

executing a single instruction that includes a source writemask register operand, a source vector register operand, a destination writemask register operand, and an opcode to:

logically AND data from the source writemask register operand with each data element of the source vector register operand,

determine of which of the logical AND operations indicate a conflict to create a conflict check result, and

logically AND the conflict check result with the data from the source writemask operand; and

storing the result of the logical ANDing of the conflict check result with the data from the source writemask operand into the destination writemask register operand.

2. The method of claim 1 , further comprising:

zero extending data of the source writemask register operand such that the zero extended data will be of the same size as each data element of the source vector register operand.

3. The method of claim 1 , further comprising:

broadcasting the zero extended data of the source writemask register operand to a temporary vector register that has a same number and size data elements as the source vector register operand.

4. The method of claim 1 , wherein the source vector register operand is of size 128-bit, 256-bit, or 512-bit.

5. The method of claim 1 , wherein the destination writemask register operand is 64 bits.

6. The method of claim 1 , wherein the destination writemask register operand is 16 bits.

7. The method of claim 1 , wherein data elements of the source vector register operand are of 8-bit, 16-bit, 32-bit, 64-bit, 128-bit, or 256-bit in size.

8. An apparatus comprising:

decode circuitry to decode a single instruction that includes a source writemask register operand, a source vector register operand, a destination writemask register operand, and an opcode;

execution circuitry to execute the decoded single vector packed conflict testing instruction to:

logically AND data from the source writemask register operand with each data element of the source vector register operand,

determine of which of the logical AND operations indicate a conflict to create a conflict check result, and

logically AND the conflict check result with the data from the source writemask operand,

store the result of the logical ANDing of the conflict check result with the data from the source writemask operand into the destination writemask register operand.

9. The apparatus of claim 8 , wherein the execution circuitry to further:

zero extend data of the source writemask register operand such that the zero extended data will be of the same size as each data element of the source vector register operand.

10. The apparatus of claim 8 , wherein the execution circuitry to further:

broadcast the zero extended data of the source writemask register operand to a temporary vector register that has a same number and size data elements as the source vector register operand.

11. The apparatus of claim 8 , wherein the source vector register operand is of size 128-bit, 256-bit, or 512-bit.

12. The apparatus of claim 8 , wherein the destination writemask register operand is 64 bits.

13. The apparatus of claim 8 , wherein the destination writemask register operand is 16 bits.

14. The apparatus of claim 8 , wherein data elements of the source vector register operand are of 8-bit, 16-bit, 32-bit, 64-bit, 128-bit, or 256-bit in size.

15. A non-transitory machine-readable medium storing an instruction which when executed by a hardware processor to cause the hardware processor to perform a method, the method comprising:

executing a single instruction that includes a source writemask register operand, a source vector register operand, a destination writemask register operand, and an opcode to:

logically AND data from the source writemask register operand with each data element of the source vector register operand,

determine of which of the logical AND operations indicate a conflict to create a conflict check result, and

logically AND the conflict check result with the data from the source writemask operand; and

storing the result of the logical ANDing of the conflict check result with the data from the source writemask operand into the destination writemask register operand.

16. The non-transitory machine-readable medium of claim 15 , wherein the method further comprises:

zero extending data of the source writemask register operand such that the zero extended data will be of the same size as each data element of the source vector register operand.

17. The non-transitory machine-readable medium of claim 15 , wherein the method further comprises:

broadcasting the zero extended data of the source writemask register operand to a temporary vector register that has a same number and size data elements as the source vector register operand.

18. The non-transitory machine-readable medium of claim 15 , wherein the source vector register operand is of size 128-bit, 256-bit, or 512-bit.

19. The non-transitory machine-readable medium of claim 15 , wherein the destination writemask register operand is 64 bits.

20. The non-transitory machine-readable medium of claim 15 , wherein the destination writemask register operand is 16 bits.

21. The non-transitory machine-readable medium of claim 15 , wherein data elements of the source vector register operand are of 8-bit, 16-bit, 32-bit, 64-bit, 128-bit, or 256-bit in size.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2016
From: HUGHES, CHRISTOPHER J.; CHARNEY, MARK J.; CORBAL, JESUS; GIRKAR, MILIND B.; OULD-AHMED-VALL, ELMOUSTAPHA; TOLL, BRET L.
To: INTEL CORPORATION
Reel/Frame 037959/0914 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 17, 2014
From: VALENTINE, ROBERT
To: INTEL CORPORATION
Reel/Frame 033332/0269 →
Continuity (1)
Related Publication 20140095843A1 · Apr 3, 2014