IP Library › Granted Patent US 11,249,755
Granted Patent B2
US 11,249,755 · App. 16/642,786 · Granted Feb 15, 2022

Vector instructions for selecting and extending an unsigned sum of products of words and doublewords for accumulation

Inventors: Venkateswara R. Madduri (Austin, TX); Carl Murray (Dublin, IE); Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Mark J. Charney (Lexington, MA); Robert Valentine (Kiryat Tivon, IL); Jesus Corbal (King City, OR)
Assignee: Intel Corporation
G06F9/30036G06F9/3001G06F9/30098
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,249,755
App. No.
16/642,786
Granted
Feb 15, 2022
Kind
B2
Abstract

Disclosed embodiments relate to executing a vector unsigned multiplication and accumulation instruction. In one example, a processor includes fetch circuitry to fetch a vector unsigned multiplication and accumulation instruction having fields for an opcode, first and second source identifiers, a destination identifier, and an immediate, wherein the identified sources and destination are same-sized registers, decode circuitry to decode the fetched instruction, and execution circuitry to execute the decoded instruction, on each corresponding pair of first and second quadwords of the identified first and second sources, to: generate a sum of products of two doublewords of the first quadword and either two lower words or two upper words of the second quadword, based on the immediate, zero-extend the sum to a quadword-sized sum, and accumulate the quadword-sized sum with a previous value of a destination quadword in a same relative register position as the first and second quadwords.

Claims (37)

1. A processor comprising:

fetch circuitry to fetch a vector unsigned multiplication and accumulation instruction having fields for an opcode, first and second source identifiers, a destination identifier, and an immediate, wherein the identified sources and destination are same-sized registers;

decode circuitry to decode the fetched instruction; and

execution circuitry to execute the decoded instruction, on each corresponding pair of first and second quadwords of the identified first and second sources, to:

generate a sum of products of two double words of the first quadword and either two lower words or two upper words of the second quadword, based on the immediate;

zero-extend the sum to a quadword-sized sum; and

accumulate the quadword-sized sum with a previous value of a destination quadword in a same relative register position as the first and second quadwords.

2. The processor of claim 1 , wherein, when the immediate has a predefined value, the generated sum uses the two upper words of the second quadword.

3. The processor of claim 1 , wherein the generated sum is represented by at least 48 bits.

4. The processor of claim 1 , wherein each of the identified sources and destination comprise one of a 64-bit general purpose register, 128-bit vector register, 256-bit vector register, or 512-bit vector register.

5. The processor of claim 1 , wherein the vector unsigned multiplication and accumulation instruction further comprises a size identifier to specify the size of the same-sized registers.

6. The processor of claim 1 , wherein the vector unsigned multiplication and accumulation instruction further includes a write mask identifier to identify a write mask to conditionally control per-element computational operation and updating of results to the identified destination.

7. A method comprising:

fetching, by fetch circuitry, a vector unsigned multiplication and accumulation instruction having fields for an opcode, first and second source identifiers, a destination identifier, and an immediate, wherein the identified sources and destination are same-sized registers;

decoding, by decode circuitry, the fetched instruction; and

executing, by execution circuitry, the decoded instruction, on each corresponding pair of first and second quadwords of the identified first and second sources, to:

generate a sum of products of two double words of the first quadword and either two lower words or two upper words of the second quadword, based on the immediate;

zero-extend the sum to a quadword-sized sum; and

accumulate the quadword-sized sum with a previous value of a destination quadword in a same relative register position as the first and second quadwords.

8. The method of claim 7 , wherein, when the immediate has a predefined value, the generated sum uses the two upper words of the second quadword.

9. The method of claim 7 , wherein the generated sum is represented by at least 48 bits.

10. The method of claim 7 , wherein each of the identified sources and destination comprise a register selected from the group consisting of a 64-bit general purpose register, 128-bit vector register, 256-bit vector register, and 512-bit vector register.

11. The method of claim 7 , wherein the vector unsigned multiplication and accumulation instruction further comprises a size identifier to specify the size of the same-sized registers.

12. The method of claim 7 , wherein the vector unsigned multiplication and accumulation instruction further includes a write mask identifier to identify a write mask to conditionally control per-element computational operation and updating of results to the identified destination.

13. A system comprising:

a memory to store a vector unsigned multiplication and accumulation instruction; and

a processor coupled to the memory, the processor comprising:

fetch circuitry to fetch the vector unsigned multiplication and accumulation instruction having fields for an opcode, first and second source identifiers, a destination identifier, and an immediate, wherein the identified sources and destination are same-sized registers;

decode circuitry to decode the fetched instruction; and

execution circuitry to execute the decoded instruction, on each corresponding pair of first and second quadwords of the identified first and second sources, to:

generate a sum of products of two double words of the first quadword and either two lower words or two upper words of the second quadword, based on the immediate;

zero-extend the sum to a quadword-sized sum; and

accumulate the quadword-sized sum with a previous value of a destination quadword in a same relative register position as the first and second quadwords.

14. The system of claim 13 , wherein, when the immediate has a predefined value, the generated sum uses the two upper words of the second quadword.

15. The system of claim 13 , wherein the generated sum is represented by at least 48 bits.

16. The system of claim 13 , wherein each of the identified sources and destination comprise one of a 64-bit general purpose register, 128-bit vector register, 256-bit vector register, or 512-bit vector register.

17. The system of claim 13 , wherein the vector unsigned multiplication and accumulation instruction further comprises a size identifier to specify the size of the same-sized registers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2020
From: MADDURI, VENKATESWARA R.; MURRAY, CARL; OULD-AHMED-VALL, ELMOUSTAPHA; CHARNEY, MARK J.; VALENTINE, ROBERT; CORBAL, JESUS
To: INTEL CORPORATION
Reel/Frame 051994/0415 →
Continuity (1)
Related Publication 20200201633A1 · Jun 25, 2020