IP Library Granted Patent US 12,197,921
Granted Patent B2
US 12,197,921 · App. 18/145,801 · Granted Jan 14, 2025

Accelerating eight-way parallel Keccak execution

Inventors: Santosh Ghosh (Hillsboro, OR); Christoph Dobraunig (St. Veit an der Glan, AT); Manoj Sastry (Portland, OR)
Assignee: Intel Corporation
G06F9/3885G06F9/3016G06F9/3802
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,197,921
App. No.
18/145,801
Granted
Jan 14, 2025
Kind
B2
Abstract

A method comprises fetching, by fetch circuitry, an encoded XOR3PP instruction comprising at least one opcode, a first source identifier to identify a first register, a second source identifier to identify a second register, a third source identifier to identifier a third register, and a fourth source identifier to identify a fourth operand, wherein the first register is to store a first value, the second register is to store a second value, and the third register is to store a third value, decoding, by decode circuitry, the encoded XOR3PP instruction to generate a decoded XOR3PP instruction; and executing, by execution circuitry, the decoded XOR3PP instruction to determine a first rotational value and a second rotational value, perform a rotate operation on at least a portion of the first value based on the first rotational value to generate a rotated third value, perform an XOR operation on at least a portion of the first value, at least a portion of the second value, and the rotated third value to generate an XOR result, perform a rotate operation on the XOR result based on the second rotational value to generate a rotated XOR; and store the rotated XOR result.

Claims (45)

1. A hardware processor, comprising:

fetch circuitry to fetch an encoded instruction comprising at least one opcode, a first source identifier to identify a first register to store a first operand value, a second source identifier to identify a second register to store a second operand value, a third source identifier to identify a third register to store a third operand value, and a fourth operand value;

decode circuitry to decode the encoded instruction to generate a decoded instruction; and

execution circuitry to execute the decoded instruction to:

determine a first rotational value and a second rotational value;

perform a first rotate operation using the first rotational value, wherein the first rotate operation is to include a rotation of a 3 operand XOR result generated based on at least a portion of the first operand value, at least a portion of the second operand value, and at least a portion of the third operand value to generate a first rotated result;

perform a second rotate operation on the first rotated result using the second rotational value to generate a second rotated result; and

store the second rotated result.

2. The hardware processor of claim 1 , further comprising:

commit circuitry to commit a result of the executed instruction.

3. The hardware processor of claim 1 , wherein the first register is a 512-bit register that stores a 64-bit operand value.

4. The hardware processor of claim 1 , wherein the first register stores eight 64-bit words that separately represent operand values.

5. The hardware processor of claim 4 , wherein the executed instruction is performed on the operand values separately represented by all eight 64-bit words.

6. The hardware processor of claim 1 , wherein the first rotational value and the second rotational value are based on a single bit of a multi-bit representation of the fourth operand value.

7. The hardware processor of claim 1 , wherein the first, the second and the third registers are separately capable of storing 512-bits of data to store respective first, second and third operand values and the fourth operand value is an 8-bit integer immediate.

8. A method, comprising:

fetching, by fetch circuitry, an encoded instruction comprising at least one opcode, a first source identifier to identify a first register to store a first operand value, a second source identifier to identify a second register to store a second operand value, a third source identifier to identify a third register to store a third operand value, and a fourth operand value;

decoding, by decode circuitry, the encoded instruction to generate a decoded instruction; and

executing, by execution circuitry, the decoded instruction to:

determine a first rotational value and a second rotational value;

perform a first rotate operation using the first rotational value, wherein the first rotate operation is to include a rotation of a 3 operand XOR result generated based on at least a portion of the first operand value, at least a portion of the second operand value, and at least a portion of the third operand value to generate a first rotated result;

perform a second rotate operation on the first rotated result using the second rotational value to generate a second rotated result; and

store the second rotated XOR result.

9. The method of claim 8 , further comprising:

committing a result of the executed instruction.

10. The method of claim 8 , wherein the first register is a 512-bit register that stores a 64-bit operand value.

11. The method of claim 8 , wherein the first register stores eight 64-bit words that separately represent operand values.

12. The method of claim 11 , wherein the executed instruction is performed on the operand values separately represented by all eight 64-bit words.

13. The method of claim 8 , wherein the first rotational value and the second rotational value are based on a single bit of a multi-bit representation of the fourth operand value.

14. The method of claim 8 , wherein the first, the second and the third registers are separately capable of storing 512-bits of data to store respective first, second and third operand values and the fourth operand value is an 8-bit integer immediate.

15. A non-transitory computer readable medium comprising instructions which, when executed by a processor, configure the processor to:

fetch an encoded instruction comprising at least one opcode, a first source identifier to identify a first register to store a first operand value, a second source identifier to identify a second register to store a second operand value, a third source identifier to identify a third register to store a third operand value, and a fourth operand value;

decode the encoded instruction to generate a decoded instruction; and

execute the decoded instruction to:

determine a first rotational value and a second rotational value;

perform a first rotate operation using the first rotational value, wherein the first rotate operation is to include a rotation of a 3 operand XOR result generated based on at least a portion of the first operand value, at least a portion of the second operand value, and at least a portion of the third operand value to generate a first rotated result;

perform a second rotate operation on the first rotated result using the second rotational value to generate a second rotated result; and

store the second rotated result.

16. The computer readable medium of claim 15 , comprising instructions to:

commit a result of the executed instruction.

17. The computer readable medium of claim 15 , wherein the first register is a 512-bit register that stores a 64-bit operand value.

18. The computer readable medium of claim 15 , wherein the first register stores eight 64-bit words that separately represent operand values.

19. The computer readable medium of claim 18 , the executed instruction is performed on the operand values separately represented by all eight 64-bit words.

20. The computer readable medium of claim 15 , wherein the first rotational value and the second rotational value are based on a single bit of a multi-bit representation of the fourth operand value.

21. The computer readable medium of claim 20 , wherein the first, the second and the third registers are separately capable of storing 512-bits of data to store respective first, second and third operand values and the fourth operand value is an 8-bit integer immediate.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2023
From: GHOSH, SANTOSH; DOBRAUNIG, CHRISTOPH; SASTRY, MANOJ
To: INTEL CORPORATION
Reel/Frame 062507/0038 →
Continuity (1)
Related Publication 20240211268A1 · Jun 27, 2024
References Cited (48)
US 5892696A · Kozu · 1999 [cited by examiner]
US 7103180B1 · McGregor, Jr. · 2006 [cited by examiner]
US 7599489B1 · Spracklen · 2009 [cited by examiner]
US 7684563B1 · Olson · 2010 [cited by examiner]
US 8838997B2 · Wolrich · 2014 [cited by examiner]
US 9027104B2 · Wolrich · 2015 [cited by examiner]
US 9251377B2 · Wolrich · 2016 [cited by examiner]
US 9495165B2 · Gopal · 2016 [cited by examiner]
US 9495166B2 · Gopal · 2016 [cited by examiner]
US 9501281B2 · Gopal · 2016 [cited by examiner]
US 9747105B2 · Gopal · 2017 [cited by examiner]
US 10684855B2 · Gopal · 2020 [cited by examiner]
US 11093213B1 · Lablans · 2021 [cited by examiner]
US 11188335B2 · Shemy · 2021 [cited by examiner]
US 11222127B2 · Ghosh · 2022 [cited by examiner]
US 11567772B2 · Shemy · 2023 [cited by examiner]
US 11681530B2 · Shemy · 2023 [cited by examiner]
US 11822901B2 · Saarinen · 2023 [cited by examiner]
US 20020032551A1 · Zakiya · 2002 [cited by examiner]
US 20100250966A1 · Olson · 2010 [cited by examiner]
US 20110153700A1 · Gopal · 2011 [cited by examiner]
US 20130132737A1 · Horsnell · 2013 [cited by examiner]
US 20140093069A1 · Wolrich · 2014 [cited by examiner]
US 20140095844A1 · Gopal et al. · 2014 [cited by applicant]
US 20140185793A1 · Wolrich · 2014 [cited by examiner]
US 20140189368A1 · Wolrich · 2014 [cited by examiner]
US 20140189369A1 · Wolrich · 2014 [cited by examiner]
US 20150089195A1 · Gopal · 2015 [cited by examiner]
US 20150089196A1 · Gopal · 2015 [cited by examiner]
US 20150089197A1 · Gopal · 2015 [cited by examiner]
US 20160139919A1 · Evans · 2016 [cited by examiner]
US 20170147340A1 · Yap et al. · 2017 [cited by applicant]
US 20170351519A1 · Gopal · 2017 [cited by examiner]
US 20200012495A1 · Aly · 2020 [cited by examiner]
US 20200104444A1 · Baratam et al. · 2020 [cited by applicant]
US 20200117811A1 · Ghosh · 2020 [cited by examiner]
US 20220066741A1 · Saarinen · 2022 [cited by examiner]
CA 2298055C · 2007 [cited by examiner]
WO 2013095648A1 · 2013 [cited by applicant]
English Abstract of Japanese Patent Application JP-2017134840-A, 2017 (Year: 2017). [cited by examiner]
‘SIMD Instruction Set Extensions for Keccak with Applications to SHA-3, Keyak and Ketje’ by Hemendra K. Rawat et al., HASP 2016. (Year: 2016). [cited by examiner]
‘Implementation of SIMD Instruction Set Extension for BLAKE2’ by Ganesh et al., ICCCNT 2019. (Year: 2019). [cited by examiner]
‘Efficient FPGA Implementation of the SHA-3 Hash Function’ by Magnus Sundal et al., 2017 IEEE Computer Society Annual Symposium on VLSI. (Year: 2017). [cited by examiner]
‘Diploma Thesis—Design of the HW accelerator of the Keccak hash function’ by Nikita Litvishko, 2020. (Year: 2020). [cited by examiner]
‘Maximizing the Potential of Custom RISC-V Vector Extensions for Speeding up SHA-3 Hash Functions’ by Huimin Li et al., 2023 Design, Automation & Test in Europe Conference. (Year: 2023). [cited by examiner]
Avanzi, Roberto, et al., “Algorithm Specifications And Supporting Documentation (version 3.02)”, Crystals—Kyber, Aug. 4, 2021, 43 pages. [cited by applicant]
FIPS PUB 202, “SHA-3 Standard: Permutation-Based Hash and Extendable-Output Functions”, Information Technology Laboratory, National Institute of Standards and Technology, Gaithersburg, MD 20899-8900, http://dx.doi.org/1… [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/145,776, Mailed Feb. 29, 2024, 9 pages. [cited by applicant]