IP Library Granted Patent US 12,700,989
Granted Patent B2
US 12,700,989 · App. 18/719,495 · Granted Aug 4, 2026

Cryptographic processor for ciphertext applications

Inventors: Shaveer Bajpeyi (Toronto, CA); Glenn Gulak (Etobicoke, CA)
Assignee: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
H04L9/008H04L9/0618
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,700,989
App. No.
18/719,495
Granted
Aug 4, 2026
Kind
B2
Abstract

Cryptographic processor chips, systems and associated methods are disclosed. In one embodiment, a cryptographic processor is disclosed. The cryptographic processor includes a first cryptographic processing module to perform a first logic operation. The first cryptographic processing module includes first input circuitry to receive ciphertext input symbols. A first pipeline stage performs a first operation on the ciphertext input symbols and generates a first stage output. On-chip memory temporarily stores the first stage output and feeds the first stage output to a second pipeline stage in a pipelined manner. The second pipeline stage is configured to perform a second operation on the first stage output in a pipelined manner with respect to the first pipeline stage.

Claims (52)

1 . A cryptographic processor, comprising:

a first cryptographic processing module to perform a first logic operation, the first cryptographic processing module including:

first input circuitry to receive ciphertext input symbols;

a first pipeline stage to perform a first operation on the ciphertext input symbols and to generate a first stage output;

storage circuitry to store constants for use by the first pipeline stage, the storage circuitry including an input interface configured with a bit-width to receive at least one entire ciphertext or key coefficient per cycle of a system clock, the storage circuitry configured to temporarily store the first stage output and to feed a second pipeline stage in a pipelined manner; and

wherein the second pipeline stage is configured to perform a second operation on the first stage output in a pipelined manner with respect to the first pipeline stage.

2 . The cryptographic processor of claim 1 , wherein:

the first input circuitry receives the ciphertext input symbols synchronous with an instruction clock signal; and

wherein a new set of input ciphertext symbols are presented to the first input circuitry each cycle of the instruction clock signal.

3 . The cryptographic processor of claim 1 , wherein:

the first cryptographic processing module is configured to perform ciphertext multiplication operations, ciphertext rotation operations, or ciphertext addition operations.

4 . The cryptographic processor of claim 1 , further comprising:

at least one Chinese Remainder Theorem (CRT) processing stage.

5 . The cryptographic processor of claim 4 , wherein:

the first cryptographic processing module is configured as a Residue Number System (RNS) architecture.

6 . The cryptographic processor of claim 5 , wherein:

the first cryptographic processing module includes multiple processing slices defining multiple processing channels, each channel to perform operations on signals 64-bits wide or less concurrently with the other processing channels.

7 . The cryptographic processor of claim 1 , wherein:

the first cryptographic processing module is configured as a Large Arithmetic Word Size (LAWS) architecture.

8 . The cryptographic processor of claim 7 , wherein:

the first cryptographic processing module includes a single processing slice defining a single processing channel to perform operations on signals that are more than 64-bits wide.

9 . The cryptographic processor of claim 1 , further comprising:

a second processing module to perform a second logic operation different than the first logic operation.

10 . The cryptographic processor of claim 1 , wherein the storage circuitry includes:

on-chip register circuitry configured to temporarily store the input of each stage.

11 . The cryptographic processor of claim 1 , wherein:

the first cryptographic processing module produces outputs with a pre-determined latency of execution cycles and constant throughput.

12 . The cryptographic processor of claim 1 , wherein:

the first pipeline stage comprises a stage of number theoretic transform (NTT) circuits to perform an NTT operation as the first operation, the stage of NTT circuits configured to exhibit a predetermined parallelism; and

wherein the second pipeline stage is configured to employ a number of inputs and outputs that match the predetermined parallelism of the stage of NTT circuits.

13 . A cryptographic processor, comprising:

a first cryptographic processing module including:

first input circuitry to receive ciphertext input symbols;

a number theoretic transform (NTT) stage to perform an NTT operation on received ciphertext input symbols and to generate an NTT stage output, the NTT stage configured to exhibit a predetermined parallelism;

a second circuit stage that receives the NTT stage output in a pipelined manner; and

wherein the second circuit stage is configured to employ a number of inputs and outputs that matches the predetermined parallelism of the NTT circuit.

14 . The cryptographic processor of claim 13 , wherein:

the first input circuitry receives the ciphertext input symbols synchronous with an instruction clock signal; and

wherein a new set of input ciphertext symbols are presented to the first input circuitry each cycle of the instruction clock signal.

15 . The cryptographic processor of claim 13 , wherein:

the first cryptographic processing module is configured to perform ciphertext addition operations or ciphertext multiplication operations or ciphertext rotation operations.

16 . The cryptographic processor of claim 13 , wherein:

the first cryptographic processing module is configured as a Residue Number System (RNS) architecture and includes multiple processing slices defining multiple processing channels, each channel to perform operations on signals 64-bits wide or less concurrently with the other processing channels.

17 . The cryptographic processor of claim 13 , wherein:

the first cryptographic processing module is configured as a Large Arithmetic Word Size (LAWS) architecture and includes a single processing slice defining a single processing channel to perform operations on signals that are more than 64-bits wide.

18 . The cryptographic processor of claim 13 , wherein:

the first cryptographic processing module produces outputs with a pre-determined latency of execution cycles and constant throughput.

19 . A method of operation in a cryptographic processor, the method comprising:

receiving ciphertext input symbols with first input circuitry;

performing a number theoretic transform (NTT) operation on the received ciphertext input symbols with an NTT stage and generating an NTT stage output, the NTT stage configured to exhibit a predetermined parallelism;

receiving the NTT stage output in a pipelined manner with a second pipeline stage; and

configuring the second pipeline stage to employ a number of inputs and outputs that matches the predetermined parallelism of the NTT circuit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2024
From: BAJPEYI, SHAVEER; GULAK, GLENN
To: GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO, THE
Reel/Frame 067871/0829 →
Continuity (3)
Continuation 18081078 · Dec 14, 2022
Provisional Application 63289783 · Dec 15, 2021
Related Publication 20250055672A1 · Feb 13, 2025
References Cited (109)
US 6804453B1 · Sasamoto · 2004 [cited by examiner]
US 7478225B1 · Brooks · 2009 [cited by examiner]
US 7953731B2 · Patel et al. · 2011 [cited by applicant]
US 8051128B2 · Seshasai · 2011 [cited by applicant]
US 8965943B2 · Mellott · 2015 [cited by examiner]
US 9135331B2 · Rosenthal et al. · 2015 [cited by applicant]
US 10298385B2 · Khedr et al. · 2019 [cited by applicant]
US 20020094083A1 · Bhattacharya et al. · 2002 [cited by applicant]
US 20020126705A1 · Gentieu · 2002 [cited by examiner]
US 20090202067A1 · Michaels et al. · 2009 [cited by applicant]
US 20110055190A1 · Alexander · 2011 [cited by applicant]
US 20110161319A1 · Rathod · 2011 [cited by applicant]
US 20110243320A1 · Halevi et al. · 2011 [cited by applicant]
US 20110270849A1 · Varma et al. · 2011 [cited by applicant]
US 20110307411A1 · Bolivar et al. · 2011 [cited by applicant]
US 20120110087A1 · Culver et al. · 2012 [cited by applicant]
US 20170155628A1 · Rohloff et al. · 2017 [cited by applicant]
US 20170293913A1 · Gulak et al. · 2017 [cited by applicant]
US 20180029495A1 · Rilling et al. · 2018 [cited by applicant]
US 20180294950A1 · Khedr et al. · 2018 [cited by applicant]
US 20200374103A1 · Cheon · 2020 [cited by examiner]
US 20220094518A1 · Ghosh et al. · 2022 [cited by applicant]
US 20230027423A1 · Rao · 2023 [cited by applicant]
US 20230053311A1 · Aharoni et al. · 2023 [cited by applicant]
US 20230163945A1 · Park · 2023 [cited by examiner]
International Search Report for PCT application No. PCT/CA2022/051827, CIPO, search completed: Jan. 18, 2023, mailed: Mar. 14, 2023. [cited by applicant]
International Search Report for PCT application No. PCT/IB2017/054919, CIPO, search completed: Apr. 5, 20183, mailed: Apr. 23, 2018. [cited by applicant]
Written Opinion of the International Searching Authority for PCT application No. PCT/CA2022/051827, CIPO, opinion completed: Jan. 31, 2023, mailed: Mar. 14, 2023. [cited by applicant]
Written Opinion of the International Searching Authority for PCT application No. PCT/IB2017/054919, CIPO, opinion completed: Apr. 5, 2018, mailed: Apr. 23, 2018. [cited by applicant]
“5 nm lithography process”, https://en.wikichip.org/wiki/5_nm_lithography_process (accessed Aug. 30, 2021). [cited by applicant]
“7 nm lithography process”, https://en.wikichip.org/wiki/7 nm_lithography_process (accessed Aug. 30, 2021). [cited by applicant]
“Barrett reduction”, Wikipedia, 2016. [cited by applicant]
“Data Protection in Virtual Environments”, Defense Advanced Research Projects Agency, Feb. 2020. [cited by applicant]
“Euclidean algorithm”, Wikipedia, 2016. [cited by applicant]
“Exponentiation by squaring”, Wikipedia, 2016. [cited by applicant]
“Fermat number”, Wikipedia, 2016. [cited by applicant]
“HBM3 Memory Subsystem”, Rambus, 2021. [cited by applicant]
“High Bandwidth Memory DRAM (HBM1, HBM2) JESD235C”, J. S. S. T. Association, 2020. [cited by applicant]
“Mersenne prime”, Wikipedia, 2016. [cited by applicant]
“Montgomery modular multiplication”, Wikipedia, 2016. [cited by applicant]
“NVIDIA Tesla P100”, Nvidia, Whitepaper 2016. [Online]. Available: https://images.nvidia.com/content/pdf/tesla/whitepaper/pascal-architecture-whitepaper.pdf. [cited by applicant]
“Snucrypto/HEAAN”, GitHub, Available: https://github.com/snucrypto/HEAAN. [cited by applicant]
“Solinas prime”, Wikipedia, 2016. [cited by applicant]
Aysu , et al., “Low-Cost and Area-Efficient FPGA Implementations of Lattice-Based Cryptography”, 2013. [cited by applicant]
Barrett, Paul , “Implementing the Rivest Shamir and Adleman Public Key Encryption Algorithm on a Standard Digital Signal Processor”, CRYPTO 1986, 1986: Springer, Berlin, Heidelberg, in Lecture Notes in Computer Science,… [cited by applicant]
Bos , et al., “Improved Security for a Ring-Based Fully Homomorphic Encryption Scheme”, 2013. [cited by applicant]
Cao , et al., “Accelerating Fully Homomorphic Encryption over the Integers with Super-size Hardware Multiplier and Modular Reduction”, 2013. [cited by applicant]
Cathebras, J , “Hardware Acceleration for Homomorphic Encryption”, Ph.D dissertation, Universite Paris-Saclay, Paris, 2018. [cited by applicant]
Chang, Jonathan , et al., “A 5nm 135Mb SRAM in EUV and High-Mobility-Channel FinFET Technology with Metal Coupling and Charge-Sharing Write-Assist Circuitry Schemes for High-Density and Low-VMIN Applications”, 2020 IEEE… [cited by applicant]
Chang, Jonathan , et al., “A 7nm 256Mb SRAM in high-k metal-gate FinFET technology with write-assist circuitry for low-VMIN applications”, 2017 IEEE International Solid-State Circuits Conference (ISSCC), 2017, pp. 206-2… [cited by applicant]
Cheon, Jung Hee , et al., “Homomorphic encryption for arithmetic of approximate numbers”, in Advances in Cryptology—ASIACRYPT 2017, T. Takagi and T. Peyrin Eds., (Lecture Notes in Computer Science. Berlin and Heidelberg… [cited by applicant]
Cooley, J, et al., “An algorithm for the machine calculation of complex Fourier series”, Mathematics of Computation, vol. 19, pp. 297-301, 1965. [cited by applicant]
Cousins, David Bruce , et al., “Designing an FPGA-Accelerated Homomorphic Encryption Co-Processor”, IEEE Transactions on Emerging Topics in Computing, vol. 5, No. 2, pp. 193-206, 2017, doi: 10.1109/tetc.2016.2619669. [cited by applicant]
Cousins , et al., “SIPHER: Scalable Implementation of Primitives For Homomorphic Encryption”, Nov. 2015. [cited by applicant]
Doroz , et al., “Accelerating Fully Homomorphic Encryption in Hardware”, 2015. [cited by applicant]
Doroz , et al., “Accelerating LTV Based Homomorphic Encryption in Reconfigurable Hardware”, 2015. [cited by applicant]
Doroz , et al., “Evaluating the Hardware Performance of a Million-bit Nultiplier”, 2013. [cited by applicant]
Du , et al., “A Family of Scalable Polynomial Multiplier Architectures for Lattice-Based Cryptography”, 2015. [cited by applicant]
Elliott, Duncan , et al., “Computational RAM: implementing processors in memory”, IEEE Design Test of Computers, vol. 16, No. 1, pp. 32-41, 1999, doi: 10.1109/54.748803. [cited by applicant]
Fu, Y , et al., “GPU Domain Specialization via Composable on-Package Architecture”, ArXiv, No. 2104.02188, 2021. [Online]. Available: https://arxiv.org/pdf/2104.02188.pdf. [cited by applicant]
Galvez, Mario Garrido , et al., “Pipelined Radix-2(k) Feedforward FFT Architectures”, IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 21, No. 1, pp. 23-32, 2013, doi: 10.1109/tvlsi.2011.2178275. [cited by applicant]
Gentleman, W. M. , et al., “Fast Fourier Transforms: for fun and profit”, Proceedings of the Nov. 7-10, 1966, fall joint computer conference, San Francisco, California, 1966: Association for Computing Machinery, pp. 563… [cited by applicant]
Gentry, C, et al., “A fully homomorphic encryption scheme”, Stanford University, 2009. [cited by applicant]
Gentry , et al., “Implementing Gentry's Fully-Homomorphic Encryption Scheme”, IBM Research, Feb. 4, 2011. [cited by applicant]
Hankerson, D. , et al., “Guide to Elliptic Curve Cryptography”, 1 ed. Springer-Verlag, New York, 2004. [cited by applicant]
Jouppi, Norman , et al., “A domain-specific supercomputer for training deep neural networks”, Communications of the ACM, vol. 63, No. 7, pp. 67-78, 2020, doi: 10.1145/3360307. [cited by applicant]
Jouppi, Norman , et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit”, ACM SIGARCH Computer Architecture News, vol. 45, No. 2, pp. 1-12, 2017, doi: 10.1145/3140659.3080246. [cited by applicant]
Jung, Won Kyung , et al., “Accelerating Fully Homomorphic Encryption Through Architecture-Centric Analysis and Optimization”, IEEE Acess, Jul. 12, 2021. [cited by applicant]
Khedr , et al., “SHIELD: Scalable Homomorphic Implementation of Encrypted Data-Classifiers”, IEEE Transactions on Computers, vol. 65, No. 9, pp. 2848-2858, Sep. 30, 2016. [cited by applicant]
Kim, Sangpyo , et al., “Accelerating Number Theoretic Transformations for Bootstrappable Homomorphic Encryption on GPUs”, 2020 IEEE International Symposium on Workload Characterization (IISWC), 2020: IEEE, pp. 264-275, … [cited by applicant]
Kim, Sunwoong , et al., “FPGA-based Accelerators of Fully Pipelined Modular Multipliers for Homomorphic Encryption”, 2019 International Conference on ReConFigurable Computing and FPGAs (ReConFig), 2019, pp. 1-8, doi: 10… [cited by applicant]
Kim, Sunwoong , et al., “Hardware Architecture of a Number Theoretic Transform for a Bootstrappable RNS-based Homomorphic Encryption Scheme”, 2020 IEEE 28th Annual International Symposium on Field-Programmable Custom Co… [cited by applicant]
Langhammer, Martin , et al., “Efficient FPGA Modular Multiplication Implementation”, The 2021 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, 2021, pp. 217-223, doi: 10.1145/3431920.3439306. [cited by applicant]
Langhammer, Martin , et al., “Folded Integer Multiplication for FPGAs”, 2021 ACM/SIGDA International Symposium on Field Programmable Gate Arrays (FPGA '21), Virtual Event, USA, Feb. 28-Mar. 2, 2021 2021: Association for… [cited by applicant]
Lindner , et al., “Better Key Sizes (and Attacks) for LWE-Based Encryption”, Nov. 30, 2010. [cited by applicant]
Lmran, M , et al., “An Experimental Study of Building Blocks of Lattice-Based NIST Post-Quantum Cryptographic Algorithms”, Electronics, vol. 9, No. 11, 1953, 2020, doi: 10.3390/electronics9111953. [cited by applicant]
Lopez-Alt , et al., “On-the-Fly Multiparty Computation on the Cloud via Multikey Fully Homomorphic Encryption”, 2012. [cited by applicant]
Montgomery , et al., “Modular Multiplication Without Trial Division”, Mathematics of Computation, vol. 44, No. 170, Apr. 1985, pp. 519-521. [cited by applicant]
Moore , et al., “Targeting FPGA DSP Slices for a Large Integer Multiplier for Integer Based FHE”, pp. 226-237, 2013. [cited by applicant]
Nejatollahi, Hamid , et al., “CryptoPIM: in-memory Acceleration for Lattice-based Cryptographic Hardware”, 2020 57th ACM/IEEE Design Automation Conference (DAC), San Francisco, CA, USA, 2020, pp. 1-6, doi: 10.1109/DAC18… [cited by applicant]
Ozturk, Erdinc , et al., “A Custom Accelerator for Homomorphic Encryption Applications”, IEEE Transactions on Computers, vol. 66, No. 1, pp. 3-16, Jan. 1, 2017, doi: 10.1109/TC.2016.2574340. [cited by applicant]
Ozturk , et al., “Accelerating Somewhat Homomorphic Evaluation using FPGAs”, 2015. [cited by applicant]
Poppelmann , et al., “Accelerating Homomorphic Evaluation on Reconfigurable Hardware”, 2015. [cited by applicant]
Poppelmann, T , et al., “High-Performance Ideal Lattice-Based Cryptography on 8-bit ATxmega Microcontrollers”, Extended Version, ed. Cryptology ePrint Archive, Report 2015/382, 2015. [cited by applicant]
Poppelmann , et al., “Towards Efficient Arithmetic for Lattice-Based Cryptography on Reconfigurable Hardware”, 2012. [cited by applicant]
Reagen, Brandon , et al., “Cheetah: Optimizations and Methods for Privacy Preserving Inference via Homomorphic Encryption”, CoRR, vol. abs/2006.00505, 2020. [Online]. Available: https://arxiv.org/abs/2006.00505. [cited by applicant]
Reis, Dayane , et al., “Computing-in-Memory for Performance and Energy-Efficient Homomorphic Encryption”, IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 28, No. 11, pp. 2300-2313, 2020, doi: 10.1… [cited by applicant]
Riazi, M. Sadegh , et al., “HEAX: an Architecture for Computing on Encrypted Data”, Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems, Lausanne, Switzerland, … [cited by applicant]
Robinson , “Mersenne and Fermat Numbers”, Proceedings of the American Mathematical Society, vol. 5, No. 5, pp. 842-846, Oct. 1954. [cited by applicant]
Rondeau, Tom , “Data Protection in Virtual Environments (DPRIVE)”, ed, 2020. [cited by applicant]
Roy , et al., “Compact Ring-LWE Cryptoprocessor”, 2014. [cited by applicant]
Roy, Sujoy Sinha , et al., “FPGA-Based High-Performance Parallel Architecture for Homomorphic Computing on Encrypted Data”, 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2019, pp. 3… [cited by applicant]
Roy , et al., “Modular hardware Architecture for Somewhat Homomorphic Function Evaluation”, International Workshop on Cryptographic Hardware and Embedded Systems, CHES 2015 pp. 164-184 Feb. 4, 2011. [cited by applicant]
Shilov , et al., “GDDR5X Standard Finalized by JEDEC: New Graphics Memory up to 14 Gbps”, Jan. 2016. [cited by applicant]
Solinas , “Generalized Mersenne Numbers”, 1999. [cited by applicant]
Turan, Furkan , et al., “HEAWS: an Accelerator for Homomorphic Encryption on the Amazon AWS FPGA”, IEEE Transactions on Computers, vol. 69, No. 8, pp. 1185-1196, 2020, doi: 10.1109/tc.2020.2988765. [cited by applicant]
Wang , et al., “Accelerating Fully Homomorphic Encryption on GPUs”, 2012. [cited by applicant]
Wang , “FPGA Implementation of a Large-Number Multiplier for Fully Homomorphic Encryption”, IEEE, 2013. [cited by applicant]
Weste, Neil H. E. , et al., “CMOS VLSI Design: a Circuits and Systems Perspective”, Fourth ed. Addison-Wesley, 2011. *** Submitted in 5 parts, using a total of 5 files ***. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/081,078, USPTO, dated: Jul. 7, 2023. [cited by applicant]
Office Action for U.S. Appl. No. 18/081,078, USPTO, notification date: Mar. 3, 2023. [cited by applicant]
Office Action for U.S. Appl. No. 18/081,078, USPTO, notification date: Mar. 27, 2025. [cited by applicant]
Supplementary European Search Report for European application No. 22905585.0, EPO, search completed: Oct. 13, 2025, communication dated: Oct. 27, 2025. [cited by applicant]
“IEEE Standard Specification for Public Key Cryptographic Techniques Based on Hard Problems over Lattices”, Mar. 2009. [cited by applicant]
Bajpeyi, Shaveer, “A Hardware Accelerator for Fully Homomorphic Encryption based Machine Learning Applications”, Nov. 1, 2021 (Nov. 1, 2021), XP093324513, H04L9/00 Retrieved from the Internet: URL:https://utoronto.schol… [cited by applicant]
Knuth, et al., “The Art of Computer Programming”, pp. xii-xiii, 1997. [cited by applicant]
Oppenheim, et al., “Discrete-Time Signal Processing”, pp. v-xiii, 2010. [cited by applicant]
Parhami, Behrooz, “Computer Arithmetic, Algorithms and Hardware Designs”, pp. xii-xiv, 2000. [cited by applicant]
Sunwoong, et al., “FPGA-based 1-15 Accelerators of Fully Pipelined Modular Multipliers for Homomorphic Encryption”, 2019 International Conference on Reconfigurable Computing and FPGAS (Reconfig), IEEE, H04L Dec. 9, 2019… [cited by applicant]