System and method for accelerating fully homomorphic encryption computations
A method and system for fully homomorphic-encryption (FHE) inference is presented. The method includes receiving, by a computing system, an encrypted request comprising a plurality of ciphertexts; loading, in response to the encrypted request, auxiliary data onto a first set of accelerators and a second set of accelerators; loading a first set of the plurality of ciphertexts onto the second set of accelerators; refreshing the first set of ciphertexts by executing, on the first set of accelerators, a bootstrapping operation that uses the auxiliary data; processing the first set of ciphertexts and a second set of the plurality of ciphertexts; and generating a response ciphertext based on the processing of the first and second sets of ciphertexts.
1 . A method for fully homomorphic-encryption (FHE) inference, the method comprising:
receiving, by a computing system, an encrypted request comprising a plurality of ciphertexts;
loading, in response to the encrypted request, auxiliary data onto a first set of accelerators and a second set of accelerators, wherein, when two or more users share an identical fully-homomorphic-encryption key set, one or more accelerators are assigned as dedicated bootstrapping accelerators for the key set and stores the auxiliary data in on-chip memory for the duration of a session;
loading a first set of the plurality of ciphertexts onto the second set of accelerators;
refreshing the first set of ciphertexts by executing, on the first set of accelerators, a bootstrapping operation that uses the auxiliary data;
processing the first set of ciphertexts and a second set of the plurality of ciphertexts; and
generating a response ciphertext based on the processing of the first and second sets of ciphertexts.
2 . The method of claim 1 , wherein the auxiliary data is session-specific auxiliary data that is loaded into on-chip memory of the first set of accelerators for a duration of a session.
3 . The method of claim 1 , wherein the processing of the first and the second set of ciphertexts comprises performing one or more non-bootstrapping homomorphic operations selected from at least one of vector-matrix multiplication, slot rotation, rescaling, and point-wise activation.
4 . The method of claim 1 , further comprising, prior to the loading of the first set of ciphertexts, determining a batch size at compile time, the batch size defining a number of ciphertexts processed in parallel by the second set of accelerators.
5 . The method of claim 1 , further comprising: transferring plaintext-weight blocks into device memory by direct-memory access while the second set of accelerators is processing ciphertexts.
6 . The method of claim 1 , wherein refreshing the first set of ciphertexts is performed entirely on-chip by the first set of accelerators with no, or reduced access to external memory.
7 . The method of claim 1 , wherein the encrypted request includes prompt-phase ciphertexts, and the first set of accelerators remains dedicated to the encrypted request until a first response ciphertext is generated.
8 . The method of claim 1 , further comprising: storing, in a double-buffer memory associated with each accelerator of the second set, a first plaintext-weight block used for computation while concurrently loading a second plaintext-weight block into an idle portion of the double-buffer memory.
9 . The method of claim 1 , further comprising, after the response ciphertext is generated, releasing at least one accelerator of the first set to process a different encrypted request.
10 . The method of claim 1 , further comprising, during a generation phase corresponding to the encrypted request, assigning a dedicated bootstrapping accelerator to a session that uses a distinct fully-homomorphic-encryption key set, wherein the dedicated bootstrapping accelerator stores the session's auxiliary data in on-chip memory for the duration of the session.
11 . The method of claim 1 , wherein each ciphertext of the first set and the second set encrypts a respective one or more tokens generated by a multi-token large-language-model.
12 . The method of claim 1 , wherein the plurality of ciphertexts represents activations of a convolutional neural network, and the first set of accelerators refreshes the ciphertexts during processing of one or more convolutional layers.
13 . A non-transitory computer-readable medium storing a set of instructions for fully homomorphic-encryption (FHE) inference, the set of instructions comprising:
one or more instructions that, when executed by one or more processing circuitries of a device, cause the device to:
receive, by a computing system, an encrypted request comprising a plurality of ciphertexts;
load, in response to the encrypted request, auxiliary data onto a first set of accelerators and a second set of accelerators, wherein, when two or more users share an identical fully-homomorphic-encryption key set, one or more accelerators are assigned as dedicated bootstrapping accelerators for the key set and stores the auxiliary data in on-chip memory for the duration of a session;
load a first set of the plurality of ciphertexts onto the second set of accelerators;
refresh the first set of ciphertexts by executing, on the first set of accelerators, a bootstrapping operation that uses the auxiliary data;
process the first set of ciphertexts and a second set of the plurality of ciphertexts; and
generate a response ciphertext based on the processing of the first and second sets of ciphertexts.
14 . A system for fully homomorphic-encryption (FHE) inference comprising:
a processing circuitry;
a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to:
receive, by a computing system, an encrypted request comprising a plurality of ciphertexts;
load, in response to the encrypted request, auxiliary data onto a first set of accelerators and a second set of accelerators, wherein, when two or more users share an identical fully-homomorphic-encryption key set, one or more accelerators are assigned as dedicated bootstrapping accelerators for the key set and stores the auxiliary data in on-chip memory for the duration of a session;
load a first set of the plurality of ciphertexts onto the second set of accelerators;
refresh the first set of ciphertexts by executing, on the first set of accelerators, a bootstrapping operation that uses the auxiliary data;
process the first set of ciphertexts and a second set of the plurality of ciphertexts; and
generate a response ciphertext based on the processing of the first and second sets of ciphertexts.
15 . The system of claim 14 , wherein the auxiliary data is session-specific auxiliary data that is loaded into on-chip memory of the first set of accelerators for a duration of a session.
16 . The system of claim 14 , wherein the processing circuitry, when processing the first and the second set of ciphertexts, is configured to perform one or more non-bootstrapping homomorphic operations selected from at least one of vectormatrix multiplication, slot rotation, rescaling, and point-wise activation.
17 . The system of claim 14 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:
determine a batch size at compile time prior to the loading of the first set of ciphertexts, wherein the batch size is defined by a number of ciphertexts processed in parallel by the second set of accelerators.
18 . The system of claim 14 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:
transfer plaintext-weight blocks into device memory by direct-memory access while the second set of accelerators is processing ciphertexts.
19 . The system of claim 14 , wherein refreshing the first set of ciphertexts is performed entirely on-chip by the first set of accelerators with no, or reduced access to external memory.
20 . The system of claim 14 , wherein the encrypted request includes prompt-phase ciphertexts, and the first set of accelerators remains dedicated to the encrypted request until a first response ciphertext is generated.
21 . The system of claim 14 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:
store, in a double-buffer memory associated with each accelerator of the second set, a first plaintext-weight block used for computation while concurrently loading a second plaintext-weight block into an idle portion of the double-buffer memory.
22 . The system of claim 14 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:
release at least one accelerator of the first set to process a different encrypted request after the response ciphertext is generated.
23 . The system of claim 14 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:
assign a dedicated bootstrapping accelerator to a session that uses a distinct fully-homomorphic-encryption key set during a generation phase corresponding to the encrypted request, wherein the dedicated bootstrapping accelerator stores the session's auxiliary data in on-chip memory for the duration of the session.
24 . The system of claim 14 , wherein each ciphertext of the first set and the second set encrypts a respective one or more tokens generated by a multi-token large-language-model.
25 . The system of claim 14 , wherein the plurality of ciphertexts represents activations of a convolutional neural network, and the first set of accelerators refreshes the ciphertexts during processing of one or more convolutional layers.