IP Library › Granted Patent US 11,573,828
Granted Patent B2
US 11,573,828 · App. 16/816,715 · Granted Feb 7, 2023

Efficient and scalable enclave protection for machine learning programs

Inventors: Chung Hwan Kim (Franklin Park, NJ); Junghwan Rhee (Princeton, NJ); Xiao Yu (Princeton, NJ); Luan Tang (Pennington, NJ); Haifeng Chen (West Windsor, NJ); Kyungtae Kim (West Lafayette, IN)
G06F9/5016G06F11/3433G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,573,828
App. No.
16/816,715
Granted
Feb 7, 2023
Kind
B2
Abstract

A computer-implemented method for efficient and scalable enclave protection for machine learning (ML) programs includes tailoring at least one ML program to generate at least one tailored ML program for execution within at least one enclave, and executing the at least one tailored ML program within the at least one enclave.

Claims (61)

1. A computer-implemented method for efficient and scalable enclave protection for machine learning (ML) programs, comprising:

tailoring at least one ML program to generate at least one tailored ML program for execution within at least one enclave, including:

allocating a shared memory for computing a plurality of layers of a neural network, the shared memory reducing total memory usage during the computation of the plurality of layers;

loading model parameter data for each of the plurality of layers onto the shared memory on-demand;

addressing memory usage dependencies of the layers using inter-layer dependency resolution; and

partitioning computation of any high memory usage layers into multiple sessions using intra-layer computation partitioning, the high memory usage layers including layers having a memory usage higher than a threshold memory usage for the at least one enclave; and

executing the at least one tailored ML program within the at least one enclave, the executing the at least one tailored ML program within the at least one enclave comprising:

computing the plurality of layers in a sequence using the shared memory, including loading weights for each layer to be executed on-demand;

adjusting an allocation of extra shared memory for additional dependencies; and

executing the at least one tailored ML program using the shared memory based on the computation, adjustment and any partitioned computations.

2. The method of claim 1 , further comprising profiling memory usage to generate memory profile results for tailoring the at least one ML program, including:

profiling I/O memory by analyzing how the at least one ML program uses I/O memory buffers; and

profiling weight memory by analyzing how the at least one ML program uses weight memory buffers.

3. The method of claim 1 , wherein the model parameter data is loaded from at least one model file associated with the at least one ML program.

4. The method of claim 1 , further comprising scheduling the at least one enclave into at least one processor such that memory usage does not exceed a memory budget.

5. The method of claim 4 , wherein scheduling the at least one enclave further includes:

receiving a new ML program execution request;

determining that launching the at least one enclave will cause page swapping using a pre-calculated memory requirement; and

scheduling the at least one enclave into the at least one processor by waiting until enough memory space becomes available for the at least one enclave.

6. The method of claim 1 , further comprising terminating the at least one enclave.

7. A computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method for efficient and scalable enclave protection for machine learning (ML) programs, the method performed by the computer comprising:

tailoring at least one ML program to generate at least one tailored ML program for execution within at least one enclave, including:

allocating a shared memory for computing a plurality of layers of a neural network, the shared memory reducing total memory usage during the computation of the plurality of layers;

loading model parameter data for each of the plurality of layers onto the shared memory on-demand;

addressing memory usage dependencies of the layers using inter-layer dependency resolution; and

partitioning computation of any high memory usage layers into multiple sessions using intra-layer computation partitioning, the high memory usage layers including layers having a memory usage higher than a threshold memory usage for the at least one enclave; and

executing the at least one tailored ML program within the at least one enclave, the executing the at least one tailored ML program within the at least one enclave comprising:

computing the plurality of layers in a sequence using the shared memory, including loading weights for each layer to be executed on-demand;

adjusting an allocation of extra shared memory for additional dependencies; and

executing the at least one tailored ML program using the shared memory based on the computation, adjustment and any partitioned computations.

8. The computer program product of claim 7 , wherein the method further includes profiling memory usage to generate memory profile results for tailoring the at least one ML program, including:

profiling I/O memory by analyzing how the at least one ML program uses I/O memory buffers; and

profiling weight memory by analyzing how the at least one ML program uses weight memory buffers.

9. The computer program product of claim 7 , wherein the model parameter data is loaded from at least one model file associated with the at least one ML program.

10. The computer program product of claim 7 , wherein the method further includes scheduling the at least one enclave into at least one processor such that memory usage does not exceed a memory budget.

11. The computer program product of claim 10 , wherein scheduling the at least one enclave further includes:

receiving a new ML program execution request;

determining that launching the at least one enclave will cause page swapping using a pre-calculated memory requirement; and

scheduling the at least one enclave into the at least one processor by waiting until enough memory space becomes available for the at least one enclave.

12. The computer program product of claim 7 , wherein the method further includes terminating the at least one enclave.

13. A system for efficient and scalable enclave protection for machine learning (ML) programs, comprising:

a memory device having program code stored thereon; and

at least one processor device operatively coupled to a memory device and configured to execute program code stored on the memory device to:

tailor at least one ML program to generate at least one tailored ML program for execution within at least one enclave by:

allocating a shared memory for computing a plurality of layers of a neural network, the shared memory reducing total memory usage during the computation of the plurality of layers;

loading model parameter data for each of the plurality of layers onto the shared memory on-demand;

addressing memory usage dependencies of the layers using inter-layer dependency resolution; and

partitioning computation of any high memory usage layers into multiple sessions using intra-layer computation partitioning, the high memory usage layers including layers having a memory usage higher than a threshold memory usage for the at least one enclave; and

execute the at least one tailored ML program within the at least one enclave, the executing the at least one tailored ML program within the at least one enclave comprising:

computing the plurality of layers in a sequence using the shared memory, including loading weights for each layer to be executed on-demand;

adjusting an allocation of extra shared memory for additional dependencies; and

executing the at least one tailored ML program using the shared memory based on the computation, adjustment and any partitioned computations.

14. The system of claim 13 , wherein the at least one processor device is further configured to execute program code stored on the memory device to profile memory usage to generate memory profile results for tailoring the at least one ML program by:

profiling I/O memory by analyzing how the at least one ML program uses I/O memory buffers; and

profiling weight memory by analyzing how the at least one ML program uses weight memory buffers.

15. The system of claim 13 , wherein the model parameter data is loaded from at least one model file associated with the at least one ML program.

16. The system of claim 13 , wherein the at least one processor device is further configured to execute program code stored on the memory device to schedule the at least one enclave into at least one processor such that memory usage does not exceed a memory budget by:

receiving a new ML program execution request;

determining that launching the at least one enclave will cause page swapping using a pre-calculated memory requirement; and

scheduling the at least one enclave into the at least one processor by waiting until enough memory space becomes available for the at least one enclave.

17. The system of claim 13 , wherein the at least one processor device is further configured to execute program code stored on the memory device to terminate the at least one enclave.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2022
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 062153/0316 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 12, 2020
From: KIM, CHUNG HWAN; RHEE, JUNGHWAN; YU, XIAO; TANG, LUAN; CHEN, HAIFENG; KIM, KYUNGTAE
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 052097/0719 →
Continuity (3)
Provisional Application 62900686 · Sep 16, 2019
Provisional Application 62927724 · Oct 30, 2019
Related Publication 20210081122A1 · Mar 18, 2021
Cited By (1)
US 12,216,758