IP Library › Granted Patent US 11,907,854
Granted Patent B2
US 11,907,854 · App. 16/910,744 · Granted Feb 20, 2024

System and method for mimicking a neural network without access to the original training dataset or the target model

Inventor: Eli David (Tel Aviv, IL)
Assignee: Nano Dimension Technologies, Ltd.
G06N3/086G06N3/04G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,907,854
App. No.
16/910,744
Granted
Feb 20, 2024
Kind
B2
Abstract

A device, system, and method is provided to mimic a pre-trained target model without access to the pre-trained target model or its original training dataset. A set of random or semi-random input data may be sent to randomly probe the pre-trained target model at a remote device. A set of corresponding output data may be received from the remote device that is generated by applying the pre-trained target model to the set of random or semi-random input data. A random probe training dataset may be generated comprising the set of random or semi-random input data and corresponding output data generated by randomly probing the pre-trained target model. A new model may be trained with the random probe training dataset so that the new model generates substantially the same corresponding output data in response to said input data to mimic the pre-trained target model.

Claims (49)

1. A method to mimic a pre-trained target model at a device without access to the pre-trained target model or its original training dataset, the method comprising, at the device:

sending a set of random or semi-random input data to a remote device to randomly probe the pre-trained target model remotely by inputting the set of random or semi-random input data into the pre-trained target model;

receiving from the remote device a set of corresponding output data generated by applying the pre-trained target model to the set of random or semi-random input data;

generating a random probe training dataset comprising the set of random or semi-random input data and corresponding output data generated by randomly probing the pre-trained target model;

training a new model with the random probe training dataset so that the new model generates substantially the same corresponding output data in response to said input data to mimic the pre-trained target model; and

removing a desired correlation in the new model based on training data linking an input to an output, without accessing at least one of the input or output, by adding to the random probe training dataset a plurality of random correlations to the output or input, respectively, to weaken or eliminate the desired correlation in the new model between the input and output diluted by the plurality of random correlations.

2. The method of claim 1 comprising adding new data to the random probe training dataset to incorporate new knowledge not present in the pre-trained target model.

3. The method of claim 1 comprising defining data to be omitted from the random probe training dataset to eliminate a category or class present in the pre-trained target model.

4. The method of claim 1 comprising re-training the new model using the random probe training dataset to mimic re-training the target pre-trained model.

5. The method of claim 4 comprising sparsifying the new model to mimic the pre-trained target model to generate a sparse new model.

6. The method of claim 4 comprising evolving the new model by applying evolutionary algorithms to mimic the pre-trained target model.

7. The method of claim 1 comprising generating or re-training the new model after all copies of the original training dataset are deleted at the remote device.

8. The method of claim 1 comprising training the new model over multiple epochs with a different random probe training dataset in each of the multiple epochs.

9. The method of claim 1 comprising setting the structure of the new model to be simpler than the structure of the pre-trained target model.

10. The method of claim 1 , wherein the models are neural networks, and comprising setting the new model to have a number of neurons, synapses, or layers, to be less than that of the pre-trained target model.

11. The method of claim 1 , wherein the models are neural networks, and comprising generating synthetic training samples by backpropagating error at an output layer backwards through the target model neural network to adjust training samples in the input layer to reduce the error.

12. The method of claim 1 , wherein the models are neural networks each comprising a plurality of layers, and comprising training the new model layer-by-layer in a plurality of sequential stages, each stage training a respective sequential layer of the new model neural network.

13. The method of claim 1 comprising measuring statistical properties of one or more sample inputs of the same type as the original training dataset or an accessible subset thereof; and semi-randomly selecting the set of input data according to those statistical properties.

14. The method of claim 1 comprising

requesting the remote device perform an initial probe of the pre-trained target model with multiple input samples of each of a plurality of data types or distributions that vary according to gaussian or uniform distributions from each other in an input space; and

selecting the data type or distribution for the random probe training dataset that has corresponding multiple target model outputs that have the smallest difference in the output space.

15. The method of claim 1 comprising, after training the new model, executing the new model in a run-time phase by inputting new data into the new model and generating corresponding data output by the new model.

16. A system for performing machine learning to generate a new model to mimic a pre-trained target model without obtaining the pre-trained target model or its original training dataset, the system comprising:

one or more processors configured to:

send a set of random or semi-random input data to a remote device to randomly probe the pre-trained target model remotely by inputting the set of random or semi-random input data into the pre-trained target model,

receive from the remote device a set of corresponding output data generated by applying the pre-trained target model to the set of random or semi-random input data,

generate a random probe training dataset comprising the set of random or semi-random input data and corresponding output data generated by randomly probing the pre-trained target model,

train a new model with the random probe training dataset so that the new model generates substantially the same corresponding output data in response to said input data to mimic the pre-trained target model, and

remove a desired correlation in the new model based on training data linking an input to an output, without accessing at least one of the input or output, by adding to the random probe training dataset a plurality of random correlations to the output or input, respectively, to weaken or eliminate the desired correlation in the new model between the input and output diluted by the plurality of random correlations.

17. The system of claim 16 comprising one or more memories to store one or more samples of the random probe training dataset.

18. The system of claim 17 , wherein the one or more memories are temporary memories that store samples of the random probe training dataset on-the-fly and delete the samples on-the-fly after the samples are used to train the new model.

19. The system of claim 16 , wherein the one or more processors are configured to add new data to the random probe training dataset to incorporate new knowledge not present in the pre-trained target model.

20. The system of claim 16 , wherein the one or more processors are configured to define data to be omitted from the random probe training dataset to eliminate a category or class present in the pre-trained target model.

21. The system of claim 16 , wherein the one or more processors are configured to re-train the new model using the random probe training dataset to mimic re-training the target pre-trained model.

22. The system of claim 16 , wherein the one or more processors are configured to generate or re-train the new model after all copies of the original training dataset are deleted at the remote device.

23. The system of claim 16 , wherein the one or more processors are configured to request the remote device perform an initial probe of the pre-trained target model with multiple input samples of each of a plurality of data types or distributions that vary according to gaussian or uniform distributions from each other in an input space; and

select the data type or distribution for the random probe training dataset that has corresponding multiple target model outputs that have the smallest difference in the output space.

24. The system of claim 16 , wherein the one or more processors are configured to train the new model over multiple epochs with a different random probe training dataset in each of the multiple epochs.

25. The system of claim 16 , wherein the one or more processors are configured to set the structure of the new model to be simpler than the structure of the pre-trained target model.

26. The system of claim 16 , wherein the models are neural networks, and the one or more processors are configured to set the new model to have a number of neurons, synapses, or layers, to be less than that of the pre-trained target model.

27. The system of claim 16 , wherein the models are neural networks, and the one or more processors are configured to generate synthetic training samples by backpropagating error at an output layer backwards through the target model neural network to adjust training samples in the input layer to reduce the error.

28. The system of claim 16 , wherein the models are neural networks each comprising a plurality of layers, and the one or more processors are configured to train the new model layer-by-layer in a plurality of sequential stages, each stage training a respective sequential layer of the new model neural network.

29. The system of claim 16 , wherein the one or more processors are configured to, after training the new model, execute the new model in a run-time phase by inputting new data into the new model and generating corresponding data output by the new model.

30. A non-transitory computer-readable medium comprising instructions which, when implemented in one or more processors in a computing device, cause the one or more processors to mimic a pre-trained target model at a device without access to the pre-trained target model or its original training dataset by:

sending a set of random or semi-random input data to a remote device to randomly probe the pre-trained target model remotely by inputting the set of random or semi-random input data into the pre-trained target model;

receiving from the remote device a set of corresponding output data generated by applying the pre-trained target model to the set of random or semi-random input data;

generating a random probe training dataset comprising the set of random or semi-random input data and corresponding output data generated by randomly probing the pre-trained target model;

training a new model with the random probe training dataset so that the new model generates substantially the same corresponding output data in response to said input data to mimic the pre-trained target model; and

removing a desired correlation in the new model based on training data linking an input to an output, without accessing at least one of the input or output, by adding to the random probe training dataset a plurality of random correlations to the output or input, respectively, to weaken or eliminate the desired correlation between the input and output diluted by the plurality of random correlations.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2023
From: DEEPCUBE LTD.
To: NANO DIMENSION TECHNOLOGIES, LTD.
Reel/Frame 063002/0579 →
CHANGE OF NAME Recorded Oct 31, 2022
From: DEEPCUBE LTD.
To: NANO DIMENSION TECHNOLOGIES, LTD.
Reel/Frame 061594/0895 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 25, 2020
From: DAVID, ELI
To: DEEPCUBE LTD.
Reel/Frame 053034/0538 →
Continuity (6)
Continuation In Part PCTIL2018051345 · Dec 10, 2018
Continuation 16211994 · Dec 6, 2018
Continuation 16910744
Continuation In Part 16211994 · Dec 6, 2018
Provisional Application 62679115 · Jun 1, 2018
Related Publication 20200320400A1 · Oct 8, 2020