IP Library › Granted Patent US 11,769,059
Granted Patent B2
US 11,769,059 · App. 18/145,004 · Granted Sep 26, 2023

Systems and methods for distributed training of deep learning models

Inventor: David Moloney (Dublin, IE)
Assignee: Movidius Limited
G06N3/084G06F18/214G06N3/045G06N3/08G06V10/454G06V10/764G06V10/82G06V10/95G06V10/96H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,769,059
App. No.
18/145,004
Granted
Sep 26, 2023
Kind
B2
Abstract

Systems and methods for distributed training of deep learning models are disclosed. An example local device to train deep learning models includes a reference generator to label input data received at the local device to generate training data, a trainer to train a local deep learning model and to transmit the local deep learning model to a server that is to receive a plurality of local deep learning models from a plurality of local devices, the server to determine a set of weights for a global deep learning model, and an updater to update the local deep learning model based on the set of weights received from the server.

Claims (41)

1. A non-transitory computer readable medium comprising instructions that, when executed, cause a server to at least:

obtain a first deep learning model from a first device;

obtain a second deep learning model from a second device;

determine that first weights of the first deep learning model are associated with a first distribution of weights from a first plurality of devices;

determine that second weights of the second deep learning model are outliers from the first distribution and are associated with a second distribution of weights, the second distribution of weights associated with a second plurality of devices;

aggregate the first weights with weights of the first distribution of weights to determine a first aggregated deep learning model;

aggregate the second weights with weights of the second distribution of weights to determine a second aggregated deep learning model; and

distribute the first aggregated deep learning model to the first device and the first plurality of devices; and

distribute the second aggregated deep learning model to the second device and the second plurality of devices.

2. The non-transitory computer readable medium of claim 1 , wherein the instructions, when executed, cause the server to train a global deep learning model based on an aggregation of the first distribution of weights and the second distribution of weights.

3. The non-transitory computer readable medium of claim 2 , wherein the server does not receive input data utilized by the first device to generate the first deep learning model.

4. The non-transitory computer readable medium of claim 2 , wherein the instructions, when executed, cause the server to aggregate the weights by averaging weights of the first distribution of weights and the second distribution of weights.

5. The non-transitory computer readable medium of claim 1 , wherein the first deep learning model is trained at the first device.

6. The non-transitory computer readable medium of claim 1 , wherein the second deep learning model is trained at the second device.

7. The non-transitory computer readable medium of claim 1 , wherein the first deep learning model is retrieved local to the first device and the second deep learning model is retrieved local to the second device.

8. A system to train deep learning models, the system comprising:

a first device to:

label input data received at the first device to generate training data;

train a first deep learning model; and

transmit the first deep learning model to an external device; and

update the first deep learning model based on a first aggregated deep learning model received from the external device; and

a server to:

obtain the first deep learning model from the first device;

obtain a second deep learning model from a second device;

determine that first weights of the first deep learning model are associated with a first distribution of weights from a first plurality of devices;

determine that second weights of the second deep learning model are outliers from the first distribution and are associated with a second distribution of weights, the second distribution of weights associated with a second plurality of devices;

aggregate the first weights with weights of the first distribution of weights to determine the first aggregated deep learning model;

aggregate the second weights with weights of the second distribution of weights to determine a second aggregated deep learning model; and

distribute the first aggregated deep learning model to the first device and the first plurality of devices; and

distribute the second aggregated deep learning model to the second device and the second plurality of devices.

9. The system of claim 8 , wherein the input data is not transmitted to the external device.

10. The system of claim 9 , wherein the first device is to sample the input data by selecting a pseudo-random portion of the input data.

11. The system of claim 9 , wherein the first device is to sample of the input data by down sampling the input data to reduce a data size of the input data.

12. The system of claim 8 , wherein the first device is to determine a difference between the label and an output of the first deep learning model.

13. The system of claim 12 , wherein the first device is to sample an output of the first deep learning model prior to the devices determining the difference.

14. The system of claim 13 , wherein the first device is to adjust the first deep learning model based on the difference.

15. The system of claim 8 , wherein the server is to train a global deep learning model based on an aggregation of the first distribution of weights and the second distribution of weights.

16. The system of claim 8 , wherein the server does not receive input data utilized by the first device to generate the first deep learning model.

17. The system of claim 8 , wherein the server is to aggregate the weights by averaging weights of the first distribution of weights and the second distribution of weights.

18. The system of claim 8 , wherein the second deep learning model is trained at the second device.

19. The system of claim 8 , wherein the first deep learning model is retrieved local to the first device and the second deep learning model is retrieved local to the second device.

Continuity (3)
Continuation 16326361
Provisional Application 62377094 · Aug 19, 2016
Related Publication 20230127542A1 · Apr 27, 2023
Cited By (1)
US 12,190,247