IP Library › Granted Patent US 11,410,024
Granted Patent B2
US 11,410,024 · App. 15/581,152 · Granted Aug 9, 2022

Tool for facilitating efficiency in machine learning

Inventors: Rajkishore Barik (Santa Clara, CA); Brian T. Lewis (Palo Alto, CA); Murali Sundaresan (Sunnyvale, CA); Jeffrey Jackson (Newbery, OR); Feng Chen (Shanghai, CN); Xiaoming Chen (Shanghai, CN); Mike Macpherson (Portland, OR)
Assignee: INTEL CORPORATION
G06N3/063G06F9/46G06N3/0445G06N3/0454G06N3/084G06N5/003
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,410,024
App. No.
15/581,152
Filed
Apr 28, 2017
Granted
Aug 9, 2022
Kind
B2
Art Unit
2125
USPC
706/23
Abstract

A mechanism is described for facilitating smart distribution of resources for deep learning autonomous machines. A method of embodiments, as described herein, includes detecting one or more sets of data from one or more sources over one or more networks, and introducing a library to a neural network application to determine optimal point at which to apply frequency scaling without degrading performance of the neural network application at a computing device.

Claims (48)

1. An apparatus comprising:

a graphics processor to:

detect one or more sets of data from one or more sources over one or more networks;

cause a neural network application to implement a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application;

implement, using the neural network application, the distributed gradient synchronization using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure;

determine, using the machine learning primitives of the library as implemented by the neural network application, a point to apply frequency scaling in the graphics processor that does not degrade performance of the neural network application, the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization implemented via the tree structure; and

determine, using the library as implemented by the neural network application, a core frequency of the frequency scaling applied at the point, wherein the library is to account for skew characteristics associated with the distributed gradient synchronization to decide the core frequency.

2. The apparatus of claim 1 , wherein the skew characteristics comprise at least the skew pattern.

3. The apparatus of claim 1 , wherein the graphics processor is further operable to introduce a sparse matrix representation for weights to overlap communication and computation across the one or more nodes associated with the neural network application to reduce communication costs.

4. The apparatus of claim 1 , wherein the graphics processor is further to automatically analyze failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.

5. The apparatus of claim 4 , wherein the graphics processor is further to provide one or more of successful execution information obtained from successful execution of programs and failed execution information obtained from failed execution of programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.

6. The apparatus of claim 1 , wherein the graphics processor is further to perform local error propagation by computing high precision and low precision for local weights and compute local errors at each of the one or more nodes, wherein performing the local error propagation further comprises facilitating weight synchronization across the one or more nodes to track the local errors for accuracy and reduced communication.

7. The apparatus of claim 1 , wherein the apparatus comprises an autonomous machine comprising one or more of a vehicle, a device, or an equipment, wherein the autonomous machine comprises one or more processors including the graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.

8. A method comprising:

detecting, by a graphics processor, one or more sets of data from one or more sources over one or more networks;

causing, by the graphics processor, a neural network application to implement a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application;

implementing, using the neural network application, the distributed gradient synchronization using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure;

determining, using the machine learning primitives of the library as implemented by the neural network application, a point to apply frequency scaling in the graphics processor that does not degrade performance of the neural network application, the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization implemented via the tree structure; and

determining, using the library as implemented by the neural network application, a core frequency of the frequency scaling applied at the point, wherein the library is to account for skew characteristics associated with the distributed gradient synchronization to decide the core frequency.

9. The method of claim 8 , wherein the skew characteristics comprise at least the skew pattern.

10. The method of claim 8 , further comprising introducing sparse matrix representation for weights to overlap communication and computation across the one or more nodes associated with the neural network application to reduce communication costs.

11. The method of claim 8 , further comprising automatically analyzing failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.

12. The method of claim 11 , further comprising providing one or more of successful execution information obtained from successful execution of programs and failed execution information obtained from failed execution of programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.

13. The method of claim 8 , further comprising performing local error propagation by computing high precision and low precision for local weights and compute local errors at each of the one or more nodes, wherein performing the local error propagation further comprises facilitating weight synchronization across the one or more nodes to track the local errors for accuracy and reduced communication.

14. The method of claim 8 , wherein the graphics processor is part of an autonomous machine comprising one or more of a vehicle, a device, or an equipment, wherein the autonomous machine comprises one or more processors including the graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.

15. At least one non-transitory machine-readable medium comprising instructions that when executed by a local computing device, cause the local computing device to perform operations comprising:

detecting, by a graphics processor of the local computing device, one or more sets of data from one or more sources over one or more networks;

causing, by the graphics processor, a neural network application to implement a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application;

implementing, using the neural network application, the distributed gradient synchronization using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure;

determining, using the machine learning primitives of the library as implemented by the neural network application, a point to apply frequency scaling in the graphics processor that does not degrade performance of the neural network application, the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization implemented via the tree structure; and

determining, using the library as implemented by the neural network application, a core frequency of the frequency scaling applied at the point, wherein the library is to account for skew characteristics associated with the distributed gradient synchronization to decide the core frequency.

16. The non-transitory machine-readable medium of claim 15 , wherein the skew characteristics comprise at least the skew pattern.

17. The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise introducing sparse matrix representation for weights to overlap communication and computation across the one or more nodes associated with neural network application to reduce communication costs.

18. The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise automatically analyzing failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.

19. The non-transitory machine-readable medium of claim 18 , wherein the operations further comprise providing one or more of successful execution information obtained from successful execution of programs and failed execution information obtained from failed execution of programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.

20. The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise performing local error propagation by computing high precision and low precision for local weights and compute local errors at each of the one or more nodes, wherein performing the local error propagation further comprises facilitating weight synchronization across the one or more nodes to track the local errors for accuracy and reduced communication, wherein the computing device comprises an autonomous machine comprising one or more of a vehicle, a device, or an equipment, wherein the autonomous machine comprises one or more processors including the graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.

21. A system comprising:

a memory; and

a graphics processor communicably coupled to the memory, the graphics processor to:

detect one or more sets of data from one or more sources over one or more networks;

cause a neural network application to implement a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application;

implement, using the neural network application, the distributed gradient synchronization using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure;

determine, using the machine learning primitives of the library as implemented by the neural network application, a point to apply frequency scaling in the graphics processor that does not degrade performance of the neural network application, the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization implemented via the tree structure; and

determine, using the library as implemented by the neural network application, a core frequency of the frequency scaling applied at the point, wherein the library is to account for skew characteristics associated with the distributed gradient synchronization to decide the core frequency.

22. The system of claim 21 , wherein the skew characteristics comprise at least the skew pattern.

23. The system of claim 21 , wherein the graphics processor is further operable to introduce a sparse matrix representation for weights to overlap communication and computation across the one or more nodes associated with the neural network application to reduce communication costs.

24. The system of claim 21 , wherein the graphics processor is further to automatically analyze failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.

25. The system of claim 24 , wherein the graphics processor is further to provide one or more of successful execution information obtained from successful execution of programs and failed execution information obtained from failed execution of programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2017
From: BARIK, RAJKISHORE; LEWIS, BRIAN T.; SUNDARESAN, MURALI; JACKSON, JEFFREY; CHEN, FENG; CHEN, XIAOMING; MACPHERSON, MIKE
To: INTEL CORPORATION
Reel/Frame 042978/0750 →
Continuity (1)
Related Publication 20180314936A1 · Nov 1, 2018
Cited By (2)
US 12,277,406 US 12,350,834