IP Library › Granted Patent US 12,198,037
Granted Patent B2
US 12,198,037 · App. 17/091,633 · Granted Jan 14, 2025

Hardware architecture determination based on a neural network and a network compilation process

Inventors: Or Davidi (Tel-Aviv, IL); Omer Shabtai (Tel-Aviv, IL); Yotam Platner (Tel-Aviv, IL)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06N3/063G06F8/41G06F30/34G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,037
App. No.
17/091,633
Granted
Jan 14, 2025
Kind
B2
Abstract

The described techniques provide for an automated process of hardware design using main applications (e.g., use cases) and a modeling heuristic of a compiler. Hardware may be optimized and designed based on the knowledge of the compiler for the hardware and firmware information. For instance, a user may define constraints (e.g., area constraints, power constraints, performance constraints, accuracy degradation constraints, etc.) and the changes of compilation heuristics and hardware parameters may be optimized to efficiently achieve the user defined constraints. Accordingly, hardware configuration parameters may be optimized based on the neural network's compilation process (e.g., actual compiler constraints) and optimization of power, performance, and area (PPA) constraints (e.g., user defined constraints). Specific neural processor (SNP) hardware may thus be designed based on the optimized hardware configuration parameters (e.g., via modeling heuristics of its compiler and main applications informed via user defined constraints or PPA constraints).

Claims (50)

1. A method comprising:

generating hardware configuration parameters for a neural network hardware accelerator based on hardware constraints and compiler constraints, wherein the hardware configuration parameters include one or more memory sizes, one or more memory form factors, a number of multipliers, a topology of multipliers, or any combination thereof;

compiling a representation of a neural network architecture to produce a representation of an instruction set for implementing the neural network architecture based on the hardware configuration parameters and the compiler constraints;

simulating performance of the neural network architecture based on the hardware configuration parameters and the representation of the instruction set to produce a performance profile;

computing a performance metric based on the performance profile and power performance and area (PPA) constraints; and

iteratively updating the hardware configuration parameters for the neural network architecture based on the performance metric, simulating the performance of the neural network architecture, and recomputing the performance metric based on the updated hardware configuration parameters.

2. The method of claim 1 , wherein:

the performance profile includes two or more parameters from a set comprising algorithm performance, power consumption, memory bandwidth, and chip area.

3. The method of claim 1 , further comprising:

multiplying each parameter of the performance profile by a weight based on the PPA constraints; and

computing the performance metric based on a sum of each parameter of the performance profile multiplied by the corresponding weight, wherein the hardware configuration parameters are updated based on the sum.

4. The method of claim 1 , further comprising:

iterating between computing the performance metric and updating the hardware configuration to produce final hardware configuration parameters.

5. The method of claim 4 , wherein:

the iterating is based on simulated annealing.

6. The method of claim 4 , wherein:

the iterating is based on stochastic gradient descent.

7. The method of claim 1 , further comprising:

prompting a user to identify the PPA constraints; and

receiving the PPA constraints from the user based on the prompting.

8. The method of claim 1 , further comprising:

generating a register-transfer level (RTL) design for the neural network architecture based on the hardware configuration parameters.

9. The method of claim 1 , further comprising:

manufacturing a specific neural processor (SNP) based on the updated hardware configuration parameters.

10. The method of claim 1 , wherein:

the neural network architecture comprises a convolutional neural network (CNN).

11. The method of claim 1 , wherein:

the neural network architecture is configured to facial recognition tasks.

12. The method of claim 1 , further comprising:

selecting the hardware configuration parameters to be optimized based on the neural network architecture.

13. A method for hardware acceleration, comprising:

generating hardware configuration parameters for a neural network architecture based on hardware constraints;

compiling a representation of the neural network architecture to produce a representation of an instruction set for implementing the neural network architecture based on the hardware configuration parameters;

simulating performance of the neural network architecture based on the hardware configuration parameters and the representation of the instruction set to produce a performance profile;

computing a performance metric by weighting parameters of the performance profile based on power performance and area (PPA) constraints; and

iteratively updating the hardware configuration parameters for the neural network architecture based on the performance metric, simulating the performance of the neural network architecture, and recomputing the performance metric based on the updated hardware configuration parameters.

14. The method of claim 13 , wherein:

the hardware configuration parameters are generated based on compiler constraints, and the instruction set is compiled based on the compiler constraints.

15. An apparatus for hardware acceleration, comprising: a processor and a memory storing instructions and in electronic communication with the processor, the processor being configured to:

generate hardware configuration parameters for a neural network architecture based on hardware constraints and compiler constraints;

compute a performance metric based on power performance and area (PPA) constraints and a performance profile;

compile a representation of a neural network architecture to produce a representation of an instruction set for implementing the neural network architecture based on the hardware configuration parameters and the compiler constraints; and

simulate performance of the neural network architecture to produce a performance profile based on the hardware configuration parameters and the representation of the instruction set to produce a performance profile,

wherein the hardware configuration parameters for the neural network architecture are iteratively updated based on computing the performance metric, simulating the performance of the neural network architecture, and recomputing the performance metric based on the updated hardware configuration parameters.

16. The apparatus of claim 15 , further comprising:

a user interface configured to receive the PPA constraints from a user.

17. The apparatus of claim 15 , further comprising:

a design component configured to design a specific neural processor (SNP) based on the hardware configuration parameters.

18. The apparatus of claim 15 , wherein:

the decision making component takes compiler constraints corresponding to the policy making component as input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2020
From: DAVIDI, OR; SHABTAI, OMER; PLATNER, YOTAM
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 054302/0059 →
Continuity (1)
Related Publication 20220147801A1 · May 12, 2022
References Cited (12)
US 20180157965A1 · Sun · 2018 [cited by examiner]
US 20180189638A1 · Nurvitadhi · 2018 [cited by examiner]
US 20190286984A1 · Vasudevan · 2019 [cited by examiner]
US 20200082247A1 · Wu · 2020 [cited by examiner]
Kolala Venkataramanaiah, Yufei Ma, Shihui Yin, Eriko Nurvithadhi*, Aravind Dasu, Yu Cao, Jae-sun Seo, “Automatic Compiler Based FPGA Accelerator for CNN Training,” 2019 29th International Conference on Field Programmabl… [cited by examiner]
Naveen Suda, Vikas Chandr, Ganesh Dasika, Abinash Mohanty, Yufei Ma, Sarma Vrudhula, Jae-sun Seo, Yu Cao, “Throughput-Optimized OpenCL-based FPGA Accelerator for Large-Scale Convolutional Neural Networks,” ACM p. 16-25 … [cited by examiner]
Rivera-Acosta, et al., “Automatic Tool for Fast Generation of Custom Convolutional Neural Networks Accelerators for FPGA”, located on the internet: https://www.mdpi.com/2079-9292/8/6/641/htm. [cited by applicant]
Du, et al., “Optimizing of Convolutional Neural Network Accelerator”, Green Electronics, Chapter 8, pp. 147-166, found on the internet: https://www.intechopen.com/books/green-electronics/optimizing-of-convolutional-neur… [cited by applicant]
Ma, et al., “ALAMO: FPGA acceleration of deep learning algorithms with a modularized RTL compiler”, Integration, the VLSI Journal 62 (2018), pp. 14-23, found on the internet: https://par.nsf.gov/servlets/purl/10087145. [cited by applicant]
Venieris,et al., “Toolflows for Mapping Convolutional Neural Networks on FPGAs:A Survey and Future Directions”, ACM Computing Surveys, vol. 0, No. 0, Article 0, Mar. 2018; found on the internet: https://arxiv.org/pdf/18… [cited by applicant]
Song, “AI for AI: Automatically Generate Deep Neural Network Accelerator in FPGA for Data-center and Edge”, IBM Research, China, 41 pages; found on the internet: http://xilinx.eetrend.com/system/files/2019-01/%E6%96%87%… [cited by applicant]
Wikipedia, “Stochastic gradient descent”, 7 pages, found on the internet: https://en.wikipedia.org/wiki/Stochastic_gradient_descent. [cited by applicant]