IP Library › Granted Patent US 12,585,443
Granted Patent B2
US 12,585,443 · App. 17/637,190 · Granted Mar 24, 2026

Compiler for neural accelerator

Inventors: Arun Chauhan (Belmont, CA); Raksit Ashok (Campbell, CA); Dong Hyuk Woo (San Jose, CA)
Assignee: Google LLC
G06F8/41G06N3/06G06N3/063G05B2219/13119G05B2219/23266G06N3/045G06N3/105
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,585,443
App. No.
17/637,190
Granted
Mar 24, 2026
Kind
B2
Abstract

A compiler of a computing device is described that identifies a sequence of neural network models frequently invoked by an application of the computing device, compiles the models in that sequence, and loads a static random access memory (SRAM) of a hardware accelerator with the compiled models only when the same compiled models—from another, but same, sequence that was previously invoked—are not already present in the SRAM. This prevents unnecessary reloading of compiled models into the SRAM, thereby increasing runtime speed and conserving computational energy.

Claims (53)

1 . A computer-implemented method comprising:

identifying a first set of neural network models that have been executed on a hardware accelerator of a computing device more than a threshold number of times in a preset amount of time in the past;

identifying a sequence in which the first set of models were executed on the hardware accelerator;

compiling each neural network model of the first set of neural network models in the sequence for execution by the hardware accelerator and assigning a first hash to the identified sequence of models;

outputting, for each neural network model of the first set of neural network models in the sequence, the compiled model to the hardware accelerator for storage in one or more memories of the hardware accelerator,

compiling a second sequence of neural network models and computing a second hash for the second sequence of neural network models;

determining that the first hash matches the second hash; and

in response, bypassing loading compiled models in the second sequence of neural network models into storage of the hardware accelerator.

2 . The method of claim 1 , further comprising:

receiving, for each neural network model of the first set of neural network models, a data structure including parameters of the neural network model,

wherein the compiling further comprises compiling the data structure for each neural network model of the first set of neural network models to generate a compiled data structure for each neural network model of the first set of neural network models, the compiled data structure being the compiled model.

3 . The method of claim 1 , further comprising:

assigning the same first hash to each compiled model in the identified sequence of models;

outputting the first hash along with each compiled model in the sequence to the hardware accelerator for storage in the one or more memories of the hardware accelerator;

assigning the same second hash to each compiled model in the previously compiled sequence of models, the second hash being same as the first hash when the second sequence is same as the first sequence, the second hash being different from the first hash when the second sequence is different from the first sequence,

wherein:

if the second hash is different from the first hash, the hardware accelerator is configured to replace each compiled model in the first sequence in the one or more memories with each compiled model in the second sequence in the one or more memories;

if the second hash is same as the first hash, the hardware accelerator is configured to prevent erasing each compiled model in the first sequence from the one or more memories.

4 . The method of claim 1 , wherein each of the first set of neural network models has been processed on the hardware accelerator more than five times in the past.

5 . The method of claim 4 , wherein the compiler compiles the first set of neural network models while the hardware accelerator simultaneously performs neural network computations of other one or more neural network models.

6 . The method of claim 1 , further comprising:

updating the identification of the first set of models and the identification of the sequence after preset intervals of time.

7 . The method of claim 6 , further comprising:

abstaining, for a preset time, from the updating in response to a failure of the compilation of the first set of neural network models.

8 . The method of claim 7 , wherein the abstaining comprises:

abstaining, for 7500 milliseconds, from the updating in response to the failure of the compilation of the first set of neural network models.

9 . The method of claim 1 , wherein the compiling of the first set of neural network models comprises:

determining that each neural network model within the first set of neural network models has a particular size that is compatible for the compiling.

10 . The device of claim 1 , wherein the compiling of the first set of neural network models comprises:

compiling only a single neural network model at any time.

11 . The method of claim 1 , wherein the sequence comprises a face recognition neural network model and one or more dependent neural network models that are to be processed after processing the face recognition neural network model.

12 . A system comprising:

a compiler configured to:

identify a first set of neural network models that have been executed on a hardware accelerator of a computing device more than a threshold number of times in a preset amount of time in the past;

identify a sequence in which the first set of models are executed on the hardware accelerator;

compile each neural network model of the first set of neural network models in the sequence for execution by the hardware accelerator and assigning a first hash to the identified sequence of models;

output, for each neural network model of the first set of neural network models in the sequence, the compiled model to the hardware accelerator for storage in one or more memories of the hardware accelerator;

compiling a second sequence of neural network models and computing a second hash for the second sequence of neural network models;

determining that the first hash matches the second hash; and

in response, bypassing loading compiled models in the second sequence of neural network models into storage of the hardware accelerator;

the hardware accelerator comprising the one or more memories to store the compiled model for each neural network model of the first set of neural network models according to the sequence, the storage according to the sequence of the compiled model for each neural network model of the first set of neural network models in the one or more memories preventing a need for reloading a compiled result recompilation of the sequence of the first set of neural network models into the one or more memories when the first set of neural network models are to be executed again on the hardware accelerator.

13 . The system of claim 12 , wherein the one or more memories is static random access memory (SRAM).

14 . The system of claim 12 , wherein the hardware accelerator further comprises a plurality of computing units, each comprising at least one processor, configured to process the first set of neural network models.

15 . The system of claim 14 , wherein:

each computing unit of the plurality of computing units further comprises a memory; and

the plurality of computing units are coupled serially via at least one bus.

16 . The system of claim 12 , wherein the first set of neural network models include at least one face recognition neural network model.

17 . The system of claim 16 , wherein the at least one face recognition neural network model is activated in response to the controller receiving an instruction from the computing device to execute the at least one face recognition neural network model.

18 . The system of claim 17 , wherein the computing device comprises an application and an application programming interface (API), the application generating the instruction to be sent via the API.

19 . The system of claim 18 , wherein the application is a camera application.

20 . The system of claim 18 , wherein:

the API is a Neural Networks API (NNAPI); and

the computing device is an Android device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2022
From: CHAUHAN, ARUN; ASHOK, RAKSIT; WOO, DONG HYUK
To: GOOGLE LLC
Reel/Frame 059197/0673 →
Continuity (1)
Related Publication 20220300826A1 · Sep 22, 2022
References Cited (44)
US 10789402B1 · Vemuri · 2020 [cited by examiner]
US 12086572B1 · Wu · 2024 [cited by examiner]
US 12093806B1 · Zejda · 2024 [cited by examiner]
US 20190286973A1 · Kovvuri et al. · 2019 [cited by applicant]
US 20190308099A1 · Lalonde et al. · 2019 [cited by applicant]
US 20190354489A1 · Gupta · 2019 [cited by examiner]
US 20190391796A1 · Brady · 2019 [cited by examiner]
US 20200082273A1 · Rossi · 2020 [cited by examiner]
US 20200241856A1 · Kim · 2020 [cited by examiner]
US 20210042259A1 · Koeplinger · 2021 [cited by examiner]
US 20210056389A1 · Yang · 2021 [cited by examiner]
US 20210081806A1 · Chai · 2021 [cited by examiner]
US 20210158131A1 · Jain · 2021 [cited by examiner]
US 20210174214A1 · Venkatesan · 2021 [cited by examiner]
CN 109858621 · 2019 [cited by applicant]
Ahn et al., “Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation” Jan. 23, 2020, arXiv: 2001.08743v1, pp. 1-17. (Year: 2020). [cited by examiner]
Jiang et al., “Hardware/Software Co-Exploration of Neural Architectures” Jan. 11, 2020, arXiv: 1907.04650v2, pp. 1-10. (Year: 2020). [cited by examiner]
Lattner et al., “MLIR: A Compiler Infrastructure for the End of Moore's Law” Mar. 1, 2020, arXiv: 2002.11054v2, pp. 1-21. (Year: 2020). [cited by examiner]
Chen et al., “TVM: An Automated End-to-End Optimizing Compiler for Deep Learning” Oct. 2018, pp. 579-594. (Year: 2018). [cited by examiner]
Ignatov et al., “AI Benchmark: Running Deep Neural Networks on Android Smartphones” Oct. 15, 2018, arXiv: 1810.01109v2, pp. 1-15. (Year: 2018). [cited by examiner]
Gong et al., “Recurrent Embedding Aggregation Network for Video Face Recognition” Jun. 25, 2019, arXiv: 1904.12019v2, pp. 1-10. (Year: 2019). [cited by examiner]
Haj-Ali et al., “NeuroVectorizer: End-to-End Vectorization with Deep Reinforcement Learning” Feb. 2020, pp. 242-255. (Year: 2020). [cited by examiner]
Huang et al., “AutoPhase: Juggling HLS Phase Orderings in Random Forests with Deep Reinforcement Learning” Mar. 4, 2020, arXiv: 2003.00671v2, pp. 1-12. (Year: 2020). [cited by examiner]
Vasilache et al., “Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions” Jun. 29, 2018, arXiv: 1802.04730v3, pp. 1-37. (Year: 2018). [cited by examiner]
Zerrell et Bruestle, “Stripe: Tensor Compilation via the Nested Polyhedral Model” Mar. 14, 2019, arXiv: 1903.06498v1, pp. 1-13. (Year: 2019). [cited by examiner]
Kaufman et al., “Learned TPU Cost Model for XLA Tensor Program” 2019, pp. 1-6. (Year: 2019). [cited by examiner]
Baghdadi et al., “Tiramisu: A Polyhedral Compiler for Dense and Sparse Deep Learning” May 7, 2020, arXiv: 2005.04091v1, pp. 1-9. (Year: 2020). [cited by examiner]
Bondhugula, Uday, “High Performance Code Generation in MLIR: An Early Case Study with GEMM” Mar. 1, 2020, arXiv: 2003.00532v1, pp. 1-23. (Year: 2020). [cited by examiner]
Jain et al., “The OoO VLIW JIT Compiler for GPU Inference” Jan. 31, 2019, arXiv: 1901.10008v2, pp. 1-7. (Year: 2019). [cited by examiner]
Venkataramani et al., “DeepTools: Compiler and Execution Runtime Extensions for RaPiD AI Accelerator” Sep. 2019, pp. 102-110. (Year: 2019). [cited by examiner]
Xing et al., “DNNVM: End-to-End Compiler Leveraging Heterogeneous Optimizations on FPGA-based CNN Accelerators” Jul. 25, 2019, arXiv: 1902.07463v2, pp. 1-18. (Year: 2019). [cited by examiner]
Einziger et al., “TinyLFU: A Highly Efficient Cache Admission Policy” Nov. 2017, pp. 1-31. (Year: 2017). [cited by examiner]
Chen et al., “SLIDE: In Defense of Smart Algorithms over Hardware Acceleration for Large-Scale Deep Learning Systems” Mar. 1, 2020, arXiv: 1903.03129v2, pp. 1-16. (Year: 2020). [cited by examiner]
Jia et al., “TASO: Optimizing Deep Learning Computations with Automatic Generation of Graph Substitutions” Oct. 2019, pp. 1-16. (Year: 2019). [cited by examiner]
Kumar et al., “Quiver: An Informed Storage Cache for Deep Learning” Feb. 2020, pp. 283-296. (Year: 2020). [cited by examiner]
International Preliminary Report on Patentability in International Appln. No. PCT/US2020/021714, dated Sep. 22, 2022, 8 pages. [cited by applicant]
Office Action in Chinese Appln. No. 202080041079.6, Mailed on Aug. 23, 2024, 13 pages (with English translation). [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2020/021714, dated Dec. 2, 2020, 13 pages. [cited by applicant]
Extended European Search Report in European Appln. No. 22207462.7, dated Mar. 24, 2023, 6 pages. [cited by applicant]
Cettei et al., “Code Cache Management in Dynamic Optimization Systems,” Thesis for the Degree of Doctor of Philosophy, Harvard University, Division of Engineering and Applied Sciences, May 2004, 111 pages. [cited by applicant]
Extended European Search Report in European Appln. No. 24212215.8, mailed on Mar. 24, 2025, 12 pages. [cited by applicant]
Lee et al., “Automated Neural Network Accelerator Generation Framework for Multiple Neural Network Applications,” Paper, Presented at Proceedings of the TENCON 2018—2018 IEEE Region 10 Conference, Jeju, South Korea, Oct… [cited by applicant]
Office Action in European Appln. No. 24212215.8, mailed on Dec. 10, 2025, 6 pages. [cited by applicant]
Taher et al., “Virtual Configuration Management: A Technique for Partial Runtime Reconfiguration,” IEEE Transactions on Computers, Oct. 2009, 58(10):1398-1410. [cited by applicant]