IP Library Granted Patent US 12,430,566
Granted Patent B2
US 12,430,566 · App. 17/356,550 · Granted Sep 30, 2025

Performance modeling and analysis of artificial intelligence (AI) accelerator architectures

Inventors: Bhaskar J Karmakar (Bangalore, IN); Anand R Kumar (Bangalore, IN); Rajagopal Hariharan (Bangalore, IN); Saurabh Tiwari (Bangalore, IN)
Assignee: Habana Labs Ltd.
G06N3/10G06F30/32G06N3/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,566
App. No.
17/356,550
Granted
Sep 30, 2025
Kind
B2
Abstract

A method includes receiving one or more input Artificial Intelligence (AI) networks, transforming the AI networks into respective graphs including interconnected logical operators, and mapping the graphs onto a design of a hardware accelerator including a plurality interconnected hardware engines. A performance of running the AI networks on the design of the hardware accelerator is simulated using a petri-net simulation.

Claims (28)

1. A method, comprising:

receiving one or more input Artificial Intelligence (AI) networks;

transforming the AI networks into respective graphs comprising interconnected logical operators;

mapping the graphs onto a design of a hardware accelerator comprising a plurality of interconnected hardware engines; and

using a petri-net simulation on the mappings, simulating an impact of the design on access to shared hardware resources to thereby simulate a performance of running the AI networks on the design of the hardware accelerator.

2. The method according to claim 1 , wherein transforming the AI networks comprises modeling the logical operators in terms of performance costs incurred by computations and data-access operations performed by the logical operators.

3. The method according to claim 1 , wherein transforming the AI networks into the graphs further comprises defining one or more tensors, which are operated on by the logical operators.

4. The method according to claim 3 , wherein mapping the graphs onto the design further comprises mapping the tensors onto memories of the hardware accelerator.

5. The method according to claim 1 , wherein mapping the logical operators onto the hardware engines comprises receiving user input that specifies properties of the design of the hardware accelerator, and mapping the logical operators onto the hardware engines in accordance with the properties.

6. The method according to claim 1 , wherein mapping the graphs onto the design, and simulating the performance of running the AI networks on the design, comprise evaluating both hardware and software components of the hardware accelerator.

7. The method according to claim 1 , wherein simulating the performance comprises evaluating multiple alternative transformations of the AI networks into graphs comprising interconnected logical operators.

8. The method according to claim 1 , wherein simulating the performance comprises evaluating multiple alternative mappings of the graphs onto designs.

9. The method according to claim 1 , wherein simulating the performance comprises evaluating the performance of a given design of the hardware accelerator over a plurality of different AI networks.

10. The method according to claim 1 , wherein simulating the performance comprises receiving user input that specifies a speed vs. accuracy setting for the petri-net simulation, and running the petri-net simulation in accordance with the specified speed vs. accuracy setting.

11. The method according to claim 1 , further comprising specifying the design of the hardware accelerator using a hierarchical hardware description model.

12. A system, comprising:

a compiler, configured to receive one or more input Artificial Intelligence (AI) networks, to transform the AI networks into respective graphs comprising interconnected logical operators, and to map the graphs onto a design of a hardware accelerator comprising a plurality of interconnected hardware engines; and

a petri-net simulator, configured to simulate an impact of the design on access to shared hardware resources to thereby simulate a performance of running the AI networks on the design of the hardware accelerator.

13. The system according to claim 12 , wherein the compiler is configured to model the logical operators in terms of performance costs incurred by computations and data-access operations performed by the logical operators.

14. The system according to claim 12 , wherein the graphs further comprise one or more tensors, which are operated on by the logical operators.

15. The system according to claim 14 , wherein the compiler is configured to map the tensors onto memories of the hardware accelerator.

16. The system according to claim 12 , wherein the compiler is configured to receive user input that specifies properties of the design of the hardware accelerator, and to map the logical operators onto the hardware engines in accordance with the properties.

17. The system according to claim 12 , wherein, in mapping the graphs onto the design and simulating the performance of running the AI networks on the design, the compiler and the petri-net simulator are configured to evaluate both hardware and software components of the hardware accelerator.

18. The system according to claim 12 , wherein, in simulating the performance, the petri-net simulator is configured to evaluate multiple alternative transformations of the AI networks into graphs comprising interconnected logical operators.

19. The system according to claim 12 , wherein, in simulating the performance, the petri-net simulator is configured to evaluate multiple alternative mappings of the graphs onto designs.

20. The system according to claim 12 , wherein the petri-net simulator is configured to evaluate the performance of a given design of the hardware accelerator over a plurality of different AI networks.

21. The system according to claim 12 , wherein the petri-net simulator is configured to receive user input that specifies a speed vs. accuracy setting, and to simulate the performance in accordance with the specified speed vs. accuracy setting.

22. The system according to claim 12 , wherein the design of the hardware accelerator is specified using a hierarchical hardware description model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2025
From: HABANA LABS LTD.
To: INTEL OVERSEAS FUNDING CORPORATION
Reel/Frame 073008/0642 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2021
From: KARMAKAR, BHASKAR J; KUMAR, ANAND R; HARIHARAN, RAJAGOPAL; TIWARI, SAURABH
To: HABANA LABS LTD.
Reel/Frame 056646/0118 →
Priority Claims (1)
IN 202141021314 · May 11, 2021 · national
Continuity (1)
Related Publication 20220366267A1 · Nov 17, 2022
References Cited (23)
US 12141683B2 · Raha · 2024 [cited by examiner]
US 20160034305A1 · Shear · 2016 [cited by examiner]
US 20160217368A1 · Ioffe et al. · 2016 [cited by applicant]
US 20200371843A1 · Garg et al. · 2020 [cited by applicant]
US 20210240526A1 · Weber · 2021 [cited by examiner]
US 20210271960A1 · Raha · 2021 [cited by examiner]
US 20210350233A1 · Saboori · 2021 [cited by examiner]
US 20220067520A1 · Dalli · 2022 [cited by examiner]
US 20220083655A1 · Yang · 2022 [cited by examiner]
US 20220147876A1 · Dalli · 2022 [cited by examiner]
US 20220156614A1 · Dalli · 2022 [cited by examiner]
US 20220374288A1 · Kibardin · 2022 [cited by examiner]
US 20240104363A1 · Klaiber · 2024 [cited by examiner]
US 20240114477A1 · Yerramalli · 2024 [cited by examiner]
US 20240160924A1 · Rackauckas · 2024 [cited by examiner]
US 20240370521A1 · Zhou · 2024 [cited by examiner]
US 20250110808A1 · Kibardin · 2025 [cited by examiner]
Synopsis Inc, “Synopsis—Platform Architect”, datasheet, pp. 1-5, year 2018. [cited by applicant]
Mathworks Inc, “Petri Net Toolbox”, pp. 1-2, years 1994-2021, as downloaded from https://uk.mathworks.com/products/connections/product_detail/petri-net-toolbox.html. [cited by applicant]
Eindhoven University of Technology, “CPN—A concrete language for high-level Petri nets”, pp. 1-168, Apr. 24, 2021, as downloaded from https://web.archive.org/web/20210424143646/http://cpntools.org/wp-content/uploads/201… [cited by applicant]
Choi et al., “Data-free Network Quantization with Adversarial Knowledge Distillation,” arXiv:2005.04136v1, pp. 1-11, May 8, 2020. [cited by applicant]
U.S. Appl. No. 18/353,128 Office Action dated Aug. 22, 2024. [cited by applicant]
U.S. Appl. No. 17/088,625 Office Action dated Jun. 4, 2024. [cited by applicant]