IP Library › Patent Application 18384023
Patent Application
App. No. 18/384,023

MODEL-SPECIFIC ASIC COMPILATION USING FUSED KERNEL REPLACEMENT

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/384,023
Abstract

Embodiments herein describe translating specialized functions for one type of hardware platform into executable code for a different type of hardware platform. For example, a specialized function developed for a first hardware platform (e.g., a GPU or CPU) can be translated into an intermediate representation (IR) containing arguments for a second hardware platform (e.g., a model-specific chipset) by a compiler. The compiler can then convert the IR into executable code for the second hardware platform.

Claims (40)

1 . A method, comprising:

receiving artificial intelligence (AI) model code containing a specialized function for a first one or more types of hardware platforms; and

converting, by a compiler, the specialized function into executable code for a second type of hardware platform.

2 . The method of claim 1 , wherein the first one or more types of hardware platforms are capable of executing different types of AI models, and the second type of hardware platform is optimized for only one type of AI model.

3 . The method of claim 2 , wherein the first one or more types of hardware platforms comprise at least one of a central processing unit (CPU) or a graphics processing unit (GPU).

4 . The method of claim 2 , wherein the second type of hardware platform comprises a model-specific chipset is optimized to execute only transformer models.

5 . The method of claim 4 , further comprising:

training a transformer model defined in the AI model code using the first one or more types of hardware platforms; and

performing inference using the trained transformer model on the model-specific chipset.

6 . The method of claim 1 , wherein the specialized function, when compiled, results in a fused kernel in the first and second type of hardware platforms.

7 . The method of claim 6 , wherein the specialized function comprises a plurality of lower-level functions defined by a machine learning (ML) or AI framework that are executed sequentially by the fused kernel.

8 . The method of claim 7 , wherein converting the specialized function into the executable code further comprises:

translating, by the compiler, the specialized function into an intermediate representation (IR); and

converting the IR into the executable code, wherein the IR comprises values of arguments that configure the second type of hardware platform to perform the plurality of lower-level functions defined in the specialized function.

9 . The method of claim 1 , wherein converting the specialized function into the executable code further comprises:

translating, by the compiler, the specialized function into an IR; and

converting the IR into the executable code, wherein the IR comprises values of arguments for performing at least one of matrix multiplication or attention operations on the second type of hardware platform.

10 . The method of claim 1 , wherein the compiler supports a plurality of specialized functions for the first one or more types of hardware platforms but supports only a limited number of lower-level functions of an ML or AI framework.

11 . A non-transitory computer readable medium having program instructions embodied therewith, the program instructions executable by a processor to perform an operation, the operation comprising:

receiving AI model code containing a specialized function for a one or more types of hardware platforms;

translating, by a compiler, the specialized function into an IR; and

converting, by the compiler, the IR into executable code for a second type of hardware platform.

12 . The non-transitory computer readable medium of claim 11 , wherein the first one or more types of hardware platforms are capable of executing different types of AI models, and the second type of hardware platform is optimized for only one type of AI model, wherein the first one or more types of hardware platforms comprise at least one of a CPU or a GPU.

13 . The non-transitory computer readable medium of claim 11 , wherein the first one or more types of hardware platforms are capable of executing different types of AI models, and wherein the second type of hardware platform comprises a model-specific chipset that is optimized to execute only transformer models.

14 . The non-transitory computer readable medium of claim 13 , wherein the operation further comprises:

training a transformer model defined in the AI model code using the first one or more types of hardware platforms; and

performing inference using the trained transformer model on the model-specific chipset.

15 . The non-transitory computer readable medium of claim 11 , wherein the specialized function, when compiled, results in a fused kernel in the first and second type of hardware platforms, wherein the specialized function comprises a plurality of lower-level functions defined by a ML or AI framework that are executed sequentially by the fused kernel.

16 . The non-transitory computer readable medium of claim 15 , wherein the IR comprises values of arguments that configure the second type of hardware platform to perform the plurality of lower-level functions defined in the specialized function.

17 . A system, comprising:

one or more processors; and

memory storing a compiler which, when executed by the one or more processors, performs an operation comprising:

receiving AI model code containing a specialized function for a first one or more types of hardware platforms;

translating the specialized function into an IR; and

converting the IR into executable code for a second type of hardware platform.

18 . The system of claim 17 , wherein the first one or more types of hardware platforms are capable of executing different types of AI models, and the second type of hardware platform is optimized for only one type of AI model, wherein the first one or more types of hardware platforms comprise at least one of a CPU or a GPU.

19 . The system of claim 17 , wherein the first one or more types of hardware platforms are capable of executing different types of AI models, and wherein the second type of hardware platform comprises a model-specific chipset that is optimized to execute only transformer models.

20 . The system of claim 19 , wherein the operation further comprises:

training a transformer model defined in the AI model code using the first one or more types of hardware platforms; and

performing inference using the trained transformer model on the model-specific chipset.

Assignments (3)
SECURITY INTEREST Recorded Jul 22, 2025
From: ETCHED.AI, INC.
To: TRIPLEPOINT CAPITAL LLC
Reel/Frame 071792/0869 →
SECURITY INTEREST Recorded Apr 24, 2024
From: ETCHED.AI, INC.
To: TRIPLEPOINT CAPITAL LLC
Reel/Frame 067204/0877 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2023
From: UBERTI, GAVIN; SHLOMI, TOM
To: ETCHED.AI INC.
Reel/Frame 065355/0527 →