IP Library Granted Patent US 12,499,362
Granted Patent B2
US 12,499,362 · App. 17/480,236 · Granted Dec 16, 2025

Model optimization in infrastructure processing unit (IPU)

Inventors: Yamini Nimmagadda (Portland, OR); Susanne M. Balle (Hudson, NH); Olugbemisola Oniyinde (Folsom, CA)
Assignee: INTEL CORPORATION
G06N3/08G06F9/5027G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,362
App. No.
17/480,236
Granted
Dec 16, 2025
Kind
B2
Abstract

An Infrastructure Processing Unit (IPU), including: a model optimization processor configured to optimize an artificial intelligence (AI) model for an accelerator managed by the IPU, and deploy the optimized AI model to the accelerator for execution of an inference; and a local memory configured to store data related to the AI model optimization.

Claims (17)

1 . At least one non-transitory computer-readable medium having stored thereon instructions which, when executed, cause a computing device to perform operations comprising:

optimizing an artificial intelligence (AI) model for an accelerator managed by a processing unit including an Infrastructure Processing Unit (IPU), wherein optimizing includes converting a high precision model to a compact low-precision model, wherein optimizing further includes identifying common processes relating to a model optimization pipeline and offloading the identified common processes to the IPU;

storing data related to the AI model optimization; and

deploying the optimized AI model to the accelerator for execution of an inference.

2 . The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise managing one or more workloads using the IPU, wherein optimizing includes converting a 32-bit floating point model to a 16-bit floating point model or an 8-bit integer type model, wherein the computing device comprises processing circuitry coupled to a memory, the processing circuitry having one or more of application processing circuitry or graphics processing circuitry.

3 . A method comprising:

optimizing, by a computing device, an artificial intelligence (AI) model for an accelerator managed by a processing unit including an Infrastructure Processing Unit (IPU), wherein optimizing includes converting a high precision model to a compact low-precision model, wherein optimizing further includes identifying common processes relating to a model optimization pipeline and offloading the identified common processes to the IPU;

storing data related to the AI model optimization; and

deploying the optimized AI model to the accelerator for execution of an inference.

4 . The method of claim 3 , further comprising managing one or more workloads using the IPU, wherein optimizing includes converting a 32-bit floating point model to a 16-bit floating point model or an 8-bit integer type model, wherein the computing device comprises processing circuitry coupled to a memory, the processing circuitry having one or more of application processing circuitry or graphics processing circuitry.

5 . An apparatus comprising:

processing circuitry coupled to a memory, the processing circuitry to:

optimize an artificial intelligence (AI) model for an accelerator managed by an Infrastructure Processing Unit (IPU), wherein to optimize includes to convert a high precision model to a compact low-precision model, wherein optimizing further includes identifying common processes relating to a model optimization pipeline and offloading the identified common processes to the IPU;

store data related to the AI model optimization; and

deploying the optimized AI model to the accelerator for execution of an inference.

6 . The apparatus of claim 5 , further comprising managing one or more workloads using the IPU, wherein to optimize includes to convert a 32-bit floating point model to a 16-bit floating point model or an 8-bit integer type model.

7 . The apparatus of claim 5 , wherein the processing circuitry comprises one or more of application processing circuitry or graphics processing circuitry.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2021
From: NIMMAGADDA, YAMINI; BALLE, SUSANNE M.; ONIYINDE, OLUGBEMISOLA
To: INTEL CORPORATION
Reel/Frame 057541/0345 →
Continuity (1)
Related Publication 20220207358A1 · Jun 30, 2022
References Cited (9)
US 10983761B2 · Svyatkovskiy · 2021 [cited by examiner]
US 20200302298A1 · Van Baalen · 2020 [cited by examiner]
US 20210350210A1 · Gong · 2021 [cited by examiner]
US 20210383206A1 · Teppoeva · 2021 [cited by examiner]
US 20220019461A1 · Balle · 2022 [cited by examiner]
US 20220091915A1 · Perneti · 2022 [cited by examiner]
US 20220207358A1 · Nimmagadda · 2022 [cited by examiner]
US 20220261685A1 · Sabat · 2022 [cited by examiner]
US 20240296283A1 · Wei · 2024 [cited by examiner]