IP Library Patent Application 18653722
Patent Application
App. No. 18/653,722

MIXED-PRECISION NEURAL NETWORKS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/653,722
Abstract

Techniques for mixed precision quantization of a machine learning (ML) model. The techniques include receiving a target performance relating to the ML model including objects of a first data type represented by a first number of bits, wherein the target performance relates to changing a first portion of the objects to a second data type represented by a second number of bits and changing a second portion of the objects to a third data type represented by a third number of bits. The techniques further include selecting the first portion and the second portion, based on maintaining a performance relating to the ML model at or below the target performance, and changing the first portion of objects from the first data type to the second data type and the second portion of objects from the first data type to the third data type.

Claims (46)

1 . A method comprising:

receiving a target performance relating to a machine learning (ML) model comprising a plurality of objects of a first data type represented by a first number of bits, wherein the target performance relates to changing a first portion of the plurality of objects to a second data type represented by a second number of bits different from the first number of bits and changing a second portion of the plurality of objects to a third data type represented by a third number of bits different from the first and second numbers of bits;

selecting the first portion and the second portion of the plurality of objects, based on maintaining a performance relating to the ML model at or below the target performance; and

changing, by a processor and based on the selecting the first portion and the second portion, the first portion of the plurality of objects from the first data type to the second data type and the second portion of the plurality of objects from the first data type to the third data type.

2 . The method of claim 1 , wherein the ML model comprises a neural network, and wherein the plurality of objects comprises a plurality of weights and a plurality of binary large objects (BLOBs).

3 . The method of claim 1 , wherein the first data type comprises a floating point data type, wherein the second data type comprises a first integer data type, and wherein the third data type comprises a second integer data type.

4 . The method of claim 1 , further comprising:

sorting the plurality of objects in the ML model based on size.

5 . The method of claim 4 , wherein the ML model comprises a neural network, and wherein sorting the plurality of objects in the ML model is further based on proximity of the respective objects to a root of a network graph representing the neural network.

6 . The method of claim 1 ,

wherein the third data type comprises more bits than the second data type and fewer bits than the first data type.

7 . The method of claim 6 ,

wherein selecting the first portion and the second portion comprises:

determining a total performance for the ML model based on changing the plurality of objects to the second data type; and

iteratively identifying one or more of the plurality of objects to change to the third data type, instead of the second data type, based on a target performance increase.

8 . The method of claim 1 , wherein the target performance comprises a target bandwidth change.

9 . A system comprising:

a processor; and

a memory storing instructions, which when executed by the processor, cause the processor to perform operations comprising:

receiving a target performance relating to a machine learning (ML) model comprising a plurality of objects of a first data type represented by a first number of bits, wherein the target performance relates to changing a first portion of the plurality of objects to a second data type represented by a second number of bits different from the first number of bits and changing a second portion of the plurality of objects to a third data type represented by a third number of bits different from the first and second numbers of bits;

selecting the first portion and the second portion of the plurality of objects, based on maintaining a performance relating to the ML model at or below the target performance; and

changing, by a processor and based on the selecting the first portion and the second portion, the first portion of the plurality of objects from the first data type to the second data type and the second portion of the plurality of objects from the first data type to the third data type.

10 . The system of claim 9 , wherein the ML model comprises a neural network, and wherein the plurality of objects comprises a plurality of weights and a plurality of binary large objects (BLOBs).

11 . The system of claim 9 , wherein the first data type comprises a floating point data type, wherein the second data type comprises a first integer data type, and wherein the third data type comprises a second integer data type.

12 . The system of claim 9 , the operations further comprising:

sorting the plurality of objects in the ML model based on size.

13 . The system of claim 12 , wherein the ML model comprises a neural network, and wherein sorting the plurality of objects in the ML model is further based on proximity of the respective objects to a root of a network graph representing the neural network.

14 . The system of claim 9 ,

wherein the third data type comprises more bits than the second data type and fewer bits than the first data type, and

wherein selecting the first portion and the second portion comprises:

determining a total performance for the ML model based on changing the plurality of objects to the second data type; and

iteratively identifying one or more of the plurality of objects to change to the third data type, instead of the second data type, based on a target performance increase.

15 . A non-transitory computer readable medium comprising stored instructions, which when executed by a processor, cause the processor to perform operations comprising:

receiving a target performance relating to a machine learning (ML) model comprising a plurality of objects of a first data type represented by a first number of bits, wherein the target performance relates to changing a first portion of the plurality of objects to a second data type represented by a second number of bits different from the first number of bits and changing a second portion of the plurality of objects to a third data type represented by a third number of bits different from the first and second numbers of bits;

selecting the first portion and the second portion of the plurality of objects, based on maintaining a performance relating to the ML model at or below the target performance; and

changing, by a processor and based on the selecting the first portion and the second portion, the first portion of the plurality of objects from the first data type to the second data type and the second portion of the plurality of objects from the first data type to the third data type.

16 . The non-transitory computer readable medium of claim 15 , wherein the ML model comprises a neural network, and wherein the plurality of objects comprises a plurality of weights and a plurality of binary large objects (BLOBs).

17 . The non-transitory computer readable medium of claim 15 , wherein the first data type comprises a floating point data type, wherein the second data type comprises a first integer data type, and wherein the third data type comprises a second integer data type.

18 . The non-transitory computer readable medium of claim 15 , the operations further comprising:

sorting the plurality of objects in the ML model based on size.

19 . The non-transitory computer readable medium of claim 18 , wherein the ML model comprises a neural network, and wherein sorting the plurality of objects in the ML model is further based on proximity of the respective objects to a root of a network graph representing the neural network.

20 . The non-transitory computer readable medium of claim 15 ,

wherein the third data type comprises more bits than the second data type and fewer bits than the first data type, and

wherein selecting the first portion and the second portion comprises:

determining a total performance for the ML model based on changing the plurality of objects to the second data type; and

iteratively identifying one or more of the plurality of objects to change to the third data type, instead of the second data type, based on a target performance increase.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2026
From: SYNOPSYS, INC.
To: MIPS HOLDING, INC.
Reel/Frame 075801/0225 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2024
From: PENNELLO, THOMAS
To: SYNOPSYS, INC.
Reel/Frame 067309/0441 →