IP Library › Granted Patent US 12,198,055
Granted Patent B2
US 12,198,055 · App. 18/532,795 · Granted Jan 14, 2025

Incremental precision networks using residual inference and fine-grain quantization

Inventors: Abhisek Kundu (Bangalore, IN); Naveen Mellempudi (Bangalore, IN); Dheevatsa Mudigere (Bangalore, IN); Dipankar Das (Pune, IN)
Assignee: Intel Corporation
G06N3/08G06F9/46G06N3/044G06N3/045G06N3/063G06N3/084G06N5/04G06T15/005G06T15/04G06T15/80G06T17/10G06T17/20G06V10/94
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,055
App. No.
18/532,795
Filed
Dec 7, 2023
Granted
Jan 14, 2025
Kind
B2
Art Unit
2124
USPC
706/25
Abstract

One embodiment provides for a computer-readable medium storing instructions that cause one or more processors to perform operations comprising determining a per-layer scale factor to apply to tensor data associated with layers of a neural network model and converting the tensor data to converted tensor data. The tensor data may be converted from a floating point datatype to a second datatype that is an 8-bit datatype. The instructions further cause the one or more processors to generate an output tensor based on the converted tensor data and the per-layer scale factor.

Claims (34)

1. A graphics processor comprising:

a memory interface; and

a processing cluster coupled with the memory interface, the processing cluster including a plurality of processing resources coupled via a data interconnect, a processing resource of the plurality of processing resources including:

first circuitry including a weight ternarization circuit and an activation quantization circuit, the weight ternarization circuit to convert a weight tensor from a floating-point representation to a ternary representation including a two-bit ternary weight and an eight-bit integer scale factor, and the two-bit ternary weight to represent a weight value of one of negative one, zero, or positive one and wherein the activation quantization circuit is to convert an activation tensor from a floating-point representation to an integer representation; and

second circuitry configured to perform a set of parallel integer compute operations on the ternary representation of the weight tensor and the integer representation of the activation tensor.

2. The graphics processor as in claim 1 , the second circuitry including one or more ternary logic units to perform one or more operations in the set of parallel integer compute operations.

3. The graphics processor as in claim 2 , wherein the second circuitry is configured to perform the set of parallel integer compute operations in response to an instruction provided by a machine learning inferencing framework.

4. The graphics processor as in claim 1 , the activation quantization circuit to convert the activation tensor from a single-precision floating-point representation to an eight-bit integer representation.

5. The graphics processor as in claim 1 , the weight ternarization circuit to convert the weight tensor from a single-precision floating-point representation to the ternary representation.

6. The graphics processor as in claim 5 , wherein the eight-bit integer scale factor is determined to minimize L2 loss between values of a group of pre-trained weights and a group of ternary weights.

7. The graphics processor as in claim 5 , wherein to convert the weight tensor from the single-precision floating-point representation, the weight ternarization circuit is to decompose the weight tensor into a set of orthogonal vectors and ternarize components of the set of orthogonal vectors.

8. The graphics processor as in claim 1 , the weight ternarization circuit to ternarize first weights having a first data distribution into first ternarized weights having a first scale factor and ternarize second weights having a second data distribution into second ternarized weights having a second scale factor different from the first scale factor.

9. A method comprising:

on a graphics processor comprising a processing cluster including a plurality of processing resources coupled via a data interconnect:

converting a weight tensor from a floating-point representation to a ternary representation including a two-bit ternary weight and an eight-bit integer scale factor using a weight ternarization circuit, the two-bit ternary weight to represent a weight value of one of negative one, zero, or positive one;

converting an activation tensor from a floating-point representation to an integer representation using an activation quantization circuit; and

performing a set of parallel integer compute operations on the ternary representation of the weight tensor and the integer representation of the activation tensor using second circuitry.

10. The method as in claim 9 , further comprising performing one or more operations in the set of parallel integer compute operations via one or more ternary logic units of the second circuitry.

11. The method as in claim 10 , further comprising performing the set of parallel integer compute operations in response to an instruction provided by a machine learning inferencing framework.

12. The method as in claim 9 , further comprising converting the activation tensor from a single-precision floating-point representation to an eight-bit integer representation.

13. The method as in claim 9 , further comprising converting the weight tensor from a single-precision floating-point representation to the ternary representation.

14. The method as in claim 13 , further comprising determining the eight-bit integer scale factor to minimize L2 loss between values of a group of pre-trained weights and a group of ternary weights.

15. The method as in claim 14 , wherein converting the weight tensor from the single-precision floating-point representation includes decomposing the weight tensor into a set of orthogonal vectors and ternarizing components of the set of orthogonal vectors.

16. The method as in claim 9 , further comprising, via the weight ternarization circuit:

ternarizing first weights having a first data distribution into first ternarized weights having a first scale factor; and

ternarizing second weights having a second data distribution into second ternarized weights having a second scale factor different from the first scale factor.

17. A data processing system comprising:

a memory device; and

graphics processor coupled with the memory device via a memory interface, the graphics processor comprising a processing cluster including a plurality of processing resources coupled via a data interconnect, wherein a processing resource of the plurality of processing resources includes:

first circuitry including a weight ternarization circuit and an activation quantization circuit, the weight ternarization circuit to convert a weight tensor from a floating-point representation to a ternary representation including a two-bit ternary weight and an eight-bit integer scale factor, and the two-bit ternary weight to represent a weight value of one of negative one, zero, or positive one and wherein the activation quantization circuit is to convert an activation tensor from a floating-point representation to an integer representation; and

second circuitry configured to perform a set of parallel integer compute operations on the ternary representation of the weight tensor and the integer representation of the activation tensor.

18. The data processing system as in claim 17 , the second circuitry including one or more ternary logic units to perform one or more operations in the set of parallel integer compute operations.

19. The data processing system as in claim 18 , wherein the second circuitry is configured to perform the set of parallel integer compute operations in response to an instruction provided by a machine learning inferencing framework.

20. The data processing system as in claim 17 , the activation quantization circuit to convert the activation tensor from a single-precision floating-point representation to an eight-bit integer representation.

Priority Claims (1)
IN 201741015052 · Apr 28, 2017 · national
Continuity (4)
Continuation 18060414 · Nov 30, 2022
Continuation 15869515 · Jan 12, 2018
Provisional Application 62501800 · May 5, 2017
Related Publication 20240160931A1 · May 16, 2024
References Cited (18)
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10643126B2 · Saldana · 2020 [cited by examiner]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
Li, Fengfu, et al. “Ternary weight networks.” arXiv preprint arXiv:1605.04711 (2016). (Year: 2016). [cited by examiner]
Mark Z. Mao. “Improving the speed of neural networks on CPUs.” Proc. deep learning and unsupervised feature learning NIPS workshop. vol. 1. No. 2011. 2011. (Year: 2011). [cited by examiner]
Dally et al. “Trained ternary quantization.” arXiv preprint arXiv:1612.01064 (2016). (Year: 2016). [cited by examiner]
Ni et al. “Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients.” arXiv preprint arXiv:1606.06160 (2016). (Year: 2016). [cited by examiner]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Stephen Junkins, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/869,515, mailed Sep. 20, 2022, 8 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/060,414, mailed Sep. 14, 2023, 8 pages. [cited by applicant]