IP Library Granted Patent US 12,430,131
Granted Patent B2
US 12,430,131 · App. 18/315,625 · Granted Sep 30, 2025

Compute optimizations for neural networks

Inventors: Kevin Nealis (San Jose, CA); Anbang Yao (Beijing, CN); Xiaoming Chen (Shanghai, CN); Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Sara S. Baghsorkhi (San Jose, CA); Eriko Nurvitadhi (Hillsboro, OR); Balaji Vembu (Folsom, CA); Nicolas C. Galoppo Von Borries (Portland, OR); Rajkishore Barik (Santa Clara, CA); Tsung-Han Lin (Campbell, CA); Kamal Sinha (Cordova, CA)
Assignee: Intel Corporation
G06F9/3001G06F9/3851G06F9/3887G06F9/3888G06F9/3893G06N3/044G06N3/045G06N3/063G06N3/084G06T1/20G06F2207/4824
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,131
App. No.
18/315,625
Granted
Sep 30, 2025
Kind
B2
Abstract

One embodiment provides for a compute apparatus comprising a decode unit to decode a single instruction into a decoded instruction that specifies multiple operands including a multi-bit input value and a one-bit weight associated with a neural network, as well as an arithmetic logic unit including a multiplier, an adder, and an accumulator register. To execute the decoded instruction, the multiplier is to perform a fused operation including an exclusive not OR (XNOR) operation and a population count operation. The adder is configured to add the intermediate product to a value stored in the accumulator register and update the value stored in the accumulator register.

Claims (29)

1. A compute apparatus comprising:

a decode unit to decode a single instruction into a decoded instruction that specifies multiple operands including a multi-bit input value and a one-bit weight associated with a neural network; and

an arithmetic logic unit including a multiplier, an adder, and an accumulator register, wherein to execute the decoded instruction, the multiplier is to perform a multiplication operation on the multi-bit input value based on the one-bit weight to generate an intermediate product, the adder is to add the intermediate product to a value stored in the accumulator register and update the value stored in the accumulator register, and the multiplication operation comprises a fused operation including an exclusive not OR (XNOR) operation and a population count operation.

2. The compute apparatus as in claim 1 , wherein the multiplier includes first circuitry to perform an XNOR operation and second circuitry to perform a population count operation on output of the first circuitry.

3. The compute apparatus as in claim 2 , wherein the multiplier includes an intermediate register to store output from the second circuitry.

4. The compute apparatus as in claim 1 , wherein the one-bit weight is a bipolar binary weight that represents a weight value of one of positive one and negative one.

5. The compute apparatus as in claim 4 , wherein the bipolar binary weight represents a weight value of negative one as a binary zero.

6. The compute apparatus as in claim 4 , wherein the value of the bipolar binary weight is referenced via an index into a multi-bit register.

7. The compute apparatus as in claim 1 , additionally including an output register to store an output value of the single instruction.

8. The compute apparatus as in claim 1 , wherein the multi-bit input value has a power of two number of bits.

9. A method comprising:

decoding a single instruction specifying multiple operands, the multiple operands including a multi-bit input value and a one-bit weight associated with a neural network;

issuing the single instruction for execution within a compute unit of a general-purpose graphics processing unit; and

responsive to the execution of the single instruction, generating a result by performing a multiplication operation on the multi-bit input value based on the one-bit weight to generate an intermediate product and updating a value stored in an accumulator register by adding the intermediate product to the value stored in the accumulator register, the multiplication operation comprising a fused operation including an exclusive not OR (XNOR) operation and a population count operation.

10. The method as in claim 9 , wherein the multiplier includes first circuitry to perform an XNOR operation and second circuitry to perform a population count operation on output of the first circuitry.

11. The method as in claim 10 , wherein the multiplier includes an intermediate register and the method additionally comprising storing output from the second circuitry to the intermediate register.

12. The method as in claim 9 , wherein the one-bit weight is a bipolar binary weight that represents a weight value of one of positive one and negative one.

13. The method as in claim 12 , wherein the bipolar binary weight represents a weight value of negative one as a binary zero.

14. The method as in claim 13 , wherein the value of the bipolar binary weight is referenced via an index into a multi-bit register.

15. The method as in claim 13 , wherein the multi-bit input value has a power of two number of bits.

16. A data processing system comprising:

a general-purpose graphics processing unit comprising:

a decode unit to decode a single instruction into a decoded instruction that specifies multiple operands including a multi-bit input value and a one-bit weight associated with a neural network, and

an arithmetic logic unit including a multiplier, an adder, and an accumulator register, wherein to execute the decoded instruction, the multiplier is to perform a multiplication operation on the multi-bit input value based on the one-bit weight to generate an intermediate product, the adder is to add the intermediate product to a value stored in the accumulator register and update the value stored in the accumulator register, and the multiplication operation comprises a fused operation including an exclusive not OR (XNOR) operation and a population count operation; and

a memory coupled with the general-purpose graphics processing unit.

17. The data processing system as in claim 16 , wherein the multiplier includes first circuitry to perform an XNOR operation and second circuitry to perform a population count operation on output of the first circuitry.

18. The data processing system as in claim 17 , wherein the multiplier includes an intermediate register to store output from the second circuitry.

19. The data processing system as in claim 16 , wherein the one-bit weight is a bipolar binary weight that represents a weight value of one of positive one and negative one, the bipolar binary weight represents a weight value of negative one as a binary zero, and the value of the bipolar binary weight is referenced via an index into a multi-bit register.

20. The data processing system as in claim 19 , wherein the multi-bit input value has a power of two number of bits.

Continuity (4)
Continuation 17443376 · Jul 26, 2021
Continuation 16505012 · Jul 8, 2019
Continuation 15494710 · Apr 24, 2017
Related Publication 20230359461A1 · Nov 9, 2023
References Cited (37)
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 11049006B2 · Langford et al. · 2021 [cited by applicant]
US 20150106414A1 · Olsen · 2015 [cited by applicant]
US 20160026912A1 · Falcon et al. · 2016 [cited by applicant]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20160093311A1 · Kim · 2016 [cited by applicant]
US 20170103304A1 · Henry · 2017 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180225116A1 · Henry · 2018 [cited by applicant]
US 20180307950A1 · Nealis · 2018 [cited by applicant]
CN 105320495 · 2016 [cited by applicant]
CN 106062786 · 2016 [cited by applicant]
DE 4228767A1 · 1993 [cited by applicant]
EP 3396535A1 · 2018 [cited by applicant]
EP 3671439A1 · 2020 [cited by applicant]
Intention to Grant for EP Application No. 20155873.1, mailed Nov. 16, 2023, 8 pages. [cited by applicant]
Extended European Search Report for Application No. 18163718.2-1221, mailed Sep. 21, 2018, 10 pages. [cited by applicant]
Soheil Hashemi et al: “Understanding the Impact of Precision Quantization on the Accuracy and Energy of Neural Networks”, Arxiv.org, Cornell University Library, Dec. 12, 2016, XP080738795. [cited by applicant]
Tianshi Chen et al: “DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning”, Mar. 1, 2014, XP055422053, Retrieved from the Internet: URL: http://novel.ict.ac.cn/ychen/pdf/DianNao.pdf [re… [cited by applicant]
Chuan Zhang Tang et al: “Multilayer Feedforward Neural Networks with Single Powers-of-Two Weights”, IEEE Transactions on Signal Processing, vol. 41, nr. 8, Aug. 1, 1993, pp. 484-487, XP55503286, Retrieved from the Inter… [cited by applicant]
Zhouhan Lin et al: “Neural Networks with 1-15 Few Multiplications”, CORR (ARXIV), vol. 1510.03009v3, Feb. 26, 2016, pp. 1-9, XP055414215. [cited by applicant]
Aojun Zhou et al: “Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights”, Arxiv.org, Cornell University Library, Feb. 10, 2017, XP080747349. [cited by applicant]
Uros Lotric et al: “Applicability of approximate multipliers in hardware neural networks”, Neurocomputing, vol. 96, Nov. 1, 2012, pp. 57-65. [cited by applicant]
Notification of Publication of CN Application No. 108734285, Nov. 2, 2018, 5 pages. [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Stephen Junkins, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
EP Allowance for Application No. 18163718.2-1221, mailed Nov. 12, 2019, 8 pages. [cited by applicant]
Extended European Search Report for EP20155873.1, 9 pages, May 28, 2020. [cited by applicant]
Renzo Andri et al., “YodaNN: An Architecture for Ultra-Low Power Binary-Weight CNN Acceleration”, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Feb. 24, 2017, XP055690573. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/443,376 mailed Feb. 24, 2023, 7 pages. [cited by applicant]
Notice of Grant of Patent Right for CN Application No. 201810367363.7, Mar. 13, 25, 4 pages. [cited by applicant]