IP Library › Granted Patent US 12,613,706
Granted Patent B2
US 12,613,706 · App. 18/380,620 · Granted Apr 28, 2026

Hardware accelerated machine learning

Inventors: Jeremy Bruestle (Seattle, WA); Choong Ng (Seattle, WA)
Assignee: Intel Corporation
G06N3/08G06F5/01G06F7/023G06N20/00G06F2205/00G06F2207/4824G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,613,706
App. No.
18/380,620
Granted
Apr 28, 2026
Kind
B2
Abstract

A machine learning hardware accelerator architecture and associated techniques are disclosed. The architecture features multiple memory banks of very wide SRAM that may be concurrently accessed by a large number of parallel operational units. Each operational unit supports an instruction set specific to machine learning, including optimizations for performing tensor operations and convolutions. Optimized addressing, an optimized shift reader and variations on a multicast network that permutes and copies data and associates with an operational unit that support those operations are also disclosed.

Claims (19)

1 . An apparatus comprising:

a hardware accelerator, comprising:

a plurality of circuitries to execute instructions associated with machine learning operations;

a local memory comprising a plurality of dual-ported static random-access memory (SRAM), banks to store data associated with one or more of the instructions;

circuitry to permute a first plurality of source data elements associated with a first multidimensional array in accordance with a permutation pattern included with a matrix multiplication instruction and route the first plurality of source data elements to the plurality of circuitries;

the plurality of circuitries, coupled to the circuitry to obtain the first plurality of source data elements, to perform a plurality of parallel multiply-accumulate (MAC) operations in accordance with the matrix multiplication instruction, each of the plurality of parallel MAC operations comprising:

multiplying a first source data element of the first plurality of source data elements provided by the circuitry to permute and a second source data element of a second plurality of source data elements associated with a second multidimensional array to generate a product, and

adding the product to an accumulation value to generate a result value, the first source data element and the second source data element each having a first bit width and the accumulation value having a second bit width at least twice the first bit width.

2 . The apparatus of claim 1 wherein the circuitry to permute comprises a plurality of configurable switches.

3 . The apparatus of claim 1 wherein the circuitry to permute is to perform a broadcast operation to route a single source data element to each of the plurality of circuitries.

4 . The apparatus of claim 1 wherein to perform the permutation, index values are to be generated by modifying initial index values in accordance with the permutation.

5 . The apparatus of claim 1 wherein the accumulation value having the second bit width is to be stored by an accumulation register.

6 . The apparatus of claim 1 wherein the first bit width is 16 bits and the second bit width is 32 bits.

7 . The apparatus of claim 6 wherein the first and second plurality of source data elements comprise 16-bit floating-point data elements and the accumulation and result values comprise 32-bit floating point data elements.

8 . The apparatus of claim 1 further comprising:

a cache shared by at least some of the plurality of circuitries.

9 . The apparatus of claim 8 wherein the cache comprises a write though cache.

10 . The apparatus of claim 1 wherein the circuitry to permute is to perform the permutation operation based, at least in part, on information included in at least one of the instructions.

11 . The apparatus of claim 10 wherein the circuitry to permute is to perform the permutation operation based, at least in part, on information included in at least one register.

Continuity (4)
Continuation 17501314 · Oct 14, 2021
Continuation 15399714 · Jan 5, 2017
Provisional Application 62276169 · Jan 7, 2016
Related Publication 20240046088A1 · Feb 8, 2024
References Cited (73)
US 5138695A · Means et al. · 1992 [cited by applicant]
US 5625825A · Rostoker et al. · 1997 [cited by applicant]
US 5751987A · Mahant-Shetti et al. · 1998 [cited by applicant]
US 5892697A · Brakefield · 1999 [cited by applicant]
US 6216167B1 · Momirov · 2001 [cited by applicant]
US 6285779B1 · Lapidous et al. · 2001 [cited by applicant]
US 6571268B1 · Giacalone et al. · 2003 [cited by applicant]
US 6768992B1 · Jolitz · 2004 [cited by applicant]
US 9747547B2 · Mccormick et al. · 2017 [cited by applicant]
US 20020062466A1 · Noguchi · 2002 [cited by applicant]
US 20020075871A1 · Blanc et al. · 2002 [cited by applicant]
US 20020126661A1 · Ngai · 2002 [cited by applicant]
US 20030206630A1 · Rarick · 2003 [cited by applicant]
US 20040078418A1 · Law et al. · 2004 [cited by applicant]
US 20060259744A1 · Matthes · 2006 [cited by applicant]
US 20070005322A1 · Patzer · 2007 [cited by examiner]
US 20070211064A1 · Buck et al. · 2007 [cited by applicant]
US 20090187746A1 · Symes et al. · 2009 [cited by applicant]
US 20090313195A1 · Mcdaid et al. · 2009 [cited by applicant]
US 20100005221A1 · Nieminen · 2010 [cited by examiner]
US 20100076915A1 · Xu et al. · 2010 [cited by applicant]
US 20110029471A1 · Chakradhar et al. · 2011 [cited by applicant]
US 20110206053A1 · Henry et al. · 2011 [cited by applicant]
US 20120005141A1 · Sasagawa · 2012 [cited by applicant]
US 20130054665A1 · Felch · 2013 [cited by applicant]
US 20140040700A1 · Kobori et al. · 2014 [cited by applicant]
US 20140136583A1 · Hyde et al. · 2014 [cited by applicant]
US 20140188968A1 · Kaul et al. · 2014 [cited by applicant]
US 20140344194A1 · Lee et al. · 2014 [cited by applicant]
US 20150199963A1 · Maaninen · 2015 [cited by applicant]
US 20150324685A1 · Bohn et al. · 2015 [cited by applicant]
US 20160344629A1 · Gray · 2016 [cited by examiner]
US 20160379137A1 · Burger et al. · 2016 [cited by applicant]
WO 2006115896A2 · 2006 [cited by applicant]
Chen et al, “DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning”, 2014, ACM SIGARCH Computer Architecture News, vol. 42, Issue 1, pp. 269-284. (Year: 2014). [cited by examiner]
Lozito et al, “FPGA Implementations of Feed Forward Neural Network by using Floating Point Hardware Accelerators”, 2014, Advances in Electrical and Electronic Engineering: Theoretical and Applied Electrical Engineering,… [cited by examiner]
“International Search Report and Written Opinion” for PCT Application No. PCT/US2017/012600 mailed Mar. 27, 2017, 7 pages. [cited by applicant]
Chen et al., “DianNao: A small-footprint high-throughput accelerator for ubiquitous machine-learning”, ACM SIGARCH Computer Architecture News, vol. 42, No. 1, Feb. 2014, pp. 269-284. [cited by applicant]
Communication under Rule 71(3) EPC, EP App. No. 17736464, Sep. 15, 2020, 6 pages. [cited by applicant]
Decision to grant a European patent, EP App. No. 17736464, Jan. 28, 2021, 2 pages. [cited by applicant]
Decision to grant a European patent, EP App. No. 21158563, Oct. 20, 2022, 2 pages. [cited by applicant]
Decision to grant, EP App. No. 21208402.4, May 11, 2023, 2 pages. [cited by applicant]
Decision to grant, EP App. No. 23163158.1, Jul. 18, 2024, 2 pages. [cited by applicant]
European Patent Office, “Invitation pursuant to Rule 63(1) EPC,” issued in connection with European Patent Application No. 21158563.3, mailed on May 18, 2021, 6 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 21158563.3, Sep. 9, 2021, 6 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 21208402.4, Feb. 28, 2022, 5 pages. [cited by applicant]
European search report and Search Opinion, PCT App. No.23163158.1, Jun. 30, 2023, 5 pages. [cited by applicant]
European Search Report from related application EP 17796609 dated Dec. 13, 2019. [cited by applicant]
European Search Report from related application EP 17796609 dated Apr. 2, 2020. [cited by applicant]
European Search Report from related application EP 17796610 dated Dec. 13, 2019. [cited by applicant]
Farabet et al., “NeuFlow: A runtime reconfigurable dataflow processor for vision”, 2011 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, CVPRW 2011, Jun. 2011, pp. 109-116. [cited by applicant]
Final Office Action, U.S. Appl. No. 15/399,714, Jan. 2, 2020, 35 pages. [cited by applicant]
Intention to grant European patent EP App. No. 21158563, Jun. 9, 2022, 6 pages. [cited by applicant]
Intention to Grant, EP App. No. 21208402.4, Jan. 5, 2023, 6 pages. [cited by applicant]
Intention to grant, EP App. No. 23163158.1, Mar. 13, 2024, 6 pages. [cited by applicant]
International Preliminary Report on Patentability, PCT App. No. PCT/US17/012600, Jul. 19, 2018, 6 pages. [cited by applicant]
International Search Report and Written Opinion dated Oct. 2, 2017, for PCT Application No. PCT/US2017/031477, 10 pages (1L.P0004PCT). [cited by applicant]
International Search Report and Written Opinion dated Oct. 24, 2017, for PCT Application No. PCT/US2017/031478, 10 pages (1L.P0005PCT). [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US17/12600, Mar. 27, 2017, 6 pages. [cited by applicant]
Jindal et al., “Exploiting GPUs with the Super Instruction Architecture”, International journal of parallel programming, vol. 44, No. 2, Aug. 20, 2014, pp. 309-324. [cited by applicant]
Jones et al., “Learning in Linear Systolic Neural Network Engines: Analysis and Implementation”, IEEE Transactions on Neural Networks, Jul. 1, 1993. [cited by applicant]
Kanoun et al., “Low power and scalable many-core architecture for big-data stream computing”, 2014 IEEE Computer Society Annual Symposium on VLSI, Jul. 2014, pp. 468-473. [cited by applicant]
Lozito et al., “FPGA Implementations of Feed Forward Neural Network by Using Floating Point Hardware Accelerators,” Theoretical and Applied Electrical Engineering vol. 12, No. 1, Mar. 2014 (http://advances.utc.sk/index.… [cited by applicant]
Lu et al., “Empirical performance model-driven data layout optimization and library call selection for tensor contraction expressions”, Journal of Parallel and Distributed Computing, vol. 72, No. 3, Mar. 2012, pp. 338-3… [cited by applicant]
Minkenberg, “On packet switch design”, Eindoven University of Technology, Jan. 1, 2001. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 15/399,714, Dec. 16, 2020, 31 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 15/399,714, Jul. 25, 2019, 31 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 15/399,714, Jul. 8, 2021, 13 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/501,314, Jul. 19, 2023, 16 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/501,314, Sep. 11, 2023, 7 pages. [cited by applicant]
Supplementary European Search Report and Search Opinion, EP App. No. 17736464.3, Jun. 21, 2019, 9 pages. [cited by applicant]
Tianshi Chen et al: “DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning”, Mar. 1, 2014 (Mar. 1, 2014), ACM SIGPLAN Notices, v.49 n.4, pp. 269-284. (Year: 2014). [cited by applicant]
Zuras et al: “IEEE Standard for Floating-Point Arithmetic”, Jun. 12, 2008, IEEE, v.754, pp. 1-52. (Year: 2008). [cited by applicant]