IP Library Granted Patent US 12,596,922
Granted Patent B2
US 12,596,922 · App. 18/612,881 · Granted Apr 7, 2026

Accelerating neural networks in hardware using interconnected crossbars

Inventors: Pierre-Luc Cantin (Palo Alto, CA); Olivier Temam (Antony, FR)
Assignee: Google LLC
G06N3/065G06F9/5027G06F17/16G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,922
App. No.
18/612,881
Granted
Apr 7, 2026
Kind
B2
Abstract

A computing unit for accelerating a neural network is disclosed. The computing unit include an input unit that includes a digital-to-analog conversion unit and an analog-to-digital conversion unit that is configured to receive an analog signal from the output of a last interconnected analog crossbar circuit of a plurality of analog crossbar circuits and convert the second analog signal into a digital output vector, and a plurality of interconnected analog crossbar circuits that include the first interconnected analog crossbar circuit and the last interconnected crossbar circuits, wherein a second interconnected analog crossbar circuit of the plurality of interconnected analog crossbar circuits is configured to receive a third analog signal from another interconnected analog crossbar circuit of the plurality of interconnected crossbar circuits and perform one or more operations on the third analog signal based on the matrix weights stored by the crosspoints of the second interconnected analog crossbar.

Claims (38)

1 . A computing unit for accelerating a neural network, comprising:

a plurality of interconnected analog crossbar circuits comprising a first analog crossbar circuit, a last analog crossbar circuit, and a set of intermediary analog crossbar circuits between the first analog crossbar circuit and the last analog crossbar circuit, wherein:

each interconnected analog crossbar circuit corresponds to a respective layer of the neural network and includes a plurality of crosspoints, each crosspoint storing a weight of a plurality of weights associated with the respective layer of the neural network,

wherein each analog crossbar circuit in the set of intermediary analog circuits receives an analog input from a preceding analog crossbar circuit and provides an analog output to a subsequent analog crossbar circuit;

a single digital-to-analog conversion (DAC) unit, wherein the single DAC unit is configured to convert a digital input vector into a first analog signal that is provided as an input to the first analog crossbar circuit; and

a single analog-to-digital conversion (ADC) unit, wherein the single ADC unit is configured to receive as an input a second analog signal from an output of the last analog crossbar circuit and convert the second analog signal into a digital output vector.

2 . The computing unit of claim 1 , wherein at least one analog crossbar circuit in the plurality of interconnected analog crossbar circuits is configured to perform matrix multiplication operations on an analog signal received by the at least one analog crossbar circuit based on the weights stored by the crosspoints of the at least one analog crossbar circuit.

3 . The computing unit of claim 1 , wherein the neural network is a fully-connected neural network.

4 . The computing unit of claim 1 , wherein each crossbar circuit of the plurality of interconnected analog crossbar circuits other than the first crossbar circuit is configured to receive as input an analog output generated by a preceding analog crossbar circuit in the plurality of interconnected analog crossbar circuits.

5 . The computing unit of claim 1 , further comprising:

at least one array of analog signal amplifiers that is positioned between a pair of analog crossbar circuits in the plurality of interconnected analog crossbar circuits.

6 . The computing unit of claim 5 , wherein the at least one array of analog signal amplifiers is configured to (i) receive as an input an analog output generated by a second analog crossbar circuit and (ii) generate as an analog output for use as an input to a third analog crossbar circuit.

7 . The computing unit of claim 5 , wherein each crossbar circuit of the plurality of interconnected analog crossbar circuits other than the first crossbar circuit is configured to receive as an input (i) an analog output generated by a preceding analog crossbar circuit or (ii) an analog output generated by the at least one array of analog signal amplifiers.

8 . A method of configuring a computing unit for accelerating a neural network, the method comprising:

configuring a plurality of interconnected analog crossbar circuits comprising a first analog crossbar circuit, a last analog crossbar circuit, and a set of intermediary analog crossbar circuits between the first analog crossbar circuit and the last analog crossbar circuit, wherein:

each interconnected analog crossbar circuit corresponds to a respective layer of the neural network and includes a plurality of crosspoints, each crosspoint storing a weight of a plurality of weights associated with the respective layer of the neural network,

wherein each analog crossbar circuit in the set of intermediary analog circuits receives an analog input from a preceding analog crossbar circuit and provides an analog output to a subsequent analog crossbar circuit;

configuring a single digital-to-analog conversion (DAC) unit to convert a digital input vector into a first analog signal that is provided as an input to the first analog crossbar circuit; and

configuring a single analog-to-digital conversion (ADC) unit to receive as an input a second analog signal from an output of the last analog crossbar circuit and convert the second analog signal into a digital output vector.

9 . The method of claim 8 , wherein at least one analog crossbar circuit in the plurality of interconnected analog crossbar circuits is configured to perform matrix multiplication operations on an analog signal received by the at least one analog crossbar circuit based on the weights stored by the crosspoints of the at least one analog crossbar circuit.

10 . The method of claim 8 , wherein the neural network is a fully-connected neural network.

11 . The method of claim 8 , wherein each crossbar circuit of the plurality of interconnected analog crossbar circuits other than the first crossbar circuit is configured to receive as input an analog output generated by a preceding analog crossbar circuit in the plurality of interconnected analog crossbar circuits.

12 . The method of claim 8 , further comprising:

at least one array of analog signal amplifiers that is positioned between a pair of analog crossbar circuits in the plurality of interconnected analog crossbar circuits.

13 . The method of claim 12 , wherein the at least one array of analog signal amplifiers is configured to (i) receive as an input an analog output generated by a second analog crossbar circuit and (ii) generate as an analog output for use as an input to a third analog crossbar circuit.

14 . The method of claim 12 , wherein each crossbar circuit of the plurality of interconnected analog crossbar circuits other than the first crossbar circuit is configured to receive as an input (i) an analog output generated by a preceding analog crossbar circuit or (ii) an analog output generated by the at least one array of analog signal amplifiers.

15 . A non-transitory computer storage medium encoded with a computer program, the program comprising instructions that are operable, when executed by a data processing apparatus, to cause the data processing apparatus to perform the operations comprising:

configuring a plurality of interconnected analog crossbar circuits comprising a first analog crossbar circuit, a last analog crossbar circuit, and a set of intermediary analog crossbar circuits between the first analog crossbar circuit and the last analog crossbar circuit, wherein:

each interconnected analog crossbar circuit corresponds to a respective layer of a neural network and includes a plurality of crosspoints, each crosspoint storing a weight of a plurality of weights associated with the respective layer of the neural network,

wherein each analog crossbar circuit in the set of intermediary analog circuits receives an analog input from a preceding analog crossbar circuit and provides an analog output to a subsequent analog crossbar circuit;

configuring a single digital-to-analog conversion (DAC) unit to convert a digital input vector into a first analog signal that is provided as an input to the first analog crossbar circuit; and

configuring a single analog-to-digital conversion (ADC) unit to receive as an input a second analog signal from an output of the last analog crossbar circuit and convert the second analog signal into a digital output vector.

16 . The media of claim 15 , wherein at least one analog crossbar circuit in the plurality of interconnected analog crossbar circuits is configured to perform matrix multiplication operations on an analog signal received by the at least one analog crossbar circuit based on the weights stored by the crosspoints of the at least one analog crossbar circuit.

17 . The media of claim 15 , wherein the neural network is a fully-connected neural network.

18 . The media of claim 15 , wherein each crossbar circuit of the plurality of interconnected analog crossbar circuits other than the first crossbar circuit is configured to receive as input an analog output generated by a preceding analog crossbar circuit in the plurality of interconnected analog crossbar circuits.

19 . The media of claim 15 , further comprising:

at least one array of analog signal amplifiers that is positioned between a pair of analog crossbar circuits in the plurality of interconnected analog crossbar circuits.

20 . The media of claim 19 , wherein the at least one array of analog signal amplifiers is configured to (i) receive as an input an analog output generated by a second analog crossbar circuit and (ii) generate as an analog output for use as an input to a third analog crossbar circuit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2024
From: CANTIN, PIERRE-LUC; TEMAM, OLIVIER
To: GOOGLE LLC
Reel/Frame 066872/0933 →
Continuity (3)
Continuation 16059578 · Aug 9, 2018
Provisional Application 62543251 · Aug 9, 2017
Related Publication 20240232603A1 · Jul 11, 2024
References Cited (31)
US 7302513B2 · Mouttet · 2007 [cited by examiner]
US 8812415B2 · Modha et al. · 2014 [cited by applicant]
US 20170109628A1 · Gokmen · 2017 [cited by applicant]
US 20170228345A1 · Gupta et al. · 2017 [cited by applicant]
US 20180005115A1 · Goknnen · 2018 [cited by applicant]
US 20180075338A1 · Goknnen · 2018 [cited by applicant]
US 20180075339A1 · Ma · 2018 [cited by examiner]
US 20180114569A1 · Strachan et al. · 2018 [cited by applicant]
US 20200159811A1 · Chatterjee · 2020 [cited by examiner]
EP 1115204 · 2009 [cited by applicant]
TW 201714120 · 2017 [cited by applicant]
TW 201723934 · 2017 [cited by applicant]
Continuous Real-World Inputs Can Open Up Alternative Accelerator Designs (Year: 2013). [cited by examiner]
Memristive_Boltzmann_machine_A_hardware_accelerator_for_combinatorial_optimization_and_deep_learning (Year: 2016). [cited by examiner]
ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars (Year: 2016). [cited by examiner]
On-chip Training of Memristor Based Deep Neural Networks (Year: 2017). [cited by examiner]
Hu—Memristor Crossbar-Based Neuromorphic Computing System a Case Study (Year: 2014). [cited by examiner]
Li—Emerging memristor technology enabled next generation cortical processor (Year: 2014). [cited by examiner]
Liu—Hardware Acceleration for Neuromorphic Computing (Year: 2015). [cited by examiner]
Tang—Spiking Neural Network with RRAM (Year: 2015). [cited by examiner]
Esteves et al., “Polar Transformer Networks,” CoRR, Submitted on Feb. 1, 2018, arXiv:1709.01889v3, 14 pages. [cited by applicant]
Hasan et al., “On-chip Training of Memristor Based Deep Neural Networks,” Presented at International Joint Conference on Neural Networks, Anchorage AK, May 14-19, 2017, 10 pages. [cited by applicant]
Hu et al., “Dot-Product Engine for Neuromorphic Computing: Programming 1T1M Crossbar to Accelerate Matrix-Vector Multiplication,” Proceedings of the 53rd Annual Design Automation Conference on DAC, Jun. 5, 2016, 6 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2018/046069, mailed on Nov. 14, 2018, 15 pages. [cited by applicant]
Li et al. “RRAM-Based Analog Approximate Computing,” IEEE Transactions on Computer Aided Design of Integrated Circuits and Systems, IEEE, Dec. 1, 2015, 34(12):13 pages. [cited by applicant]
Li et al., “Dense Transformer Networks,” CoRR, Submitted on Jun. 8, 2017, arXiv:1705.08881v2, 10 pages. [cited by applicant]
Li et al., “Training itself: Mixed-signal training acceleration for memristor-based neural network,” 19th Asia and South Pacific Design Automation Conference, Jan. 20, 2014, 6 pages. [cited by applicant]
Lin et al., “Inverse Compositional Spatial Transformer Networks,” CoRR, Submitted on Dec. 12, 2016, arXiv:1612.03897v1, 9 pages. [cited by applicant]
Office Action in Taiwan Application No. 107127883, mailed on Jun. 21, 2019, 11 pages (with English translation). [cited by applicant]
Shafiee et al., “ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars,” ACM SIGARCH Computer Architecture News, Oct. 12, 2016, 44(3):13 pages. [cited by applicant]
Tang et al., “Binary Convolutional Neural Network on RRAM,” Presented at 22nd Asia and South Pacific Design Automation Conference, Chiba, Japan, Jan. 16-19, 2017, pp. 782-787. [cited by applicant]