IP Library Granted Patent US 12,316,318
Granted Patent B2
US 12,316,318 · App. 17/818,924 · Granted May 27, 2025

Multiplier-accumulator circuitry, and processing pipeline including same

Inventor: Cheng C. Wang (San Jose, CA)
Assignee: Analog Devices, Inc.
H03K19/1776H03K19/17724
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,316,318
App. No.
17/818,924
Granted
May 27, 2025
Kind
B2
Abstract

An integrated circuit comprising a plurality of MACs, connected to form a pipeline, to perform a plurality of multiply and accumulate operations, wherein each MAC includes: (A) a multiplier, coupled to memory to (i) receive the multiplier weight data, (ii) multiply first data and the multiplier weight data and (iii) output product data, (B) an accumulator, coupled to the multiplier of the MAC, to add second data and the first product data and output sum data, and (C) a load-store register, coupled to: (i) an output of the accumulator of the associated MAC and (ii) an input of the load-store register of an immediately successive MAC. Each load-store register may include two interconnected registers, and is configurable to, on the same clock cycle, (a) load the initialization data into the accumulator of the immediately successive MAC and (b) store the sum data from the associated MAC into the load-store register.

Claims (74)

1. An integrated circuit comprising:

first memory to store a plurality of multiplier weight data including multiplier weight data;

a plurality of multiplier-accumulator circuits, connected to form a pipeline, to perform a plurality of multiply and accumulate operations, wherein each multiplier-accumulator circuit includes:

a multiplier, coupled to the first memory to (i) receive the multiplier weight data, (ii) multiply first data and the multiplier weight data and (iii) output product data, and

an accumulator, coupled to the multiplier of the multiplier-accumulator circuit, to add second data and the first product data and output sum data; and

a load-store register, coupled to: (i) an output of the accumulator of the associated multiplier-accumulator circuit and (ii) an input of the load-store register of an immediately successive multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of the pipeline via an electrical path that bypasses an adder of the immediately successive multiplier-accumulator circuit.

2. The integrated circuit of claim 1 wherein:

the load-store register of each multiplier-accumulator circuit is configurable to: (i) temporarily store initialization data and (ii) load the initialization data into the accumulator of the immediately successive multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of the pipeline.

3. The integrated circuit of claim 2 wherein:

the load-store register of each multiplier-accumulator circuit is configurable to: (i) temporarily store initialization data, (ii) load the initialization data into the accumulator of the immediately successive multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of the pipeline, and (iii) store the data from the associated multiplier-accumulator circuit into the load-store register.

4. The integrated circuit of claim 3 wherein:

each load-store register:

includes two interconnected registers, and

is configurable to, on the same clock cycle, (a) load the initialization data into the accumulator of the immediately successive multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of the pipeline and (b) store the sum data from the associated multiplier-accumulator circuit into the load-store register, wherein data in the two interconnected registers are swapped therebetween.

5. The integrated circuit of claim 1 wherein:

each load-store register includes two interconnected registers.

6. The integrated circuit of claim 5 wherein:

each load-store register further includes two multiplexers.

7. The integrated circuit of claim 1 wherein the load-store register includes:

two interconnected registers, and

two multiplexers, wherein each multiplexer includes (i) an output connected to an input of one of the interconnected registers and (ii) a first input connected to the output of the other interconnected register.

8. The integrated circuit of claim 1 wherein:

each multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits further includes second memory, coupled to the first memory, to receive multiplier weight data from the first memory and store multiplier weight data for use by the multiplier of the multiplier-accumulator circuit.

9. The integrated circuit of claim 8 wherein:

the second memory is ROM and, in operation, the multiplier weight data is read from the ROM and output to the multiplier of the multiplier-accumulator circuit.

10. The integrated circuit of claim 1 wherein:

each multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits further includes a by-pass circuit, coupled between an output of the accumulator of an immediately preceding multiplier-accumulator circuit and the load-store register of the multiplier-accumulator circuit.

11. An integrated circuit comprising:

memory to store a plurality of multiplier weight data including multiplier weight data;

a plurality of multiplier-accumulator circuits, connected to form a pipeline, to perform a plurality of multiply and accumulate operations, wherein each multiplier-accumulator circuit includes:

a multiplier, coupled to the memory to (i) receive the multiplier weight data, (ii) multiply first data and the multiplier weight data and (iii) output product data, and

an accumulator, coupled to the multiplier of the multiplier-accumulator circuit, to add second data and the first product data and output sum data, and

a load-store register, coupled to: (i) an output of the accumulator of the associated multiplier-accumulator circuit and (ii) an input of the load-store register of an immediately successive multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of the pipeline; and

input selection circuit, connected to a first multiplier-accumulator circuit of the pipeline and a last multiplier-accumulator circuit of the pipeline, wherein the input selection circuit is configurable to connect the last multiplier-accumulator circuit to the first multiplier-accumulator circuit to form a ring.

12. The integrated circuit of claim 11 wherein:

the input selection circuit includes a multiplexer.

13. The integrated circuit of claim 11 wherein:

input selection circuit is configurable to disconnect the last multiplier-accumulator circuit from the first multiplier-accumulator circuit.

14. The integrated circuit of claim 11 wherein:

input selection circuit is configurable to receive sum data generated via a different pipeline of multiplier-accumulator circuits.

15. The integrated circuit of claim 11 wherein:

the load-store register of each multiplier-accumulator circuit is configurable to: (i) temporarily store initialization data and (ii) load the initialization data into the accumulator of the immediately successive multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of the pipeline.

16. The integrated circuit of claim 11 wherein:

the load-store register of each multiplier-accumulator circuit is configurable to: (i) temporarily store initialization data, (ii) load the initialization data into the accumulator of the immediately successive multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of the pipeline, and (iii) store the sum data from the associated multiplier-accumulator circuit into the load-store register.

17. The integrated circuit of claim 16 wherein:

each load-store register:

includes two interconnected registers, and

is configurable to, on the same clock cycle, (a) load the initialization data into the accumulator of the immediately successive multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of the pipeline, and (b) store the sum data from the associated multiplier-accumulator circuit into the load-store register, wherein the data in the two interconnected registers are swapped therebetween.

18. The integrated circuit of claim 11 wherein:

each load-store register includes two interconnected registers.

19. The integrated circuit of claim 11 wherein:

each load-store register further includes two multiplexers.

20. The integrated circuit of claim 11 wherein:

the load-store register includes two interconnected registers; and

the load-store register further includes two multiplexers, wherein each multiplexer includes: (i) an output connected to an input of one of the interconnected registers and (ii) a first input connected to the output of the other interconnected register.

21. An integrated circuit comprising:

memory to store a plurality of multiplier weight data; and

a plurality of multiplier-accumulator circuits, connected to form a processing pipeline, to perform a plurality of multiply and accumulate operations, wherein each multiplier-accumulator circuit includes:

a multiplier, coupled to the memory to (i) receive the multiplier weight data, (ii) multiply first data and the multiplier weight data and (iii) output product data, and

an accumulator, coupled to the multiplier of the multiplier-accumulator circuit, to add second data and the first product data and output sum data, and

a load-store register (i) directly connected to an output of the accumulator of the associated multiplier-accumulator circuit and (ii) coupled to an input of the load-store register of an immediately successive multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of the processing pipeline via an electrical path that bypasses an adder of the immediately successive multiplier-accumulator circuit.

22. The integrated circuit of claim 21 wherein:

each load-store register includes two interconnected registers.

23. The integrated circuit of claim 21 wherein:

the load-store register includes two interconnected registers; and

the load-store register further includes two multiplexers, wherein each multiplexer includes: (i) an output connected to an input of one of the interconnected registers and (ii) a first input connected to the output of the other interconnected register.

24. The integrated circuit of claim 21 wherein:

the load-store register of each multiplier-accumulator circuit is configurable to: (i) temporarily store initialization data and (ii) load the initialization data into the accumulator of the immediately successive multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of the pipeline.

25. The integrated circuit of claim 21 wherein:

the load-store register of each multiplier-accumulator circuit is configurable to: (i) temporarily store initialization data, (ii) load the initialization data into the accumulator of the immediately successive multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of the pipeline, and (iii) store the sum data from the associated multiplier-accumulator circuit into the load-store register.

26. The integrated circuit of claim 25 wherein:

each load-store register:

includes two interconnected registers, and

is configurable to, on the same clock cycle, (a) load the initialization data into the accumulator of the immediately successive multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of the pipeline, and (b) store the sum data from the associated multiplier-accumulator circuit into the load-store register, wherein the data in the two interconnected registers are swapped therebetween.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2024
From: FLEX LOGIX TECHNOLOGIES, INC.
To: ANALOG DEVICES, INC.
Reel/Frame 069409/0035 →
Continuity (5)
Division 17219952 · Apr 1, 2021
Division 16887265 · May 29, 2020
Division 16545345 · Aug 20, 2019
Provisional Application 62725306 · Aug 31, 2018
Related Publication 20220385291A1 · Dec 1, 2022
References Cited (65)
US 4958312A · Ang et al. · 1990 [cited by applicant]
US 5278781A · Aono · 1994 [cited by examiner]
US 6115729A · Matheny et al. · 2000 [cited by applicant]
US 6148101A · Tanaka et al. · 2000 [cited by applicant]
US 6298366B1 · Gatherer et al. · 2001 [cited by applicant]
US 6557022B1 · Sih et al. · 2003 [cited by applicant]
US 7107305B2 · Deng et al. · 2006 [cited by applicant]
US 7299342B2 · Nilsson et al. · 2007 [cited by applicant]
US 7346644B1 · Langhammer et al. · 2008 [cited by applicant]
US 7698358B1 · Langhammer et al. · 2010 [cited by applicant]
US 8051124B2 · Salama et al. · 2011 [cited by applicant]
US 8645450B1 · Choe et al. · 2014 [cited by applicant]
US 8751551B2 · Streicher et al. · 2014 [cited by applicant]
US 8788562B2 · Langhammer et al. · 2014 [cited by applicant]
US 9600278B1 · Langhammer · 2017 [cited by applicant]
US 10693469B2 · Wang · 2020 [cited by applicant]
US 10972103B2 · Wang · 2021 [cited by applicant]
US 11476854B2 · Wang · 2022 [cited by examiner]
US 20030172101A1 · Liao et al. · 2003 [cited by applicant]
US 20070239967A1 · Dally et al. · 2007 [cited by applicant]
US 20090094303A1 · Katayama · 2009 [cited by applicant]
US 20140019727A1 · Zhu et al. · 2014 [cited by applicant]
US 20170115958A1 · Langhammer · 2017 [cited by applicant]
US 20170116693A1 · Rae et al. · 2017 [cited by applicant]
US 20170214929A1 · Susnow et al. · 2017 [cited by applicant]
US 20170322813A1 · Langhammer · 2017 [cited by applicant]
US 20170344876A1 · Brothers · 2017 [cited by applicant]
US 20180052661A1 · Langhammer · 2018 [cited by applicant]
US 20180081632A1 · Langhammer · 2018 [cited by applicant]
US 20180081633A1 · Langhammer · 2018 [cited by applicant]
US 20180157961A1 · Henry et al. · 2018 [cited by applicant]
US 20180189651A1 · Henry et al. · 2018 [cited by applicant]
US 20180300105A1 · Langhammer · 2018 [cited by applicant]
US 20180321909A1 · Langhammer · 2018 [cited by applicant]
US 20180321910A1 · Langhammer et al. · 2018 [cited by applicant]
US 20180341460A1 · Langhammer · 2018 [cited by applicant]
US 20180341461A1 · Langhammer · 2018 [cited by applicant]
US 20190042191A1 · Langhammer · 2019 [cited by applicant]
US 20190079728A1 · Langhammer et al. · 2019 [cited by applicant]
US 20190196786A1 · Langhammer · 2019 [cited by applicant]
US 20190250886A1 · Langhammer · 2019 [cited by applicant]
US 20190286417A1 · Langhammer · 2019 [cited by applicant]
US 20190310828A1 · Langhammer et al. · 2019 [cited by applicant]
US 20190324722A1 · Langhammer · 2019 [cited by applicant]
US 20200004506A1 · Langhammer et al. · 2020 [cited by applicant]
US 20200026493A1 · Streicher et al. · 2020 [cited by applicant]
US 20200174750A1 · Langhammer · 2020 [cited by applicant]
US 20200326948A1 · Langhammer · 2020 [cited by applicant]
CN 1439126A · 2003 [cited by applicant]
CN 106610813A · 2019 [cited by applicant]
EP 0405726 · 1999 [cited by applicant]
EP 1229438A2 · 2002 [cited by applicant]
EP 1229438A3 · 2005 [cited by applicant]
EP 2280341 · 2013 [cited by applicant]
European Patent Office, Examination Report in European Patent Application No. 19855457.8, dated Sep. 9, 2023, 8 pages. [cited by applicant]
CNIPA, Office Action in China Patent Application No. 201980055487.4 with English-language translation, dated Aug. 9, 2023, 13 pages. [cited by applicant]
Priyanka Nain, “Multiplier-Accumulator (MAC) Unit”, IJDACR, vol. 5, Issue 3, Oct. 2016, 4 pages. [cited by applicant]
Jebashini et al., “A Survey and Comparative Analysis of Multiply-Accumulate (MAC) Block for Digital Signal Processing Application on ASIC and FPGA”, Journal of Applied Science, vol. 15, Issue 7, pp. 934-946, Jul. 2015. [cited by applicant]
Agrawal et al., “A 7nm 4-Core AI Chip with 25.6TFLOPS Hybrid FP8 Training, 102.4TOPS INT4 Inference and Workload-Aware Throttling”, ISSCC, pp. 144-145, 2021. [cited by applicant]
Linley Gwennap, “IBM Demonstrates New AI Data Types”, Microprocessor Report, Apr. 2021. [cited by applicant]
Choi et al., “Accurate and Efficient 2-Bit Quantized Neural Networks”, Proceedings of 2 [cited by applicant]
Sun et al., “Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks”, NeurIPS 2019, 10 pages. [cited by applicant]
“NVidia A100 Tensor Core GPU Architecture”, v1.0, 2020, 82 pages. [cited by applicant]
Papadantonakis et al., “Pipelining Saturated Accumulation”, IEEE, vol. 58, No. 2, pp. 208-219, Feb. 2009. [cited by applicant]
EPO Supplementary Search Report, mailed Nov. 9, 2021, 9 pages. [cited by applicant]