IP Library Granted Patent US 12,461,713
Granted Patent B2
US 12,461,713 · App. 17/683,284 · Granted Nov 4, 2025

MAC processing pipelines, circuitry to configure same, and methods of operating same

Inventors: Frederick A. Ware (Los Altos Hills, CA); Cheng C. Wang (San Jose, CA)
Assignee: Analog Devices, Inc.
G06F7/5443G06F9/3893G06F2207/3884
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,461,713
App. No.
17/683,284
Granted
Nov 4, 2025
Kind
B2
Abstract

An integrated circuit comprising a plurality MAC processors, interconnected into a linear pipeline, configurable to process input data, wherein each MAC processor includes (A) a multiplier and (B) an accumulator circuit, and (C) a plurality of rotate input data paths, wherein each rotate input data path couples two sequential MAC processors of the linear pipeline including an input of the multiplier circuit of a first MAC processor of sequential MAC processors to an input of the multiplier circuit of the immediately following MAC processor of the associated sequential MAC processors of the pipeline—wherein each rotate input data path is configurable to provide rotate input data from a first MAC processor of sequential MAC processors of the linear pipeline to the immediately following MAC processor of the associated sequential MAC processors thereby forming a serial circular path via the plurality of rotate input data paths.

Claims (54)

1 . An integrated circuit comprising:

a plurality MAC processors, interconnected into a linear pipeline, configurable to process input data during a plurality of execution cycles of an execution sequence, wherein each MAC processor of the plurality of the MAC processors includes:

a multiplier circuit to (i) receive first input data at a first input and (ii) receive multiplier weight data at a second input, (iii) multiply first input data by multiplier weight data to generate product data, and (iv) output the product data,

an accumulator circuit, coupled to the multiplier circuit of the associated MAC processor, to (i) receive accumulation data at a first input, (ii) receive the product data output by the associated multiplier circuit, (iii) add the accumulation data and the product data output by the associated multiplier circuit to generate second accumulation data, and (iv) output the second accumulation data via an output of the accumulation circuit, and

an accumulation data path to couple the output of the accumulation circuit of the associated MAC processor to the first input thereof;

a plurality of rotate input data paths, wherein each rotate input data path couples two sequential MAC processors of the linear pipeline including the first input of the multiplier circuit of a first MAC processor of sequential MAC processors to the first input of the multiplier circuit of the immediately following MAC processor of the associated sequential MAC processors of the linear pipeline; and

wherein each rotate input data path is configurable to provide rotate input data from a first MAC processor of sequential MAC processors of the linear pipeline to the immediately following MAC processor of the associated sequential MAC processors thereby forming a serial circular path, including the plurality of MAC processors of the linear pipeline, via the plurality of rotate input data paths.

2 . The integrated circuit of claim 1 wherein:

each MAC processor of the plurality of the MAC processors further includes an accumulation register disposed in the accumulation data path of the associated MAC processor between the output of the accumulation circuit of the associated MAC processor and the first input thereof.

3 . The integrated circuit of claim 2 wherein:

accumulation register, in operation, stores the second accumulation data output from the accumulation circuit of the associated MAC processor on a first execution cycle and the accumulator circuit inputs second accumulation data at the first input as accumulation data on an execution cycle that follows the first execution cycle.

4 . The integrated circuit of claim 2 wherein:

after a multiply-accumulate operation of the linear pipeline, the accumulation register of each MAC processor stores a different second accumulation data output from the accumulation circuit of the associated MAC processor.

5 . The integrated circuit of claim 1 wherein:

each MAC processor of the plurality of the MAC processors further includes an input data register connected to an output of the rotate input data path to store rotate input data provided from the immediately preceding MAC processor of the plurality of the MAC processors of the linear pipeline, via the associated rotate input data path, thereby forming a circular serial shift path between the plurality of MAC processors of the linear pipeline.

6 . The integrated circuit of claim 1 wherein:

each MAC processor of the plurality of the MAC processors further includes an input data register disposed between the first input of the multiplier circuit of the first MAC processor and the first input of the multiplier circuit of the second MAC processor of each associated pair of MAC processors to store rotate input data provided from the immediately preceding MAC processor, via the associated rotate input data path, after a multiply-accumulate operation of the plurality of the MAC processors of the linear pipeline.

7 . The integrated circuit of claim 1 wherein:

each MAC processor of the plurality of the MAC processors further includes a data input register, connected to an output of the associated rotate input data path, to store rotate input data.

8 . The integrated circuit of claim 1 wherein:

each MAC processor of the plurality of the MAC processors further includes a memory dedicated to the associated MAC processor, wherein the memory is configurable to store filter weight data corresponding to the associated MAC processor.

9 . The integrated circuit of claim 8 wherein:

the memory includes a plurality of memory banks, each bank is configurable to store filter weight data for each execution cycle of the execution sequence.

10 . An integrated circuit comprising:

a plurality MAC processors, interconnected into a linear pipeline, configurable to process input data during a plurality of execution cycles of an execution sequence, wherein each MAC processor of the plurality of the MAC processors includes:

a multiplier circuit to (i) receive first input data at a first input and (ii) receive multiplier weight data at a second input, (iii) multiply first input data by multiplier weight data to generate product data, and (iv) output the product data,

an accumulator circuit, coupled to the multiplier circuit of the associated MAC processor, to (i) receive accumulation data at a first input, (ii) receive the product data output by the associated multiplier circuit, (iii) add the accumulation data and the product data output by the associated multiplier circuit to generate second accumulation data, and (iv) output the second accumulation data via an output of the accumulation circuit, and

an accumulation data path to couple the output of the accumulation circuit of the associated MAC processor to the first input thereof;

a plurality of rotate input data paths, wherein each rotate input data path couples two sequential MAC processors of the linear pipeline including the first input of the multiplier circuit of a first MAC processor of sequential MAC processors to the first input of the multiplier circuit of the immediately following MAC processor of the associated sequential MAC processors of the linear pipeline;

a plurality of rotate accumulation data paths, wherein each rotate accumulation data path couples two sequential MAC processors of the linear pipeline including the output of the accumulator circuit of a first MAC processor of sequential MAC processors to the first input of the accumulator circuit of the immediately following MAC processor of the associated sequential MAC processors of the linear pipeline;

wherein each rotate input data path is configurable to provide rotate input data from a first MAC processor of sequential MAC processors to the immediately following MAC processor of the associated sequential MAC processors thereby forming a first configurable serial circular path, including the plurality of MAC processors of the linear pipeline, via the plurality of rotate input data paths; and

wherein each rotate accumulation data path is configurable to provide rotate accumulation data from the first MAC processor of sequential MAC processors to the immediately following MAC processor of the associated sequential MAC processors thereby forming a second configurable serial circular path, including the plurality of MAC processors of the linear pipeline, via the plurality of rotate accumulation data paths.

11 . The integrated circuit of claim 10 wherein:

each MAC processor includes circuitry to responsively connect (1) the rotate input data path between each of the sequential MAC processors of the linear pipeline or (2) the rotate accumulation data path between each of the sequential MAC processors of the linear pipeline.

12 . The integrated circuit of claim 10 wherein:

each MAC processor of the plurality of the MAC processors further includes a multiplexer having (i) a first input connected to an output of the rotate input data path from the immediately preceding MAC processor, (ii) a second input connected to a shift in data path of the linear pipeline, and (iii) and output coupled to the first input of the multiplier circuit of the associated MAC processor.

13 . The integrated circuit of claim 10 wherein:

each MAC processor of the plurality of the MAC processors further includes a multiplexer having (i) a first input connected to an output of the rotate accumulation data path from the immediately preceding MAC processor, (ii) a second input connected to the output of the accumulator circuit of the associated MAC processor, and (iii) and output coupled to the first input of the accumulator circuit of the associated MAC processor.

14 . The integrated circuit of claim 10 wherein:

each MAC processor of the plurality of the MAC processors further includes:

a first multiplexer having (i) a first input connected to an output of the rotate input data path from the immediately preceding MAC processor, (ii) a second input connected to a shift in data path of the linear pipeline, and (iii) and output coupled to the first input of the multiplier circuit of the associated MAC processor; and

a second multiplexer having (i) a first input connected to an output of the rotate accumulation data path from the immediately preceding MAC processor, (ii) a second input connected to the output of the accumulator circuit of the associated MAC processor, and (iii) and output coupled to the first input of the accumulator circuit of the associated MAC processor.

15 . The integrated circuit of claim 10 wherein:

each MAC processor of the plurality of the MAC processors further includes an accumulation register disposed in the accumulation data path of the associated MAC processor between the output of the accumulation circuit of the associated MAC processor and the first input thereof.

16 . The integrated circuit of claim 15 wherein:

accumulation register, in operation, stores the second accumulation data output from the accumulation circuit of the associated MAC processor on a first execution cycle and the accumulator circuit inputs second accumulation data at the first input as accumulation data on an execution cycle that follows the first execution cycle.

17 . The integrated circuit of claim 15 wherein:

after a multiply-accumulate operation, the accumulation register of each MAC processor stores a different second accumulation data output from the accumulation circuit of the associated MAC processor.

18 . The integrated circuit of claim 10 wherein:

each MAC processor of the plurality of the MAC processors further includes an input data register connected to an output of the rotate input data path to store rotate input data provided from the immediately preceding MAC processor of the plurality of the MAC processors of the linear pipeline, via the associated rotate input data path, thereby forming a circular serial shift path between the plurality of MAC processors of the linear pipeline.

19 . The integrated circuit of claim 10 wherein:

each MAC processor of the plurality of the MAC processors further includes an input data register disposed between the first input of the multiplier circuit of the first MAC processor and the first input of the multiplier circuit of the second MAC processor of each associated pair of MAC processors to store rotate input data provided to the first MAC processor after a multiply-accumulate operation of the plurality of the MAC processors of the linear pipeline.

20 . The integrated circuit of claim 10 wherein:

each MAC processor of the plurality of the MAC processors further includes an input data register to store rotate input data after a multiply-accumulate operation of the plurality of the MAC processors of the linear pipeline.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2024
From: FLEX LOGIX TECHNOLOGIES, INC.
To: ANALOG DEVICES, INC.
Reel/Frame 069409/0035 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2022
From: WARE, FREDERICK A; WANG, CHENG C
To: FLEX LOGIX TECHNOLOGIES, INC.
Reel/Frame 059458/0396 →
Continuity (2)
Provisional Application 63156263 · Mar 3, 2021
Related Publication 20220283779A1 · Sep 8, 2022
References Cited (88)
US 4958312A · Ang et al. · 1990 [cited by applicant]
US 6115729A · Matheny et al. · 2000 [cited by applicant]
US 6148101A · Tanaka et al. · 2000 [cited by applicant]
US 6298366B1 · Gatherer et al. · 2001 [cited by applicant]
US 6538470B1 · Langhammer et al. · 2003 [cited by applicant]
US 7107305B2 · Deng et al. · 2006 [cited by applicant]
US 7299342B2 · Nilsson et al. · 2007 [cited by applicant]
US 7346644B1 · Langhammer et al. · 2008 [cited by applicant]
US 7698358B1 · Langhammer et al. · 2010 [cited by applicant]
US 8051124B2 · Salama et al. · 2011 [cited by applicant]
US 8266199B2 · Langhammer et al. · 2012 [cited by applicant]
US 8645450B1 · Choe et al. · 2014 [cited by applicant]
US 8751551B2 · Streicher et al. · 2014 [cited by applicant]
US 8788562B2 · Langhammer et al. · 2014 [cited by applicant]
US 9600278B1 · Langhammer · 2017 [cited by applicant]
US 11314504B2 · Ware et al. · 2022 [cited by applicant]
US 20030172101A1 · Liao et al. · 2003 [cited by applicant]
US 20050144215A1 · Simkins · 2005 [cited by examiner]
US 20070239967A1 · Dally et al. · 2007 [cited by applicant]
US 20080211827A1 · Donovan et al. · 2008 [cited by applicant]
US 20090094303A1 · Katayama · 2009 [cited by applicant]
US 20140019727A1 · Zhu et al. · 2014 [cited by applicant]
US 20140281370A1 · Kahn · 2014 [cited by applicant]
US 20170011288A1 · Brothers et al. · 2017 [cited by applicant]
US 20170115958A1 · Langhammer · 2017 [cited by applicant]
US 20170116693A1 · Rae et al. · 2017 [cited by applicant]
US 20170214929A1 · Susnow et al. · 2017 [cited by applicant]
US 20170315778A1 · Sano · 2017 [cited by applicant]
US 20170322813A1 · Langhammer · 2017 [cited by examiner]
US 20170344876A1 · Brothers · 2017 [cited by applicant]
US 20180052661A1 · Langhammer · 2018 [cited by applicant]
US 20180081632A1 · Langhammer · 2018 [cited by applicant]
US 20180081633A1 · Langhammer · 2018 [cited by applicant]
US 20180157961A1 · Henry et al. · 2018 [cited by applicant]
US 20180189651A1 · Henry et al. · 2018 [cited by applicant]
US 20180300105A1 · Langhammer · 2018 [cited by applicant]
US 20180314492A1 · Fais et al. · 2018 [cited by applicant]
US 20180321909A1 · Langhammer · 2018 [cited by applicant]
US 20180321910A1 · Langhammer et al. · 2018 [cited by applicant]
US 20180341460A1 · Langhammer · 2018 [cited by applicant]
US 20180341461A1 · Langhammer · 2018 [cited by applicant]
US 20190042191A1 · Langhammer · 2019 [cited by applicant]
US 20190042244A1 · Henry et al. · 2019 [cited by applicant]
US 20190042544A1 · Kashyap et al. · 2019 [cited by applicant]
US 20190079728A1 · Langhammer et al. · 2019 [cited by applicant]
US 20190196786A1 · Langhammer · 2019 [cited by applicant]
US 20190243610A1 · Lin et al. · 2019 [cited by applicant]
US 20190250886A1 · Langhammer · 2019 [cited by applicant]
US 20190286417A1 · Langhammer · 2019 [cited by applicant]
US 20190310828A1 · Langhammer et al. · 2019 [cited by applicant]
US 20190324722A1 · Langhammer · 2019 [cited by applicant]
US 20190340489A1 · Mills · 2019 [cited by applicant]
US 20190392297A1 · Lau et al. · 2019 [cited by applicant]
US 20200004506A1 · Langhammer et al. · 2020 [cited by applicant]
US 20200026493A1 · Streicher et al. · 2020 [cited by applicant]
US 20200076435A1 · Wang · 2020 [cited by applicant]
US 20200097799A1 · Divakar et al. · 2020 [cited by applicant]
US 20200174750A1 · Langhammer · 2020 [cited by applicant]
US 20200202198A1 · Lee et al. · 2020 [cited by applicant]
US 20200310818A1 · Ware et al. · 2020 [cited by applicant]
US 20200326939A1 · Ware et al. · 2020 [cited by applicant]
US 20200326948A1 · Langhammer · 2020 [cited by applicant]
US 20200401414A1 · Ware et al. · 2020 [cited by applicant]
US 20210064568A1 · Wang et al. · 2021 [cited by applicant]
US 20210081211A1 · Wang · 2021 [cited by applicant]
US 20210103630A1 · Ware et al. · 2021 [cited by applicant]
US 20210132905A1 · Ware et al. · 2021 [cited by applicant]
US 20210173617A1 · Ware et al. · 2021 [cited by applicant]
US 20210263993A1 · Urbanski · 2021 [cited by examiner]
US 20210326286A1 · Ware et al. · 2021 [cited by applicant]
US 20220027152A1 · Ware et al. · 2022 [cited by applicant]
US 20220057994A1 · Ware et al. · 2022 [cited by applicant]
US 20220171604A1 · Ware et al. · 2022 [cited by applicant]
US 20220244917A1 · Ware · 2022 [cited by examiner]
EP 0405726 · 1999 [cited by applicant]
EP 2280341 · 2013 [cited by applicant]
WO WO2018126073 · 2018 [cited by applicant]
Priyanka Nain, “Multiplier-Accumulator (MAC) Unit”, IJDACR, vol. 5, Issue 3, Oct. 2016, 4 pages. [cited by applicant]
Jebashini et al., “A Survey and Comparative Analysis of Multiply-Accumulate (MAC) Block for Digital Signal Processing Application on ASIC and FPGA”, Journal of Applied Science, vol. 15, Issue 7, pp. 934-946, Jul. 2015. [cited by applicant]
Agrawal et al., “A 7nm 4-Core Al Chip with 25.6TFLOPS Hybrid FP8 Training, 102.4TOPS INT4 Inference and Workload-Aware Throttling”, ISSCC, pp. 144-145, 2021. [cited by applicant]
Choi et al., “Accurate and Efficient 2-Bit Quantized Neural Networks”, Proceedings of 2 [cited by applicant]
Sun et al., “Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks”, NeurIPS 2019, 10 pages. [cited by applicant]
Nvidia A100 Tensor Core GPU Architecture, v1.0, 2020, 82 pages. [cited by applicant]
Papadantonakis et al., “Pipelining Saturated Accumulation”, IEEE, vol. 58, No. 2, pp. 208-219, Feb. 2009. [cited by applicant]
Linley Gwennap, “IBM Demonstrates New AI Data Types”, Microprocessor Report, Apr. 2021. [cited by applicant]
Liang Yun, et al., “Evaluating Fast Algorithms for Convolutional Neural Networks on FPGAs”, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 39, No. 4, Feb. 5, 2019, 14 pages (Note: Th… [cited by applicant]
Zhao et al., “A Fast Algorithm for Reducing the Computation Complexity of Convolutional Neural Networks”, Algorithms 2018, 11, 159; doi:10.3390/a11100159; www. Mdpi.com/journal/algorithms, 11 pages, Oct. 2018. [cited by applicant]
International Search Report and Written Opinion of International Searching Authority re: re: PCT/US2022/018239, mailed Jun. 24, 2022, 18 pages. [cited by applicant]