IP Library Granted Patent US 12,282,749
Granted Patent B2
US 12,282,749 · App. 17/391,082 · Granted Apr 22, 2025

Configurable MAC pipelines for finite-impulse-response filtering, and methods of operating same

Inventors: Frederick A. Ware (Los Altos Hills, CA); Cheng C. Wang (San Jose, CA)
Assignee: Analog Devices, Inc.
G06F7/5443G06F2207/3884G06F2207/4812H03H2017/0081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,282,749
App. No.
17/391,082
Granted
Apr 22, 2025
Kind
B2
Abstract

An integrated circuit comprising a plurality MAC pipelines wherein each MAC pipeline includes: (i) a plurality of MACs connected in series and (ii) a plurality of data paths including an accumulation data path, wherein each MAC includes a multiplier to multiply to generate product data and an accumulator to generate sum data. The integrated circuit further comprises a plurality of control/configure circuits, wherein each control/configure circuit connects directly to and is associated with a MAC pipeline, wherein each control/configure circuit includes an accumulation data path which is configurable to directly connect to the accumulation data path of the MAC pipeline to form an accumulation ring when the control/configure circuit is configured in an accumulation mode, and an output data path configurable to directly connect to the output of the accumulation data path of the MAC pipeline when the control/configure circuit is configured in an output data mode.

Claims (51)

1. An integrated circuit comprising:

a plurality MAC pipelines wherein each MAC pipeline includes: (i) a plurality of multiplier-accumulator circuits connected in series to perform a plurality of concatenated multiply and accumulate operations and (ii) a plurality of data paths including an accumulation data path, and wherein each multiplier-accumulator circuit of the plurality of multiplier-accumulator circuits of each MAC pipeline includes:

a multiplier to multiply first data by multiplier weight data and generate product data, and

an accumulator, coupled to the multiplier of the associated multiplier-accumulator circuit, to add second data and the product data of the associated multiplier to generate sum data;

a plurality of control/configure circuits, wherein each control/configure circuit connects directly to and is associated with one of the MAC pipelines, wherein each control/configure circuit includes a plurality of configurable data paths including:

an accumulation data path which is configurable to directly connect to an input and an output of the accumulation data path of the associated MAC pipeline to form an accumulation ring when the associated control/configure circuit is configured in an accumulation mode, and

an output data path which is configurable to directly connect to the output of the accumulation data path of the associated MAC pipeline to receive final accumulation data from the associated MAC pipeline when the associated control/configure circuit is configured in an output data mode.

2. The integrated circuit of claim 1 wherein:

each control/configure circuit includes: (a) a first output data port disposed at a first interface of the associated control/configure circuit, (b) a second output data port disposed at a second interface of the associated control/configure circuit wherein the second interface is an interface with the associated MAC pipeline, and (c) a first multiplexer having:

a first input directly connected to the second output data port,

a second input connected to the output of the accumulation data path of the associated MAC pipeline, and

and an output connected to the output data path of the associated control/configure circuit; and

the first multiplexer responsively connects the second input of the first multiplexer, and the output of the accumulation data path of the associated MAC pipeline, to the output of the first multiplexer when the associated control/configure circuit is configured in the output data mode.

3. The integrated circuit of claim 2 wherein:

the plurality of configurable data paths of each control/configure circuit further includes an input path configurable to connect to the accumulation data path of the associated MAC pipeline to input initial values into the plurality of multiplier-accumulator circuits of the associated MAC pipeline when the associated control/configure circuit is configured in a data load mode.

4. The integrated circuit of claim 3 wherein:

each control/configure circuit includes: (a) an input port disposed at the first interface of the associated control/configure circuit, (b) a first accumulation data port connected to an input of the accumulation data path of the associated MAC pipeline and disposed at the second interface of the associated control/configure circuit, and (c) a second multiplexer having:

a first input connected to the accumulation data path of the control/configure circuit,

a second input connected to an input path of the control/configure circuit, and

an output connected to the input of the accumulation data path of the associated MAC pipeline; and

the second multiplexer responsively connects the second input of the second multiplexer, and the input path of the associated MAC pipeline, to the output of the second multiplexer when the associated control/configure circuit is configured in the data load mode.

5. The integrated circuit of claim 4 wherein:

the second multiplexer responsively connects the first input, and the accumulation data path of the associated MAC pipeline, to the output of the second multiplexer when the associated control/configure circuit is configured in the accumulation mode.

6. The integrated circuit of claim 5 wherein:

the first multiplexer responsively connects the first input, and the accumulation data path of the associated MAC pipeline, to the output of the second multiplexer when the associated control/configure circuit is configured in an accumulation mode.

7. The integrated circuit of claim 1 further including:

configurable processing circuitry, coupled to the output data path of at least one control/configure circuit, wherein the configurable processing circuitry is programmable to receive and post-process the final accumulation data from the MAC pipeline which is associated with the at least one control/configure circuit.

8. An integrated circuit comprising:

a plurality MAC pipelines wherein each MAC pipeline includes: (i) a plurality of multiplier-accumulator circuits connected in series to perform a plurality of concatenated multiply and accumulate operations and (ii) a plurality of data paths including an accumulation data path, and wherein each multiplier-accumulator circuit of the plurality of multiplier- accumulator circuits of each MAC pipeline includes:

a multiplier to multiply first data by multiplier weight data and generate product data, and

an accumulator, coupled to the multiplier of the associated multiplier-accumulator circuit, to add second data and the product data of the associated multiplier to generate sum data;

a plurality of control/configure circuits, wherein each control/configure circuit connects directly to and is associated with one of the MAC pipelines, wherein each control/configure circuit includes a plurality of configurable data paths including:

an input path configurable to connect to the accumulation data path of the associated MAC pipeline when the associated control/configure circuit is configured in a data load mode, and

an accumulation data path configurable to connect to an input and an output of the accumulation data path of the associated MAC pipeline to form an accumulation ring when the associated control/configure circuit is configured in an accumulation mode.

9. The integrated circuit of claim 8 wherein:

each control/configure circuit further includes a first multiplexer having:

a first input connected to the accumulation data path of the control/configure circuit,

a second input connected to the input path of the control/configure circuit, and

an output connected to the input of the accumulation data path of the associated MAC pipeline; and

the first multiplexer responsively connects the second input, and the input path of the associated control/configure circuit, to the output of the first multiplexer when the associated control/configure circuit is configured in the data load mode.

10. The integrated circuit of claim 9 wherein:

when the control/configure circuit is configured in the data load mode, the associated MAC pipeline receives initial values via the input path of the associated control/configure circuit.

11. The integrated circuit of claim 10 wherein:

the initial values are partial accumulation values.

12. The integrated circuit of claim 9 wherein:

the first multiplexer responsively connects the first input of the first multiplexer, and the accumulation data path of the associated control/configure circuit, to the output of the first multiplexer when the associated control/configure circuit is configured in an accumulation mode.

13. The integrated circuit of claim 8 wherein:

the plurality of configurable data paths of each control/configure circuit further includes:

an output data path is configurable to directly connect to the output of the accumulation data path of the associated MAC pipeline to receive final accumulation data from the associated MAC pipeline when the associated control/configure circuit is configured in an output data mode.

14. The integrated circuit of claim 13 further including:

configurable processing circuitry, coupled to the output data path of at least one control/configure circuit, wherein the configurable processing circuitry is programmable to receive and post-process the final accumulation data from the MAC pipeline which is associated with the at least one control/configure circuit.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2024
From: FLEX LOGIX TECHNOLOGIES, INC.
To: ANALOG DEVICES, INC.
Reel/Frame 069409/0035 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 20, 2021
From: WARE, FREDERICK A; WANG, CHENG C
To: FLEX LOGIX TECHNOLOGIES, INC.
Reel/Frame 057238/0601 →
Continuity (2)
Provisional Application 63067979 · Aug 20, 2020
Related Publication 20220057994A1 · Feb 24, 2022
References Cited (85)
US 4958312A · Ang et al. · 1990 [cited by applicant]
US 6115729A · Matheny et al. · 2000 [cited by applicant]
US 6148101A · Tanaka et al. · 2000 [cited by applicant]
US 6298366B1 · Gatherer et al. · 2001 [cited by applicant]
US 6538470B1 · Langhammer et al. · 2003 [cited by applicant]
US 7107305B2 · Deng et al. · 2006 [cited by applicant]
US 7299342B2 · Nilsson et al. · 2007 [cited by applicant]
US 7346644B1 · Langhammer et al. · 2008 [cited by applicant]
US 7698358B1 · Langhammer et al. · 2010 [cited by applicant]
US 8051124B2 · Salama et al. · 2011 [cited by applicant]
US 8266199B2 · Langhammer et al. · 2012 [cited by applicant]
US 8645450B1 · Choe et al. · 2014 [cited by applicant]
US 8751551B2 · Streicher et al. · 2014 [cited by applicant]
US 8788562B2 · Langhammer et al. · 2014 [cited by applicant]
US 9600278B1 · Langhammer · 2017 [cited by applicant]
US 11314504B2 · Ware et al. · 2022 [cited by applicant]
US 20030172101A1 · Liao et al. · 2003 [cited by applicant]
US 20050144215A1 · Simkins et al. · 2005 [cited by applicant]
US 20070239967A1 · Dally et al. · 2007 [cited by applicant]
US 20080211827A1 · Donovan et al. · 2008 [cited by applicant]
US 20090094303A1 · Katayama · 2009 [cited by applicant]
US 20140019727A1 · Zhu et al. · 2014 [cited by applicant]
US 20140281370A1 · Kahn · 2014 [cited by applicant]
US 20170011288A1 · Brothers et al. · 2017 [cited by applicant]
US 20170115958A1 · Langhammer · 2017 [cited by applicant]
US 20170116693A1 · Rae et al. · 2017 [cited by applicant]
US 20170214929A1 · Susnow et al. · 2017 [cited by applicant]
US 20170315778A1 · Sano · 2017 [cited by applicant]
US 20170322813A1 · Langhammer · 2017 [cited by applicant]
US 20170344876A1 · Brothers · 2017 [cited by applicant]
US 20180052661A1 · Langhammer · 2018 [cited by applicant]
US 20180081632A1 · Langhammer · 2018 [cited by applicant]
US 20180081633A1 · Langhammer · 2018 [cited by applicant]
US 20180157961A1 · Henry et al. · 2018 [cited by applicant]
US 20180189651A1 · Henry et al. · 2018 [cited by applicant]
US 20180300105A1 · Langhammer · 2018 [cited by applicant]
US 20180314492A1 · Fais et al. · 2018 [cited by applicant]
US 20180321909A1 · Langhammer · 2018 [cited by applicant]
US 20180321910A1 · Langhammer et al. · 2018 [cited by applicant]
US 20180341460A1 · Langhammer · 2018 [cited by applicant]
US 20180341461A1 · Langhammer · 2018 [cited by applicant]
US 20190042191A1 · Langhammer · 2019 [cited by applicant]
US 20190042244A1 · Henry et al. · 2019 [cited by applicant]
US 20190042544A1 · Kashyap et al. · 2019 [cited by applicant]
US 20190079728A1 · Langhammer et al. · 2019 [cited by applicant]
US 20190196786A1 · Langhammer · 2019 [cited by applicant]
US 20190243610A1 · Lin et al. · 2019 [cited by applicant]
US 20190250886A1 · Langhammer · 2019 [cited by applicant]
US 20190286417A1 · Langhammer · 2019 [cited by applicant]
US 20190310828A1 · Langhammer et al. · 2019 [cited by applicant]
US 20190324722A1 · Langhammer · 2019 [cited by applicant]
US 20190340489A1 · Mills · 2019 [cited by applicant]
US 20190392297A1 · Lau et al. · 2019 [cited by applicant]
US 20200004506A1 · Langhammer et al. · 2020 [cited by applicant]
US 20200026493A1 · Streicher et al. · 2020 [cited by applicant]
US 20200076435A1 · Wang · 2020 [cited by applicant]
US 20200097799A1 · Divakar et al. · 2020 [cited by applicant]
US 20200174750A1 · Langhammer · 2020 [cited by applicant]
US 20200202198A1 · Lee et al. · 2020 [cited by applicant]
US 20200310818A1 · Ware et al. · 2020 [cited by applicant]
US 20200326939A1 · Ware et al. · 2020 [cited by applicant]
US 20200326948A1 · Langhammer · 2020 [cited by applicant]
US 20200401414A1 · Ware et al. · 2020 [cited by applicant]
US 20210073171A1 · Master · 2021 [cited by examiner]
US 20210081211A1 · Wang · 2021 [cited by applicant]
US 20210103630A1 · Ware et al. · 2021 [cited by applicant]
US 20210132905A1 · Ware et al. · 2021 [cited by applicant]
US 20210173617A1 · Ware et al. · 2021 [cited by applicant]
US 20210326286A1 · Ware et al. · 2021 [cited by applicant]
US 20220027152A1 · Ware et al. · 2022 [cited by applicant]
EP 0405726 · 1999 [cited by applicant]
EP 2280341 · 2013 [cited by applicant]
WO WO2018126073 · 2018 [cited by applicant]
International Bureau of World Intellectual Property Organization (WIPO), International Preliminary Report on Patentability dated Mar. 2, 2023 in International Application No. PCT/US2021/044118, 8 pages. [cited by applicant]
International Search Report and Written Opinion of International Searching Authority re: re: PCT/US2020/024808, mailed Dec. 30, 2021, 11 pages. [cited by applicant]
Zhao et al., “A Fast Algorithm for Reducing the Computation Complexity of Convolutional Neural Networks”, Algorithms 2018, 11, 159; doi:10.3390/a11100159; www. Mdpi.com/journal/algorithms, 11 pages, Oct. 2018. [cited by applicant]
Liang Yun, et al., “Evaluating Fast Algorithms for Convolutional Neural Networks on FPGAs”, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 39, No. 4, Feb. 5, 2019, 14 pages (Note: Th… [cited by applicant]
Priyanka Nain, “Multiplier-Accumulator (MAC) Unit”, IJDACR, vol. 5, Issue 3, Oct. 2016, 4 pages. [cited by applicant]
Jebashini et al., “A Survey and Comparative Analysis of Multiply-Accumulate (MAC) Block for Digital Signal Processing Application on ASIC and FPGA”, Journal of Applied Science, vol. 15, Issue 7, pp. 934-946, Jul. 2015. [cited by applicant]
Agrawal et al., “A 7nm 4-Core AI Chip with 25.6TFLOPS Hybrid FP8 Training, 102.4TOPS INT4 Inference and Workload-Aware Throttling”, ISSCC, pp. 144-145, 2021. [cited by applicant]
Linley Gwennap, “IBM Demonstrates New AI Data Types”, Microprocessor Report, Apr. 2021. [cited by applicant]
Choi et al., “Accurate and Efficient 2-Bit Quantized Neural Networks”, Proceedings of 2 [cited by applicant]
Sun et al., “Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks”, NeurIPS 2019, 10 pages. [cited by applicant]
“NVidia A100 Tensor Core GPU Architecture”, v1.0, 2020, 82 pages. [cited by applicant]
Papadantonakis et al., “Pipelining Saturated Accumulation”, IEEE, vol. 58, No. 2, pp. 208-219, Feb. 2009. [cited by applicant]