IP Library Patent Application 18144288
Patent Application
App. No. 18/144,288

BROADCAST DATA MULTIPLY-ACCUMULATE WITH SHARED UNLOAD

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/144,288
Abstract

An integrated circuit device includes broadcast data paths, a weighting-value memory, multiply-accumulate (MAC) units, and shared shift-out circuitry. The MAC units are coupled in common to each of the broadcast data paths and coupled to receive respective weighting values from the weighting-value memory via respective weighting-value paths. Each of the MAC units includes MAC circuits that each receive an input data value via a respective one of the broadcast data paths and a shared one of the weighting values via a shared one of the respective weighting-value paths; generate a sequence of multiplication products by multiplying the input data value with the shared one of the weighting values; accumulate a sum of the multiplication products; and output the sum of the multiplication products to a respective one of a plurality of serially coupled storage elements within the shared shift-out path.

Claims (45)

1 . An integrated circuit device comprising:

a plurality of broadcast data paths;

a weighting-value memory;

a plurality of multiply-accumulate (MAC) units coupled in common to each of the broadcast data paths and coupled to receive respective weighting values from the weighting-value memory via respective weighting-value paths, each of the MAC units having a plurality of MAC circuits coupled respectively to the broadcast data paths, each of the MAC circuits within a given one of the MAC units having:

a data input coupled to receive, during each of a plurality of timing cycles, an input data value via a respective one of the broadcast data paths;

a weighting-value input coupled to receive, during each of the plurality of timing cycles, a shared one of the weighting values via a shared one of the respective weighting-value paths;

a multiplier circuit to generate a sequence of multiplication products by multiplying the input data value received during each of the plurality of timing cycles with the shared one of the weighting values received during each of the plurality of timing cycles; and

an accumulator circuit to accumulate a sum of constituent multiplication products within the sequence of multiplication products; and

shift-out circuitry having storage elements coupled to receive respective instances of the sum of constituent multiplication products from the accumulator circuits within the plurality of MAC units and to sequentially output the respective instances of the sum of constituent multiplication products over a number of the timing cycles that exceeds a number of MAC units included within the plurality of MAC units.

2 . The integrated circuit device of claim 1 wherein the shift-out circuitry having the storage elements to sequentially output the respective instances of the sum of constituent multiplication products comprises a quantity N of the storage elements coupled to form a serial shift register, and wherein the plurality of MAC units is constituted by a quantity L of the MAC units, where L is less than N.

3 . The integrated circuit device of claim 2 wherein each of the MAC units comprises a number M of the MAC circuits such that the plurality of MAC units comprises L*M MAC circuits, wherein a ratio of N to L*M is less than or equal to one.

4 . The integrated circuit device of claim 2 wherein each of the MAC units comprises a number M of the MAC circuits, wherein N is equal to a product of L and M.

5 . The integrated circuit device of claim 1 wherein the number of timing cycles corresponds to a collective number of the MAC circuits included within the plurality of MAC units.

6 . The integrated circuit device of claim 1 wherein the storage elements within the shift-out circuitry are configured to sequentially output respective instances of the sum of constituent multiplication products from the accumulator circuits within one or more of the plurality of MAC units.

7 . The integrated circuit device of claim 1 wherein the storage elements within the shift-out circuitry are configured to sequentially output all instances of the sum of constituent multiplication products from the accumulator circuits within a first MAC unit of the plurality of MAC units before outputting any instances of the sum of constituent multiplication products from the accumulator circuits within a second MAC unit of the plurality of MAC units.

8 . The integrated circuit device of claim 1 wherein the storage elements within the shift-out circuitry are configured to sequentially output all instances of the sum of constituent multiplication products from the accumulator circuits within a first MAC unit of the plurality of MAC units before outputting all instances of the sum of constituent multiplication products from the accumulator circuits within a second MAC unit of the plurality of MAC units.

9 . The integrated circuit device of claim 1 wherein each of the MAC circuits further comprises a data operand register, coupled between the data input and the multiplier circuit, to store the input data value received during each of the plurality of timing cycles and to output the data input value received during each of the plurality of timing cycles to the multiplier circuit.

10 . The integrated circuit device of claim 1 wherein the multiplier circuit within at least one of the MAC circuits within the given one of the MAC units is coupled to output a carry value to the multiplier circuit of at least one other of the MAC circuits within the given one of the MAC units.

11 . A method of operation with an integrated-circuit (IC) device having a plurality of broadcast data paths, a weighting-value memory, shift-out circuitry, and a plurality of multiply-accumulate (MAC) units coupled in common to each of the broadcast data paths and coupled to receive respective weighting values from the weighting-value memory via respective weighting-value paths, each of the MAC units having a plurality of MAC circuits with inputs coupled respectively to the broadcast data paths and outputs coupled to respective storage elements within the shift-out circuitry, the method comprising:

executing the following operations in parallel within each of the MAC circuits of a given one of the MAC units:

receiving, during each of a plurality of timing cycles, an input data value via a respective one of the broadcast data paths;

receiving, during each of the plurality of timing cycles, a shared one of the weighting values via a shared one of the respective weighting-value paths;

multiplying the input data value received during each of the plurality of timing cycles with the shared one of the weighting values received during each of the plurality of timing cycles to generate a sequence of multiplication products; and

accumulating a sum of constituent multiplication products within the sequence of multiplication products;

loading respective instances of the sum of constituent multiplication products accumulated within the plurality of MAC units into the storage elements within the shift-out circuitry; and

sequentially outputting the respective instances of the sum of constituent multiplication products via the shift-out circuitry over a number of the timing cycles that exceeds a number of MAC units included within the plurality of MAC units.

12 . The method of claim 11 wherein sequentially outputting the respective instances of the sum of constituent multiplication products comprises sequentially outputting the respective instances of the sum of constituent multiplication products via a quantity N of the storage elements coupled to form a serial shift register, and wherein the plurality of MAC units is constituted by a quantity L of the MAC units, where L is less than N.

13 . The method of claim 12 wherein each of the MAC units comprises a number M of the MAC circuits such that the plurality of MAC units comprises L*M MAC circuits, wherein a ratio of N to L*M is less than or equal to one.

14 . The method of claim 12 wherein each of the MAC units comprises a number M of the MAC circuits, wherein N is equal to a product of L and M.

15 . The method of claim 11 wherein sequentially outputting the respective instances of the sum of constituent multiplication products via the shift-out circuitry comprises sequentially outputting the respective instances of the sum of constituent multiplication products from the shift-out circuitry over a number of timing cycles that corresponds to a collective number of the MAC circuits included within the plurality of MAC units.

16 . The method of claim 11 wherein sequentially outputting the respective instances of the sum of constituent multiplication products via the shift-out circuitry comprises sequentially outputting all instances of the sum of constituent multiplication products from the accumulator circuits within a first MAC unit of the plurality of MAC units before outputting any instances of the sum of constituent multiplication products from the accumulator circuits within a second MAC unit of the plurality of MAC units.

17 . The method of claim 11 wherein sequentially outputting the respective instances of the sum of constituent multiplication products via the shift-out circuitry comprises sequentially outputting all instances of the sum of constituent multiplication products from the accumulator circuits within a first MAC unit of the plurality of MAC units before outputting all instances of the sum of constituent multiplication products from the accumulator circuits within a second MAC unit of the plurality of MAC units.

18 . The method of claim 11 further comprising:

storing the input data value received via the respective one of the broadcast data paths the during each of the plurality of timing cycles within a respective data operand register within each of the MAC circuits of the given one of the MAC units; and

storing a respective one of the weighting values received via a respective one of the weighting-value paths within a weighting-value register included in the given one of the MAC units.

19 . The method of claim 11 wherein multiplying the input data value with the shared one of the weighting values comprises generating a carry value within a first one of the MAC circuits of the given one of the MAC units, the method further comprising receiving the carry value within a second one of the MAC circuits of the given one of the MAC units.

20 . The method of claim 11 wherein multiplying the input data value received during each of the timing cycles with the shared one of the weighting data values comprises multiplying the input data value with the shared one of the weighting data values during a timing cycle that transpires after reception of the input data value and shared one of the weighting values.

21 . An integrated circuit component comprising:

a plurality of broadcast data paths;

a weighting-value memory;

a plurality of multiply-accumulate (MAC) units coupled in common to each of the broadcast data paths and coupled to receive respective weighting values from the weighting-value memory via respective weighting-value paths, each of the MAC units having a plurality of MAC circuits coupled respectively to the broadcast data paths, each of the MAC circuits within a given one of the MAC units having:

means for receiving, during each of a plurality of timing cycles, (i) an input data value via a respective one of the broadcast data paths, and (ii) a shared one of the weighting values via a shared one of the respective weighting-value paths;

means for generating a sequence of multiplication products by multiplying the input data value received during each of the plurality of timing cycles with the shared one of the weighting values received during each of the plurality of timing cycles; and

means for accumulating a sum of constituent multiplication products within the sequence of multiplication products; and

means for receiving respective instances of the sum of constituent multiplication products from the accumulator circuits within the plurality of MAC units and for sequentially outputting the respective instances of the sum of constituent multiplication products over a number of the timing cycles that exceeds a number of MAC units included within the plurality of MAC units.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2024
From: FLEX LOGIX TECHNOLOGIES, INC.
To: ANALOG DEVICES, INC.
Reel/Frame 069409/0035 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2023
From: WARE, FREDERICK A.; WANG, CHENG C.
To: FLEX LOGIX TECHNOLOGIES, INC.
Reel/Frame 063897/0837 →