IP Library › Granted Patent US 12,333,405
Granted Patent B2
US 12,333,405 · App. 17/035,879 · Granted Jun 17, 2025

Methods and apparatus for matrix processing in a convolutional neural network

Inventors: Mihir Narendra Mody (Bangalore, IN); Shyam Jagannathan (Bangalore, IN); Manu Mathew (Bangalore, IN); Jason T. Jones (Richmond, TX)
Assignee: TEXAS INSTRUMENTS INCORPORATED
G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,405
App. No.
17/035,879
Granted
Jun 17, 2025
Kind
B2
Abstract

Described examples include an integrated circuit including a vector multiply unit including a plurality of multiply/accumulate nodes, in which the vector multiply unit is operable to provide an output from the multiply/accumulate nodes, a first data feeder operable to provide first data to the vector multiply unit in vector format, and a second data feeder operable to provide second data to the vector multiply unit in vector format.

Claims (46)

1. A method comprising:

receiving first data from a memory;

providing, by a first data feeder, a portion of the first data to a vector multiply unit in vector format by providing a number of feature planes K to the vector multiply unit, the feature planes including at least a first feature plane and a second feature plane;

providing, by a second data feeder, second data including a first set of weight values to the vector multiply unit by at least providing a corresponding weight from each convolution filter of a set of convolution filters to the vector multiply unit in vector format, the first set of weight values including multiple sets of weight da each set of weight data including L weight coefficients;

multiplying and accumulating the first data and the second data as a partial product in at least one node of the vector multiply unit, such that a product is accumulated in the at least one node of the vector multiply unit after a plurality of iterations by at least:

multiplying the first set of weight values by each pixel value in a first subset of pixel positions for the first feature plane to generate a first result;

storing the first result in a first accumulator of the node;

multiplying the first set of weight values by each pixel value in the first subset of pixel positions for the second feature plane to generate a second result; and

storing the second result in a second accumulator of the node; and

rearranging, using the first data feeder and the second data feeder, the portion of the first data and the second data so that partial products computed by the vector multiply unit for one output feature value are directed to one multiplier and one accumulator within one node of the vector multiply unit;

reading the first and second results from the node after a delay of at least K×L cycles after a previous iteration of the reading step; and

storing the first and second results in the memory.

2. The method of claim 1 in which the multiplying and accumulating provides an outer product of the first data and the second data in a clock cycle.

3. The method of claim 1 in which the first data is K feature planes and N values are read from the K feature planes in a clock cycle of the vector multiply unit.

4. The method of claim 3 in which the second data is multiple sets of weight data with L weight coefficients and output from the vector multiply unit is provided after K×L cycles of the vector multiply unit.

5. The method of claim 1 in which an output of the vector multiply unit is a convolution layer and is provided after a fixed number of cycles of the vector multiply unit.

6. The method of claim 1 in which the first data is one row of length N, in which the second data is one column of length M that are read in a first cycle and multiplied and accumulated in N×M nodes in the vector multiply unit in a second cycle and in which, after the plurality of iterations, the N×M nodes are a results matrix that is at least part of a complete results matrix.

7. The method of claim 1 , wherein, after the first and second results are read from the node, the first and second results are processed by one or more convolution layers or output layers prior to being stored in the memory.

8. A device comprising:

a memory;

a vector multiply unit coupled to the memory;

a first data feeder configured to provide first data to the vector multiply unit in vector format by providing a number of feature planes K to the vector multiply unit, wherein the feature planes include at least a first feature plane and a second feature plane;

a second data feeder configured to provide second data including a first set of weight values to the vector multiply unit by at least providing a corresponding weight from each convolution filter of a set of convolution filters to the vector multiply unit in vector format,

wherein the first set of weight values includes multiple sets of weight data, each set of weight data including L weight coefficients,

wherein the first data feeder and the second data feeder are configured to rearrange the first data and the second data so that partial products computed by the vector multiply unit for one output feature value are directed to one multiplier and one accumulator within one node of the vector multiply unit,

wherein the vector multiply unit is configured to multiply and accumulate the first data and the second data as a partial product in at least one node of the vector multiply unit, such that a product is accumulated in the at least one node of the vector multiply unit after a plurality of iterations by at least:

multiplying the first set of weight values by each pixel value in a first subset of pixel positions for the first feature plane to generate a first result;

storing the first result in a first accumulator of the node;

multiplying the first set of weight values by each pixel value in the first subset of pixel positions for the second feature plane to generate a second result; and

storing the second result in a second accumulator of the node; and

wherein the device is configured to:

read the first and second results from the node after a delay of at least K×L cycles after a previous iteration of the read action; and

store the first and second results in the memory.

9. The device of claim 8 , wherein the vector multiply unit is configured to perform the multiplying and accumulating to provide an outer product of the first data and the second data in a clock cycle.

10. The device of claim 8 ,

wherein the first data is K feature planes, and

wherein the device is configured to read N values from the K feature planes in a clock cycle of the vector multiply unit.

11. The device of claim 10 ,

wherein the second data is multiple sets of weight data with L weight coefficients, and

wherein the vector multiply unit is configured to provide an output after K×L cycles of the vector multiply unit.

12. The device of claim 8 , wherein the vector multiply unit is configured to provide an output after a fixed number of cycles of the vector multiply unit.

13. The device of claim 8 ,

wherein the first data is one row of length N,

wherein the second data is one column of length M that is read in a first cycle and multiplied and accumulated in N×M nodes in the vector multiply unit in a second cycle, and

wherein, after the plurality of iterations, the N×M nodes are a results matrix that is at least part of a complete results matrix.

14. The device of claim 8 , wherein, after the first and second results are read from the node, one or more convolution layers or output layers are configured to process the first and second results prior to being stored in the memory.

Continuity (3)
Continuation 15784588 · Oct 16, 2017
Provisional Application 62445493 · Jan 12, 2017
Related Publication 20220101083A1 · Mar 31, 2022
References Cited (16)
US 9384168B2 · Mortensen · 2016 [cited by applicant]
US 20040151356A1 · Li et al. · 2004 [cited by applicant]
US 20050125369A1 · Buck et al. · 2005 [cited by applicant]
US 20050197977A1 · Buck et al. · 2005 [cited by applicant]
US 20070028076A1 · Wezelenburg · 2007 [cited by applicant]
US 20160259826A1 · Acar · 2016 [cited by applicant]
US 20160342890A1 · Young · 2016 [cited by applicant]
US 20160342891A1 · Ross et al. · 2016 [cited by applicant]
US 20160358069A1 · Brothers et al. · 2016 [cited by applicant]
US 20170221176A1 · Munteanu · 2017 [cited by examiner]
Hijazi, Samer, Rishi Kumar, and Chris Rowen. “Using convolutional neural networks for image recognition.” Cadence Design Systems Inc.: San Jose, CA, USA 9 (2015): 1. (Year: 2015). [cited by examiner]
Moons et al. A 0.3-2.6 TOPS/W Precision-Scalable Processor for Real-Time Large-Scale ConvNets. 2016 (Year: 2016). [cited by examiner]
Athale et al. Optical Processing Using Outer-Product Concepts. 1984. (Year: 1984). [cited by examiner]
Rahman et al. Efficient FPGA Acceleration of Convolutional Neural Networks Using Logical-3D Compute Array. IEEE 2016 (Year: 2016). [cited by examiner]
Sankaradas et al., A Massively Parallel Coprocessor for Convolutional Neural Networks. 2009 20th IEEE International Conference on Application-specific Systems, Architectures and Processors (Year: 2009). [cited by examiner]
International Search Report for PCT/US2018/013586 mailed Jun. 28, 2018. [cited by applicant]
Cited By (1)
US 12,579,611