IP Library › Granted Patent US 12,530,170
Granted Patent B2
US 12,530,170 · App. 18/638,441 · Granted Jan 20, 2026

Vector operation acceleration with convolution computation unit

Inventors: Xiaoqian Zhang (San Jose, CA); Zhibin Xiao (Los Altos, CA); Changxu Zhang (Santa Clara, CA); Renjie Chen (Mountain View, CA)
Assignee: Moffett International Co., Limited
G06F7/5443G06F7/50
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,170
App. No.
18/638,441
Granted
Jan 20, 2026
Kind
B2
Abstract

This application describes hybrid hardware accelerators, systems, and apparatus for performing various computations in neural network applications using the same set of hardware resources. An example accelerator may include weight selectors, activation input interfaces, and a plurality of Multiplier-Accumulation (MAC) circuits organized as a plurality of MAC lanes Each of the plurality of MAC lanes may be configured to: receive a control signal indicating whether to perform convolution or vector operations; receive one or more weights according to the control signal; receive one or more activations according to the control signal; and generate output data based on the one or more weights and the one or more input activations according to the control signal and feed the output data into an output buffer. Each of the plurality of MAC lanes includes a plurality of multiplier circuits and a plurality of adder-subtractor circuits.

Claims (72)

1 . A dual-operation neural network accelerator, comprising:

a plurality of Multiplier-Accumulation (MAC) circuits organized as a plurality of MAC lanes, wherein each of the plurality of MAC lanes comprises:

a plurality of first circuits for performing multiplication operations; and

a plurality of second circuits organized as a tree comprising a root level and a leaf level, the root level comprising one second circuit and the leaf level comprising multiple second circuits, wherein:

the multiple second circuits on the leaf level are configured to receive data from the plurality of first circuits and perform addition or subtraction operations on the received data, and feed results of the addition or subtraction operations towards the root level; and

the second circuit on the root level is configured to generate an output based on results of the addition or subtraction operations, and send the output to an adder-subtractor circuit outside of the tree of the plurality of second circuits, wherein the adder-subtractor circuit is configured to read previous data from an output buffer, compute a new output based on the previous data and the output received from the second circuit on the root level, and write the new output to the output buffer;

wherein each of the plurality of second circuits is a combination of an accumulator and a subtractor, and is configured to perform both addition and subtraction using a same circuit;

wherein:

each of the plurality of MAC lanes is configured to (1) fetch weights from a weight cache for convolution operations and (2) obtain weights generated by a weight matrix generating circuit for vector operations, and

the weight matrix generating circuit is configured to internally generate a row of weights for one of the plurality of MAC lanes.

2 . The dual-operation neural network accelerator of claim 1 , wherein each of the plurality of second circuits comprises an adder-subtractor logic gate configurable to perform the addition or subtraction operations.

3 . The dual-operation neural network accelerator of claim 2 , wherein the plurality of second circuits are configured to receive a first control signal and a second control signal, wherein:

the first control signal configures the adder-subtractor logic gate to perform addition or subtraction, and

the second control signal configures the adder-subtractor logic gate to post-process a result of the addition or subtraction.

4 . The dual-operation neural network accelerator of claim 1 , wherein the second circuit on the root level is further configured to:

read a previous output from a buffer of the dual-operation neural network accelerator that was generated from a previous iteration, and

generating the output based on the previous output and the results of the addition or subtraction operations.

5 . The dual-operation neural network accelerator of claim 1 , further comprising:

a plurality of weight selectors configured to obtain weights, and

a plurality of activation input interfaces configured to obtain activations,

wherein each of the plurality of MAC lanes is configured to:

receive a control signal indicating whether to perform the convolution or the vector operations;

receive one or more weights from at least one of the plurality of weight selectors according to the control signal;

receive one or more activations from at least one of the plurality of activation input interfaces according to the control signal; and

input the one or more weights and the one or more activations into the MAC lane of the plurality of MAC lanes to generate output data according to the control signal.

6 . The dual-operation neural network accelerator of claim 5 , wherein the vector operations comprise one or more of reduce mean, reduce minimum, reduce maximum, reduce average, reduce add, or pooling.

7 . The dual-operation neural network accelerator of claim 6 , wherein each of the plurality of weight selectors comprises a multiplexer coupled with the weight matrix generating circuit and the weight cache.

8 . The dual-operation neural network accelerator of claim 7 , wherein each of the plurality of weight selectors is configured to:

in response to the control signal indicating performing the convolution operations, obtain a weight from the weight cache; and

in response to the control signal indicating performing vector computation, obtain a weight from the weight matrix generating circuit.

9 . The dual-operation neural network accelerator of claim 7 , wherein a first subset of the plurality of MAC lanes are configured to receive weights from the weight cache and a second subset of the plurality of MAC lanes are configured to receive weights generated by the weight matrix generating circuit, and

the first subset of the plurality of MAC lanes are further configured to perform convolution computations and the second subset of the plurality of MAC lanes are further configured to perform vector operations, and the convolution computations and the vector operations are performed in parallel.

10 . The dual-operation neural network accelerator of claim 1 , wherein each of the plurality of second circuits comprises a plurality of multiplexers for selecting signals from multiple inputs based on control signals, wherein the control signals comprise a first control signal indicating whether an addition or a subtraction is being performed, and a second control signal indicating whether to obtain a min or max value from input values.

11 . A Multiplier-Accumulation (MAC) lane in a neural network accelerator, comprising:

a plurality of first circuits for performing multiplication operations; and

a plurality of second circuits organized as a tree comprising a root level and a leaf level, the root level comprising one second circuit and the leaf level comprising multiple second circuits, wherein:

the multiple second circuits on the leaf level are configured to receive data from the plurality of first circuits and perform addition or subtraction operations on the received data, and feed results of the addition or subtraction operations towards the root level; and

the second circuit on the root level is configured to generate an output based on results of the addition or subtraction operations, and send the output to an adder-subtractor circuit outside of the tree of the plurality of second circuits, wherein the adder-subtractor circuit is configured to read previous data from an output buffer, compute a new output based on the previous data and the output received from the second circuit on the root level, and write the new output to the output buffer;

wherein each of the plurality of second circuits is a combination of an accumulator and a subtractor, and is configured to perform both addition and subtraction using a same circuit;

wherein:

the MAC lane is configured to (1) fetch weights from a weight cache for convolution operations and (2) obtain weights generated by a weight matrix generating circuit for vector operations, and

the weight matrix generating circuit is configured to internally generate a row of weights for the MAC lane.

12 . The MAC lane of claim 11 , wherein each of the plurality of second circuits comprises an adder-subtractor logic gate configurable to perform the addition or subtraction operations.

13 . The MAC lane of claim 12 , wherein the plurality of second circuits are configured to receive a first control signal and a second control signal, wherein:

the first control signal configures the adder-subtractor logic gate to perform addition or subtraction, and

the second control signal configures the adder-subtractor logic gate to post-process a result of the addition or subtraction.

14 . The MAC lane of claim 11 , wherein the second circuit on the root level is further configured to:

read a previous output from a buffer of the neural network accelerator that was generated from a previous iteration, and

generating the output based on the previous output and the results of the addition or subtraction operations.

15 . The MAC lane of claim 11 , wherein the MAC lane is coupled with a plurality of weight selectors configured to obtain weights, and a plurality of activation input interfaces configured to obtain activations.

16 . The MAC lane of claim 15 , wherein the plurality of first circuits are configured to:

receive a control signal indicating whether to perform convolution or vector operations;

receive one or more weights from at least one of the plurality of weight selectors according to the control signal; and

receive one or more activations from at least one of the plurality of activation input interfaces according to the control signal.

17 . A computer-implemented method, comprising:

receiving a control signal indicating whether to perform convolution or vector operations;

receiving one or more weights from at least one of a plurality of weight selectors according to the control signal;

receiving one or more activations from at least one of a plurality of activation input interfaces according to the control signal; and

inputting the one or more weights and the one or more activations into a Multiplier-Accumulation (MAC) lane to generate output data according to the control signal,

wherein the MAC lane comprises:

a plurality of first circuits for performing multiplication operations,

a plurality of second circuits organized as a tree comprising a root level and a leaf level, the root level comprising one second circuit and the leaf level comprising multiple second circuits,

the multiple second circuits on the leaf level are configured to receive data from the plurality of first circuits and perform addition or subtraction operations on the received data, and feed results of the addition or subtraction operations towards the root level, and

the second circuit on the root level is configured to generate the output data based on results of the addition or subtraction operations, and send the output to an adder-subtractor circuit outside of the tree of the plurality of second circuits, wherein the adder-subtractor circuit is configured to read previous data from an output buffer, compute a new output based on the previous data and the output received from the second circuit on the root level, and write the new output to the output buffer;

wherein each of the plurality of second circuits is a combination of an accumulator and a subtractor, and is configured to perform both addition and subtraction using a same circuit;

wherein:

the MAC lane is configured to (1) fetch weights from a weight cache for convolution operations and (2) obtain weights generated by a weight matrix generating circuit for vector operations, and

the weight matrix generating circuit is configured to internally generate a row of weights for the MAC lane.

18 . The computer-implemented method of claim 17 , wherein each of the plurality of second circuits comprises an adder-subtractor logic gate configurable to perform the addition or subtraction operations.

19 . The computer-implemented method of claim 18 , wherein the plurality of second circuits are configured to receive a first control signal and a second control signal, wherein:

the first control signal configures the adder-subtractor logic gate to perform addition or subtraction, and

the second control signal configures the adder-subtractor logic gate to post-process a result of the addition or subtraction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2024
From: ZHANG, XIAOQIAN; XIAO, ZHIBIN; ZHANG, CHANGXU; CHEN, RENJIE
To: MOFFETT INTERNATIONAL CO., LIMITED
Reel/Frame 067185/0520 →
Continuity (3)
Continuation 18130311 · Apr 3, 2023
Continuation 17944772 · Sep 14, 2022
Related Publication 20240264802A1 · Aug 8, 2024
References Cited (41)
US 11347916B1 · Desai et al. · 2022 [cited by applicant]
US 11494627B1 · Li et al. · 2022 [cited by applicant]
US 11726746B1 · Zhang et al. · 2023 [cited by applicant]
US 20150130819A1 · Peng · 2015 [cited by examiner]
US 20170316312A1 · Goyal et al. · 2017 [cited by applicant]
US 20170357891A1 · Judd · 2017 [cited by examiner]
US 20180307976A1 · Fang et al. · 2018 [cited by applicant]
US 20190065188A1 · Shippy et al. · 2019 [cited by applicant]
US 20200050918A1 · Chen · 2020 [cited by examiner]
US 20200074293A1 · Chin · 2020 [cited by examiner]
US 20200089506A1 · Power et al. · 2020 [cited by applicant]
US 20200104167A1 · Chen et al. · 2020 [cited by applicant]
US 20200394516A1 · Chen et al. · 2020 [cited by applicant]
US 20210208884A1 · Song · 2021 [cited by applicant]
US 20210295140A1 · Holm et al. · 2021 [cited by applicant]
US 20210319290A1 · Mills et al. · 2021 [cited by applicant]
US 20210326683A1 · Narayanaswami et al. · 2021 [cited by applicant]
US 20210357748A1 · Shin · 2021 [cited by examiner]
US 20220035629A1 · Kwon · 2022 [cited by examiner]
US 20220036243A1 · Das et al. · 2022 [cited by applicant]
US 20220101083A1 · Mody et al. · 2022 [cited by applicant]
US 20220129320A1 · Mohapatra · 2022 [cited by examiner]
US 20220180187A1 · Kim · 2022 [cited by examiner]
US 20220188073A1 · Bowman et al. · 2022 [cited by applicant]
US 20220253716A1 · Choudhury et al. · 2022 [cited by applicant]
US 20230012553A1 · Ahmadi · 2023 [cited by examiner]
US 20230062217A1 · Cassidy et al. · 2023 [cited by applicant]
US 20230418559A1 · Rossi · 2023 [cited by examiner]
CN 112204579A · 2021 [cited by applicant]
CN 113692592A · 2021 [cited by applicant]
CN 114402337A · 2022 [cited by applicant]
EP 3555814A1 · 2019 [cited by applicant]
Non-Final Office Action dated Dec. 16, 2022, issued in related U.S. Appl. No. 17/944,772 (16 pages). [cited by applicant]
Notice of Allowance mailed Mar. 24, 2023, issued in related U.S. Appl. No. 17/944,772 (13 pages). [cited by applicant]
PCT International Search Report and the Written Opinion mailed Dec. 1, 2023, issued in related International Application No. PCT/CN2023/116979 (10 pages). [cited by applicant]
Non-Final Office Action dated Dec. 22, 2023, issued in related U.S. Appl. No. 18/130,311 (18 pages). [cited by applicant]
Wu, Ning. “A Reconfigurable Convolutional Neural Network-Accelerated Coprocessor Based on RISC-V Instruction Set.” Www. Researchgate. Net/, 2020, www.researchgate.net/publication/342227564_A_Reconfigurable_Convolutional… [cited by applicant]
Extender European Search Report dated Sep. 26, 2025, issued in European Patent Application No. 23864604.6 (11 pages). [cited by applicant]
Tianshi Chen et al., “DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning,” ASPLOS '14, Mar. 1, 2024, pp. 269-283. [cited by applicant]
Li Du et al., “A Reconfigurable Streaming Deep Convolutional Neural Network Accelerator for Internet of Things,” Aug 16, 2017, Retrieved from the Internet: URL:https://arxiv.org/pdf/1707.02973, retrieved on Sep. 8, 2025… [cited by applicant]
Zidong Du et al., “ShiDianNao: Shifting Vision Processing Closer to the Sensor,” Jun. 17, 2015, Retrieved from the Internet: URL:https://ieeexplore.ieee.org/document/7284058, retrieved on Sep. 8, 2025, pp. 92-104. [cited by applicant]