IP Library › Granted Patent US 12,705,468
Granted Patent B2
US 12,705,468 · App. 17/813,834 · Granted Aug 11, 2026

Hybrid machine learning architecture with neural processing unit and compute-in-memory processing elements

Inventors: Mustafa Badaroglu (Leuven, BE); Zhongze Wang (San Diego, CA); Titash Rakshit (Austin, TX)
Assignee: QUALCOMM Incorporated
G06N3/063G06F15/80G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,468
App. No.
17/813,834
Filed
Jul 20, 2022
Granted
Aug 11, 2026
Kind
B2
Art Unit
2128
USPC
706/15
Abstract

Methods and apparatus for performing machine learning tasks, and in particular, a hybrid architecture that includes both neural processing unit (NPU) and compute-in-memory (CIM) elements. One example neural-network-processing circuit generally includes a plurality of CIM processing elements (PEs), a plurality of neural processing unit (NPU) PEs, and a bus coupled to the plurality of CIM PEs and to the plurality of NPU PEs. One example method for neural network processing generally includes processing data in a neural-network-processing circuit comprising a plurality of CIM PEs, a plurality of NPU PEs, and a bus coupled to the plurality of CIM PEs and to the plurality of NPU PEs; and transferring the processed data between at least one of the plurality of CIM PEs and at least one of the plurality of NPU PEs via the bus.

Claims (24)

1 . A neural-network-processing circuit comprising:

a plurality of compute-in-memory (CIM) processing elements (PEs);

a plurality of neural processing unit (NPU) PEs;

a bus coupled to the plurality of CIM PEs and to the plurality of NPU PEs;

a global memory, wherein at least one of the plurality of CIM PEs is configured to transfer data to or receive data from at least one of the plurality of NPU PEs without the data being written to or being read from the global memory;

bus arbitration logic coupled between the bus and the plurality of CIM PEs and between the bus and the plurality of NPU PEs;

a first first-in, first-out (FIFO) circuit coupled between the bus arbitration logic and the plurality of CIM PEs; and

a second FIFO circuit coupled between the bus arbitration logic and the plurality of NPU PEs.

2 . The neural-network-processing circuit of claim 1 , wherein the at least one of the plurality of CIM PEs is in a same neural network layer as the at least one of the plurality of NPU PEs.

3 . The neural-network-processing circuit of claim 1 , wherein the at least one of the plurality of CIM PEs is in a first neural network layer and wherein the at least one of the plurality of NPU PEs is in a second neural network layer, different from the first neural network layer.

4 . The neural-network-processing circuit of claim 3 , wherein the second neural network layer is adjacent to the first neural network layer.

5 . The neural-network-processing circuit of claim 1 , wherein the plurality of CIM PEs are configured as pseudo-weight-stationary PEs.

6 . The neural-network-processing circuit of claim 1 , wherein the plurality of CIM PEs are configured as digital compute-in-memory (DCIM) PEs.

7 . The neural-network-processing circuit of claim 1 , wherein the plurality of NPU PEs are configured as output-stationary PEs.

8 . The neural-network-processing circuit of claim 1 , further comprising a digital processing circuit coupled between the bus arbitration logic and the plurality of CIM PEs and between the bus arbitration logic and the plurality of NPU PEs.

9 . A neural-network-processing circuit comprising:

a plurality of compute-in-memory (CIM) processing elements (PEs);

a plurality of neural processing unit (NPU) PEs;

a bus coupled to the plurality of CIM PEs and to the plurality of NPU PES;

a global memory, wherein at least one of the plurality of CIM PEs is configured to transfer data to or receive data from at least one of the plurality of NPU PEs without the data being written to or being read from the global memory;

bus arbitration logic coupled between the bus and the plurality of CIM PEs and between the bus and the plurality of NPU PEs;

a digital processing circuit coupled between the bus arbitration logic and the plurality of CIM PEs and between the bus arbitration logic and the plurality of NPU PEs;

a first first-in, first-out (FIFO) circuit coupled between the digital processing circuit and the plurality of CIM PEs; and

a second FIFO circuit coupled between the digital processing circuit and the plurality of NPU PEs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2022
From: BADAROGLU, MUSTAFA; WANG, ZHONGZE; RAKSHIT, TITASH
To: QUALCOMM INCORPORATED
Reel/Frame 060812/0254 →
Continuity (2)
Provisional Application 63224155 · Jul 21, 2021
Related Publication 20230025068A1 · Jan 26, 2023
References Cited (18)
US 10725946B1 · Berke · 2020 [cited by examiner]
US 10839894B2 · Chen et al. · 2020 [cited by applicant]
US 20130107872A1 · Lovett · 2013 [cited by examiner]
US 20180315473A1 · Yu · 2018 [cited by examiner]
US 20190042199A1 · Sumbul · 2019 [cited by examiner]
US 20190042949A1 · Young · 2019 [cited by examiner]
US 20190102170A1 · Chen · 2019 [cited by examiner]
US 20190281587A1 · Zhang · 2019 [cited by examiner]
US 20210005230A1 · Wang · 2021 [cited by examiner]
US 20210073619A1 · Wang · 2021 [cited by examiner]
US 20210073650A1 · Reisser · 2021 [cited by examiner]
US 20210089865A1 · Wang et al. · 2021 [cited by applicant]
US 20220351032A1 · Chou · 2022 [cited by examiner]
Chih, Y-D., et al., “16.4 An 89TOPS/W and 16.3TOPS/mm2 All-Digital SRAM-Based Full-Precision Compute-In Memory Macro in 22nm for Machine-Learning Edge Applications”, 2021 IEEE International Solid-State Circuits Conferen… [cited by applicant]
Kang M., et al., “Deep In-Memory Architectures in SRAM: An Analog Approach to Approximate Computing”, Proceedings of the IEEE, vol. 108, No. 12, Dec. 2020, DOI: 10.1109/JPROC.2020.3034117, pp. 2251-2275. [cited by applicant]
International Search Report and Written Opinion—PCT/US2022/073979—ISA/EPO—Nov. 24, 2022. [cited by applicant]
Zhou K., et al., “Domino: A Tailored Network-on-Chip Architecture to Enable Highly Localized Inter- and Intra-Memory DNN Computing”, arxiv.org, Cornell University Library, 201 OLIN Library Cornell University Ithaca, NY … [cited by applicant]
Zhou K., et al., “Domino: A Tailored Network-on-Chip Architecture to Enable Highly Localized Inter-and Intra-Memory DNN Computing”, ArXiv:2107.09500v1 [cs.AR], Jul. 18, 2021, pp. 1-13. [cited by applicant]