IP Library Granted Patent US 12,705,468
Granted Patent B2
US 12,705,468 · App. 17/813,834 · Granted Aug 11, 2026

Hybrid machine learning architecture with neural processing unit and compute-in-memory processing elements

Inventors: Mustafa Badaroglu (Leuven, BE); Zhongze Wang (San Diego, CA); Titash Rakshit (Austin, TX)
Assignee: QUALCOMM Incorporated
G06N3/063G06F15/80G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,468
App. No.
17/813,834
Granted
Aug 11, 2026
Kind
B2
Abstract

Methods and apparatus for performing machine learning tasks, and in particular, a hybrid architecture that includes both neural processing unit (NPU) and compute-in-memory (CIM) elements. One example neural-network-processing circuit generally includes a plurality of CIM processing elements (PEs), a plurality of neural processing unit (NPU) PEs, and a bus coupled to the plurality of CIM PEs and to the plurality of NPU PEs. One example method for neural network processing generally includes processing data in a neural-network-processing circuit comprising a plurality of CIM PEs, a plurality of NPU PEs, and a bus coupled to the plurality of CIM PEs and to the plurality of NPU PEs; and transferring the processed data between at least one of the plurality of CIM PEs and at least one of the plurality of NPU PEs via the bus.

Claims (24)

1 . A neural-network-processing circuit comprising:

a plurality of compute-in-memory (CIM) processing elements (PEs);

a plurality of neural processing unit (NPU) PEs;

a bus coupled to the plurality of CIM PEs and to the plurality of NPU PEs;

a global memory, wherein at least one of the plurality of CIM PEs is configured to transfer data to or receive data from at least one of the plurality of NPU PEs without the data being written to or being read from the global memory;

bus arbitration logic coupled between the bus and the plurality of CIM PEs and between the bus and the plurality of NPU PEs;

a first first-in, first-out (FIFO) circuit coupled between the bus arbitration logic and the plurality of CIM PEs; and

a second FIFO circuit coupled between the bus arbitration logic and the plurality of NPU PEs.

2 . The neural-network-processing circuit of claim 1 , wherein the at least one of the plurality of CIM PEs is in a same neural network layer as the at least one of the plurality of NPU PEs.

3 . The neural-network-processing circuit of claim 1 , wherein the at least one of the plurality of CIM PEs is in a first neural network layer and wherein the at least one of the plurality of NPU PEs is in a second neural network layer, different from the first neural network layer.

4 . The neural-network-processing circuit of claim 3 , wherein the second neural network layer is adjacent to the first neural network layer.

5 . The neural-network-processing circuit of claim 1 , wherein the plurality of CIM PEs are configured as pseudo-weight-stationary PEs.

6 . The neural-network-processing circuit of claim 1 , wherein the plurality of CIM PEs are configured as digital compute-in-memory (DCIM) PEs.

7 . The neural-network-processing circuit of claim 1 , wherein the plurality of NPU PEs are configured as output-stationary PEs.

8 . The neural-network-processing circuit of claim 1 , further comprising a digital processing circuit coupled between the bus arbitration logic and the plurality of CIM PEs and between the bus arbitration logic and the plurality of NPU PEs.

9 . A neural-network-processing circuit comprising:

a plurality of compute-in-memory (CIM) processing elements (PEs);

a plurality of neural processing unit (NPU) PEs;

a bus coupled to the plurality of CIM PEs and to the plurality of NPU PES;

a global memory, wherein at least one of the plurality of CIM PEs is configured to transfer data to or receive data from at least one of the plurality of NPU PEs without the data being written to or being read from the global memory;

bus arbitration logic coupled between the bus and the plurality of CIM PEs and between the bus and the plurality of NPU PEs;

a digital processing circuit coupled between the bus arbitration logic and the plurality of CIM PEs and between the bus arbitration logic and the plurality of NPU PEs;

a first first-in, first-out (FIFO) circuit coupled between the digital processing circuit and the plurality of CIM PEs; and

a second FIFO circuit coupled between the digital processing circuit and the plurality of NPU PEs.