IP Library Granted Patent US 12699521
Granted Patent B2
US 12699521 · App. 18/534,487 · Granted Aug 4, 2026

Core group memory processing chip design

Inventors: Timothy Wesley (Ann Arbor, MI); Jacob Botimer (Ann Arbor, MI); Mohammed Zidan (Ann Arbor, MI); Chester Liu (Ann Arbor, MI); Wei Lu (Ann Arbor, MI)
Assignee: MemryX Incorporated
G06F3/064G06F3/0604G06F3/0679
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12699521
App. No.
18/534,487
Granted
Aug 4, 2026
Kind
B2
Abstract

A memory processing unit (MPU) includes processing regions and a first memory having memory regions configured in memory blocks. The processing regions are interleaved between the memory regions and include core groups including compute cores. One or more of the core groups of a respective one of the processing regions are coupled between adjacent ones of the memory blocks of adjacent ones of the memory regions. Components of the MPU and dataflow between adjacent compute cores in respective ones of the processing regions through one or more corresponding memory blocks of corresponding memory regions, are configured based on one or more of an impact of a number of parameters and a number of layers of one or more neural network models on memory balance, core utilization and memory reuse on bandwidth balance, design points, supported operations, target models, and supported memory organizations.

Claims (25)

1 . A method of designing a memory processing unit (MPU), wherein the MPU includes:

a first memory including a plurality of memory regions, wherein one or more of the plurality of memory regions are configured in a corresponding pluralities of memory blocks; and

a plurality of processing regions, wherein each processing region of the plurality of processing regions is interleaved between two adjacent memory regions of the plurality of memory regions of the first memory, wherein the processing regions include a plurality of core groups, wherein the core groups include one or more compute cores, and wherein one or more of the plurality of core groups of a respective one of the plurality of processing regions are coupled between adjacent ones of the memory blocks of adjacent ones of the plurality of memory regions of the first memory;

the method comprising configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of corresponding memory regions, based on one or more of an impact of a number of parameters and a number of layers of one or more neural network models on memory balance, core utilization and memory reuse on bandwidth balance, design points, supported operations, target models, and supported memory organizations.

2 . The method according to claim 1 , wherein:

the MPU further includes a second memory coupled to the plurality of processing regions configured to store weight values; and

configuring the second memory, and the dataflow between the second memory and compute cores in respective ones of the plurality of processing regions based on one or more of the impact of the number of parameters and the number of layers of one or more neural network models on memory balance, includes determining the number of layers based on utilization of the second memory storing weight values.

3 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and the dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on one or more of the impact of the number of parameters and the number of layers of one or more neural network models on memory balance, includes determining a number of dies for implementing the MPU based on the impact of the number of parameters and the number of layers.

4 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and the dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on one or more of the impact of the number of parameters and the number of layers of one or more neural network models on memory balance, includes determining a number of physical channels of the first memory based on a width of the first memory, a granularity of the one or more compute cores, and the number of cores per core groups.

5 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and the dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on one or more of the impact of core utilization and memory reuse on bandwidth balance, includes determining weight data sharing between cores, feature map data reuse, amount of the second memory and number of cores per core groups.

6 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on design points, includes determining a number of processing regions, a number of memory regions of the first memory, a number of core groups per processing regions, the number of memory blocks per memory regions, a number of inter-layer communication units per memory regions, a number of input ports coupled to a first predetermined first memory region, a number of output ports coupled to a second predetermined first memory region, a total memory size of the first memory and a total memory size of the second memory.

7 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on design points, includes determining one or more parameters of the core groups including a number of near memory cores, a number of arithmetic cores, a size of the second memory, a data bus size of the second memory, an address size of the second memory, a number of arbiters of the second memory, a number of read ports of the memory regions of the first memory, a number of write ports of the memory regions of the first memory, a data size per port of the memory regions of the first memory, an address size per port of the memory regions of the first memory, a number of arbiters of the first memory, a number of ports of the inter-layer communication unit, and a number of arbiters of the inter-layer communication unit.

8 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on design points, includes determining one or more parameters of the memory regions of the first memory including a size of the memory regions, a number of read ports of the memory regions of the first memory, a number of parallel accesses of the memory regions of the first memory, a number of banks of the memory regions of the first memory, an address domain of the memory regions of the first memory, an address granularity of the memory regions of the first memory, an address width of the memory regions of the first memory, a data width of the memory regions of the first memory, and a latency of the memory regions of the first memory.

9 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on design points, includes determining one or more parameters of the memory regions of the second memory including a size of the second memory, a number of entries of the second memory, a number of read ports of the second memory, a number of write ports of the second memory, a number of parallel access ports, a number of banks of the second memory, an address domain of the second memory, an address granularity of the second memory, an address width of the ports of the second memory, a data width of the ports of the second memory, and a latency of the second memory.

10 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on design points, includes determining one or more parameters of the near memory cores including a number of multiply accumulation units, a number of accumulators, a number of physical channels, a second memory buffer size, a first memory buffer size, a number of ports for the second memory, a number of fetch ports for the first memory, a number of writeback ports for the first memory, a number of inter-layer communication ports, a multiply accumulator sequence, a zero skipping, and a group B-float encoding size.

11 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on design points, includes determining one or more parameters of the arithmetic cores, including determining a number of arithmetic logic units, a number of physical channels, a first memory output buffer entry size, a first memory input buffer entry size, a number of fetch ports for the first memory, and a number of inter-layer communication ports.

12 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on design points, includes determining one or more parameters of the inter-layer communication units including a number of port, a number of parallel accesses, a number of layers per inter-layer-communication unit, and bit size of an index, an index domain, a counter bit width, and an access type.

13 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on supported operations, includes determining supported two-dimensional (2D) convolution operations, supported fully connected convolution operations, supported scaling operations, supported branch and merge operations, supported activation operations, and flattening, reshape, batch normalization, rescaling, constate add, constant multiply and clip operations.

14 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on one or more of target models, includes determining MobileNet models, ResNet Models, SSD model, Tiny YOLO models, YOLO models, OpenPose model, Xception model, Inception model, and EfficientNet model.

15 . The method according to claim 2 , wherein configuring one or more of the plurality of memory regions of the first memory, the memory blocks of the memory regions, the plurality of processing regions, the plurality of core groups, the compute cores, and dataflow between adjacent compute cores in respective ones of the plurality of processing regions through one or more corresponding memory blocks of the corresponding memory regions based on memory organizations, includes determining feature map value encoding, feature map value packing in the first memory, weight value encoding and weigh value packing in the second memory.

16 . The MPU of claim 1 , wherein a given core group of a respective processing region is coupled to a set of memory blocks that are proximate to the given core group, while not coupled to memory blocks in the adjacent memory regions that are distal from the given core group.

17 . The MPU of claim 1 , wherein the plurality of core groups of respective ones of the plurality of processing regions are coupled between adjacent ones of the plurality of memory regions of the first memory.

18 . The MPU of claim 1 , wherein each of the plurality of memory blocks are arranged in a plurality of columns and rows.

19 . The MPU of claim 18 , wherein each of the plurality of memory regions are a first plurality of blocks wide and a second plurality of blocks long.

20 . The MPU of claim 1 , wherein compute functions performed by the one or more compute cores and dataflow between the compute cores and the plurality of memory blocks are mapped based on adjacency so that dataflow of shared data is synchronized.