IP Library › Granted Patent US 12,405,866
Granted Patent B2
US 12,405,866 · App. 18/406,346 · Granted Sep 2, 2025

High performance processor for low-way and high-latency memory instances

Inventors: Elad Sity (Kfar Saba, IL); Eliad Hillel (Kfar Saba, IL)
Assignee: NeuroBlade Ltd.
G06F11/1658G06F9/3001G06F9/3885G06F9/3889G06F9/3895G06F11/1016G06F11/102G06F11/16G06F11/1616G06F13/1657G06F15/8038G06N3/04G11C7/1072G11C11/1655G11C11/1657G11C11/1675G11C11/4076G11C11/408G11C11/4093G06F2015/765
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,405,866
App. No.
18/406,346
Granted
Sep 2, 2025
Kind
B2
Abstract

Distributed processors and methods for compiling code for execution by distributed processors are disclosed. In one implementation, a distributed processor may include a substrate; a memory array disposed on the substrate; and a processing array disposed on the substrate. The memory array may include a plurality of discrete memory banks, and the processing array may include a plurality of processor subunits, each one of the processor subunits being associated with a corresponding, dedicated one of the plurality of discrete memory banks. The distributed processor may further include a first plurality of buses, each connecting one of the plurality of processor subunits to its corresponding, dedicated memory bank, and a second plurality of buses, each connecting one of the plurality of processor subunits to another of the plurality of processor subunits.

Claims (54)

1. A method performed for operating a distributed memory device comprising:

determining a number of words that are required simultaneously to perform a task, the task requiring at least one computation, and

providing instructions for writing words that need to be accessed simultaneously in a plurality of memory banks when a number of words that can be accessed simultaneously from one of the plurality of memory banks is lower than the number of words that are required simultaneously;

receiving, by a configuration manager, an indication to perform the task; and

in response to receiving the indication, configuring a memory controller to:

within a first line access cycle:

access at least one first word from a first memory bank from the plurality of memory banks using a first memory line,

send the at least one first word to at least one processing unit connected to the memory controller, and

open a first memory line in a second memory bank to access a second address from the second memory bank from the plurality of memory banks, and

within a second line access cycle:

access at least one second word from the second memory bank using the first memory line,

send the at least one second word to at least one processing unit connected to the memory controller, and

access a third address from the first memory bank using a second memory line in the first memory bank.

2. The method of claim 1 , further comprising:

determining a number of cycles necessary to perform the task; and

writing words that are needed in sequential cycles in a single memory bank of the plurality of memory banks.

3. The method of claim 1 , further configuring a selected processing unit to transfer data to the second memory bank during the first line access cycle.

4. The method of claim 1 , wherein

the memory controller comprises at least two data inputs from the plurality of memory banks and at least two data outputs connected to each one of the at least one processing unit; and

further comprising configuring the memory controller to:

simultaneously receive data from two memory banks via the two data inputs; and

simultaneously transmit data received via the two data inputs to at least one selected processing unit via the two data outputs.

5. The method of claim 1 , further comprising:

when a number a number of words that can be accessed simultaneously from one of the plurality of memory banks is lower than the number of words that are required simultaneously, divide the number of words required simultaneously between multiple memory banks.

6. The method of claim 1 , wherein the configuration manager comprises a local memory that stores a command to be transmitted to at least one of a plurality of processing units.

7. The method of claim 1 , further comprising configuring the memory controller to interrupt the task in response to receiving a request from an external interface.

8. The method of claim 1 , further comprising configuring the configuration manger and the at least one processing unit to hand over access to the memory controller between each other after finalizing a task.

9. A non-transitory computer-readable medium that stores instructions that, when executed by at least one processor, cause the at least one processor to:

determine a number of words that are required simultaneously to perform a task, the task requiring at least one computation;

write words that need to be accessed simultaneously in a plurality of memory banks when a number of words that can be accessed simultaneously from one of the plurality of memory banks is lower than the number of words that are required simultaneously;

transmit an indication to perform the task to a configuration manager; and

transmit instructions to configure a memory controller to:

within a first line access cycle:

access at least one first word from a first memory bank from the plurality of memory banks using a first memory line,

send the at least one first word to at least one processing unit connected to the memory controller, and

open a first memory line in a second memory bank to access a second address from the second memory bank from the plurality of memory banks, and

within a second line access cycle:

access at least one second word from the second memory bank using the first memory line,

send the at least one second word to at least one processing unit connected to the memory controller, and

access a third address from the first memory bank using a second memory line in the first memory bank.

10. The non-transitory computer-readable medium of claim 9 , wherein a selected processing unit is configured to transfer data to the second memory bank during the first line access cycle.

11. The non-transitory computer-readable medium of claim 9 , wherein

the memory controller comprises at least two data inputs from the plurality of memory banks and at least two data outputs connected to each one of the at least one processing unit;

the memory controller is configured to simultaneously receive data from two memory banks via the two data inputs; and

the memory controller is configured to simultaneously transmit data received via the two data inputs to at least one selected processing unit via the two data outputs.

12. The non-transitory computer-readable medium of claim 9 , wherein the at least one processing unit comprises a plurality of accelerators configured for pre-defined tasks.

13. The non-transitory computer-readable medium of claim 12 , wherein the plurality of accelerators comprise at least one of a vector multiply accumulate unit or a direct memory access.

14. The non-transitory computer-readable medium of claim 9 , wherein the configuration manager comprises at least one of a RISC processor or a micro-controller.

15. The non-transitory computer-readable medium of claim 9 , further comprising an external interface connected to the plurality of memory banks.

16. The non-transitory computer-readable medium of claim 9 , when a number a number of words that can be accessed simultaneously from one of the plurality of memory banks is lower than the number of words that are required simultaneously, divide the number of words required simultaneously between multiple memory banks.

17. The non-transitory computer-readable medium of claim 9 , wherein the words comprise machine instructions.

18. The non-transitory computer-readable medium of claim 9 , wherein the configuration manager comprises a local memory that stores a command to be transmitted to at least one of the at least one processing unit.

19. The non-transitory computer-readable medium of claim 9 , wherein the plurality of memory banks includes at least one of DRAM mats, DRAM, banks, flash mats, or SRAM mats.

20. The non-transitory computer-readable medium of claim 9 , wherein the at least one processing unit comprises at least one arithmetic logic unit, at least one vector handling logic unit, at least one register, and at least one direct memory access.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2024
From: SITY, ELAD; HILLEL, ELIAD
To: NEUROBLADE, LTD.
Reel/Frame 066047/0076 →
Continuity (7)
Continuation 17397061 · Aug 9, 2021
Division 16512622 · Jul 16, 2019
Continuation PCTIB2018000995 · Jul 30, 2018
Provisional Application 62548990 · Aug 23, 2017
Provisional Application 62538722 · Jul 30, 2017
Provisional Application 62538724 · Jul 30, 2017
Related Publication 20240143457A1 · May 2, 2024
References Cited (187)
US 5155729A · Rysko et al. · 1992 [cited by applicant]
US 5179702A · Spix et al. · 1993 [cited by applicant]
US 5214747A · Cok · 1993 [cited by applicant]
US 5239654A · Ing-Simmons et al. · 1993 [cited by applicant]
US 5297260A · Kametani · 1994 [cited by applicant]
US 5345552A · Brown · 1994 [cited by applicant]
US 5502728A · Smith, III · 1996 [cited by applicant]
US 5506992A · Saxenmeyer · 1996 [cited by applicant]
US 5517600A · Shimokawa · 1996 [cited by applicant]
US 5568617A · Kametani · 1996 [cited by applicant]
US 5590345A · Barker et al. · 1996 [cited by applicant]
US 5625796A · Kaczmarczyk et al. · 1997 [cited by applicant]
US 5710935A · Barker et al. · 1998 [cited by applicant]
US 5751987A · Mahant-Shetti et al. · 1998 [cited by applicant]
US 5752036A · Nakamura et al. · 1998 [cited by applicant]
US 5752067A · Wikinson et al. · 1998 [cited by applicant]
US 5822608A · Dieffenderfer et al. · 1998 [cited by applicant]
US 5956703A · Turner et al. · 1999 [cited by applicant]
US 6026464A · Cohen · 2000 [cited by applicant]
US 6041400A · Ozcelik et al. · 2000 [cited by applicant]
US 6064621A · Tanizaki et al. · 2000 [cited by applicant]
US 6067260A · Ooishi et al. · 2000 [cited by applicant]
US 6067262A · Irrinki et al. · 2000 [cited by applicant]
US 6096094A · Kay et al. · 2000 [cited by applicant]
US 6145069A · Dye · 2000 [cited by applicant]
US 6216178B1 · Stracovsky et al. · 2001 [cited by applicant]
US 6363453B1 · Esposito et al. · 2002 [cited by applicant]
US 6389497B1 · Koslawsky et al. · 2002 [cited by applicant]
US 6396760B1 · Behera et al. · 2002 [cited by applicant]
US 6449732B1 · Rasmussen et al. · 2002 [cited by applicant]
US 6453398B1 · McKenzie · 2002 [cited by applicant]
US 6523018B1 · Louis et al. · 2003 [cited by applicant]
US 6553355B1 · Arnoux et al. · 2003 [cited by applicant]
US 6584033B2 · Ayukawa et al. · 2003 [cited by applicant]
US 6601126B1 · Zaidi et al. · 2003 [cited by applicant]
US 6640283B2 · Naffziger et al. · 2003 [cited by applicant]
US 6668308B2 · Barroso et al. · 2003 [cited by applicant]
US 6678801B1 · Greim et al. · 2004 [cited by applicant]
US 6708254B2 · Lee et al. · 2004 [cited by applicant]
US 6717834B2 · Zagorianakos et al. · 2004 [cited by applicant]
US 6785780B1 · Klein et al. · 2004 [cited by applicant]
US 6785841B2 · Akrout et al. · 2004 [cited by applicant]
US 6798420B1 · Xie · 2004 [cited by examiner]
US 6823429B1 · Olnowich · 2004 [cited by applicant]
US 6877046B2 · Singh · 2005 [cited by applicant]
US 6894392B1 · Gudesen et al. · 2005 [cited by applicant]
US 6988154B2 · Latta · 2006 [cited by applicant]
US 6999950B1 · Linneberg et al. · 2006 [cited by applicant]
US 7107285B2 · von Kaenel et al. · 2006 [cited by applicant]
US 7107412B2 · Klein et al. · 2006 [cited by applicant]
US 7159141B2 · Lakhani et al. · 2007 [cited by applicant]
US 7194568B2 · Jeter, Jr. et al. · 2007 [cited by applicant]
US 7290109B2 · Horii et al. · 2007 [cited by applicant]
US 7325232B2 · Liem · 2008 [cited by applicant]
US 7450410B2 · Klein · 2008 [cited by applicant]
US 7549081B2 · Robbins et al. · 2009 [cited by applicant]
US 7557605B2 · D'Souza et al. · 2009 [cited by applicant]
US 7558934B2 · Kondo et al. · 2009 [cited by applicant]
US 7562271B2 · Shaeffer et al. · 2009 [cited by applicant]
US 7610537B2 · Dickinson et al. · 2009 [cited by applicant]
US 7657712B2 · Lentz et al. · 2010 [cited by applicant]
US 7772880B2 · Solomon · 2010 [cited by applicant]
US 7782703B2 · Oh · 2010 [cited by applicant]
US 7836168B1 · Vasko et al. · 2010 [cited by applicant]
US 7881321B2 · Deneroff et al. · 2011 [cited by applicant]
US 7882307B1 · Wentzlaff et al. · 2011 [cited by applicant]
US 7882320B2 · Caulkins · 2011 [cited by applicant]
US 7949820B2 · Caulkins · 2011 [cited by applicant]
US 8028124B2 · Ban et al. · 2011 [cited by applicant]
US 8042082B2 · Solomon · 2011 [cited by applicant]
US 8055601B2 · Pandya · 2011 [cited by applicant]
US 8074031B2 · Bekooij · 2011 [cited by applicant]
US 8171233B2 · Kwon · 2012 [cited by applicant]
US 8369123B2 · Kang et al. · 2013 [cited by applicant]
US 8442927B2 · Chakradhar et al. · 2013 [cited by applicant]
US 8648403B2 · Luk et al. · 2014 [cited by applicant]
US 8677306B1 · Andreev et al. · 2014 [cited by applicant]
US 8738860B1 · Griffin et al. · 2014 [cited by applicant]
US 8984256B2 · Fish · 2015 [cited by applicant]
US 9003109B1 · Lam · 2015 [cited by applicant]
US 9053762B2 · Hirobe · 2015 [cited by applicant]
US 9098209B2 · Gopalakrishnan et al. · 2015 [cited by applicant]
US 9164807B2 · Blanc et al. · 2015 [cited by applicant]
US 9177611B2 · D'Abreu · 2015 [cited by examiner]
US 9239691B2 · Lam · 2016 [cited by applicant]
US 9245222B2 · Modha · 2016 [cited by applicant]
US 9305635B2 · Lin et al. · 2016 [cited by applicant]
US 9342474B2 · Macri et al. · 2016 [cited by applicant]
US 9348385B2 · De Rochemont et al. · 2016 [cited by applicant]
US 9378003B1 · Sundararajan et al. · 2016 [cited by applicant]
US 9432298B1 · Smith · 2016 [cited by applicant]
US 9449257B2 · Shi et al. · 2016 [cited by applicant]
US 9477636B2 · Walker et al. · 2016 [cited by applicant]
US 9570132B2 · Kim et al. · 2017 [cited by applicant]
US 9653151B1 · Ong et al. · 2017 [cited by applicant]
US 9665503B2 · Dalal · 2017 [cited by applicant]
US 9672169B2 · Winderweedle · 2017 [cited by applicant]
US 9760827B1 · Lin et al. · 2017 [cited by applicant]
US 9904874B2 · Shoaib et al. · 2018 [cited by applicant]
US 9978014B2 · Lupon et al. · 2018 [cited by applicant]
US 9984337B2 · Kadav et al. · 2018 [cited by applicant]
US 10032110B2 · Young et al. · 2018 [cited by applicant]
US 10067681B2 · Park et al. · 2018 [cited by applicant]
US 10078620B2 · Farabet et al. · 2018 [cited by applicant]
US 10114795B2 · Cargnini et al. · 2018 [cited by applicant]
US 10140252B2 · Fowers et al. · 2018 [cited by applicant]
US 10157309B2 · Molchanov et al. · 2018 [cited by applicant]
US 10423876B2 · Henry · 2019 [cited by examiner]
US 20050138325A1 · Hofstee et al. · 2005 [cited by applicant]
US 20050223275A1 · Jardine et al. · 2005 [cited by applicant]
US 20070022063A1 · Lightowler · 2007 [cited by applicant]
US 20070214335A1 · Bellows et al. · 2007 [cited by applicant]
US 20070220369A1 · Fayad et al. · 2007 [cited by applicant]
US 20070276977A1 · Coteus et al. · 2007 [cited by applicant]
US 20080046665A1 · Kim · 2008 [cited by applicant]
US 20090083263A1 · Felch et al. · 2009 [cited by applicant]
US 20090158010A1 · Gartner et al. · 2009 [cited by applicant]
US 20090196116A1 · Oh · 2009 [cited by applicant]
US 20090217222A1 · Yasuda et al. · 2009 [cited by applicant]
US 20100131827A1 · Sokolov et al. · 2010 [cited by applicant]
US 20110016278A1 · Ware et al. · 2011 [cited by applicant]
US 20110296078A1 · Khan et al. · 2011 [cited by applicant]
US 20110314473A1 · Yang et al. · 2011 [cited by applicant]
US 20120102576A1 · Chew · 2012 [cited by applicant]
US 20120255018A1 · Sallam · 2012 [cited by applicant]
US 20120291131A1 · Nagpal et al. · 2012 [cited by applicant]
US 20130339589A1 · Qawami · 2013 [cited by examiner]
US 20140040622A1 · Kendall et al. · 2014 [cited by applicant]
US 20140136915A1 · Hyde et al. · 2014 [cited by applicant]
US 20140143470A1 · Dobbs et al. · 2014 [cited by applicant]
US 20140310495A1 · Michelogiannakis et al. · 2014 [cited by applicant]
US 20140328103A1 · Arsovski · 2014 [cited by applicant]
US 20140359199A1 · Huppenthal et al. · 2014 [cited by applicant]
US 20150089162A1 · Ahsan et al. · 2015 [cited by applicant]
US 20150146491A1 · Akerib · 2015 [cited by examiner]
US 20150212861A1 · Canoy et al. · 2015 [cited by applicant]
US 20150293859A1 · Huang et al. · 2015 [cited by applicant]
US 20150309779A1 · Baskaran et al. · 2015 [cited by applicant]
US 20150324690A1 · Chilimbi et al. · 2015 [cited by applicant]
US 20160034809A1 · Trenholm et al. · 2016 [cited by applicant]
US 20160232445A1 · Srinivasan et al. · 2016 [cited by applicant]
US 20160260024A1 · Campos et al. · 2016 [cited by applicant]
US 20160379109A1 · Chung et al. · 2016 [cited by applicant]
US 20170103298A1 · Ling et al. · 2017 [cited by applicant]
US 20170194045A1 · Kang · 2017 [cited by applicant]
US 20170200094A1 · Bruestle et al. · 2017 [cited by applicant]
US 20170286825A1 · Akopyan et al. · 2017 [cited by applicant]
US 20170301389A1 · Kajigaya · 2017 [cited by examiner]
US 20180052766A1 · Mehra et al. · 2018 [cited by applicant]
US 20180144244A1 · Masoud et al. · 2018 [cited by applicant]
US 20180189215A1 · Boesch et al. · 2018 [cited by applicant]
US 20180189645A1 · Chen et al. · 2018 [cited by applicant]
US 20180210830A1 · Malladi et al. · 2018 [cited by applicant]
US 20180285254A1 · Baum et al. · 2018 [cited by applicant]
US 20180285718A1 · Baum et al. · 2018 [cited by applicant]
US 20180336035A1 · Choi et al. · 2018 [cited by applicant]
US 20190102325A1 · Natu et al. · 2019 [cited by applicant]
CA 2149479C · 2001 [cited by applicant]
International Searching Authority, Notification of Transmittal of the International Search Report and the Written Opinion of the International Search Authority, International Application No. PCT/IB2020/00665, dated May … [cited by applicant]
Ahn et al., “A Scalable Processing-in-Memory Accelerator for Parallel Graph Processing,” ISCA '15 (Jun. 13-17, 2015), pp. 105-117. [cited by applicant]
Stripf et al., “Compiling Scilab to high performance embedded multicore systems,” Microprocessors and Microsystems (2013), 37:1033-49. [cited by applicant]
Chandramoorthy, “Design and Exploration of Accelerator-Rich Multi Core Architectures,” The Pennsylvania State University (Dec. 2016), pp. title page, ii-xii, and 1-106. [cited by applicant]
Chaturvedi, “Techniques to Improve the Performance of Cache Memory for Multi-Core Processors,” Birla Institute of Technology and Science (2015), pp. title page, i-xv, and 1-185. [cited by applicant]
Chetlur et al., “cuDNN: Efficient Primitives for Deep Learning,” arXiv:1410.0759v3 (Dec. 18, 2014), pp. 1-9. [cited by applicant]
Chi et al., “Prime: A Novel Processing-in-memory Architecture for Neural Network Computation in ReRAM-based Main Memory,” 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (2016), pp. 27-39. [cited by applicant]
Chung, “CoRAM: An In-Fabric Memory Architecture for FPGA-Based Computing,” Carnegie Mellon University (Aug. 2011), pp. title page, i-xiv, 1-179. [cited by applicant]
Cristian, “Understanding Fault-Tolerant Distributed Systems,” University of California, San Diego (May 25, 1993), pp. 1-45. [cited by applicant]
De Dinechin et al., “A Clustered Manycore Processor Architecture for Embedded and Accelerated Applications,” ResearchGate (Jan. 2013), pp. 1-6. [cited by applicant]
Gupta et al., “The StageNet fabric for constructing resilient multicore systems,” Proceedings of the 41 [cited by applicant]
HP Proliant Servers Troubleshooting Guide, Hewlett Packard (Oct. 30, 2011), Edition 12. [cited by applicant]
Draper et al., “Implementation of a 32-bit RISC Processor for the Data-Intensive Architecture Processing-In-Memory Chip,” Proceedings of the IEEE International Conference on Application-Specific Systems, Architectures, … [cited by applicant]
Malki et al., “A CNN-Specific Integrated Processor,” EURASIP Journal on Advances in Signal Processing (2009), 2009: pp. 1-14. [cited by applicant]
Meridian 1 Engineering Handbook, Update Package, Meridian 1 Communication Systems (Apr. 30, 1990), pp. 1-642. [cited by applicant]
Nanyang Technological University, “Scientists turn memory chips into processors to speed up computing tasks,” Nanyang Technological University (Jan. 3, 2017), pp. 1-5. [cited by applicant]
Shafiee et al., “ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars,” ACM SIGARCH Computer Architecture News 44 (2016), pp. 14-26. [cited by applicant]
Weninger et al., “Introducing CURRENNT: The Munich Open-Source CUDA RecurREnt Neural Network Toolkit,” Journal of Machine Learning Research (Mar. 2015), pp. 547-551. [cited by applicant]
Wu et al., “Demonstration and architectural analysis of complementary metal-oxide semiconductor/multiple-quantum-well smart-pixel array cellular logic processors for single-instruction multiple-data parallel-pipeline pr… [cited by applicant]
Yuan et al., “Complexity Effective Memory Access Scheduling for Many-Core Accelerator Architectures,” MICRO '09 (Dec. 12-16, 2009), pp. 34-44. [cited by applicant]
Udipi et al., “Rethinking DRAM Design and Organization for Energy-Constrained Multi-Cores,” ISCA '10 (Jun. 19-23, 2010), pp. 175-186. [cited by applicant]
Chen et al., “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” 2016 ACM/IEEE 43 [cited by applicant]
Gao et al., “TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory,” ASPLOS '17 (Apr. 8-12, 2017), pp. 1-14. [cited by applicant]
Chen et al., “DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning,” ASPLOS '14 (Mar. 1-5, 2014), pp. 1-1-15. [cited by applicant]
Azarkhish et al., “Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cubes,” arXiv:1701.06420v3 (Sep. 24, 2017), pp. 1-15. [cited by applicant]
Farabet et al., “NeuFlow: A Runtime Reconfigurable Dataflow Processor Vision,” CVPR 2011 Workshops (Jun. 20-25, 2011), pp. 109-116. [cited by applicant]
Wang et al., “DLAU: A Scalable Deep Learning Accelerator Unit on Fpga,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (2016), pp. 1-5. [cited by applicant]
Han et al., “EIE: Efficient Inference Engine on Compressed Deep Neural Network,” arXiv:1602.01528v2 (May 3, 2016), pp. 1-12. [cited by applicant]
European Search Report mailed Apr. 28, 2023, by the European Patent Office in European Application No. EP 23 15 1586, 15 pages. [cited by applicant]