IP Library Granted Patent US 12,625,745
Granted Patent B2
US 12,625,745 · App. 17/771,606 · Granted May 12, 2026

Optimized placement for efficiency for accelerated deep learning

Inventors: Vladimir Kibardin (Palo Alto, CA); Michael Edwin James (San Carlos, CA); Michael Morrison (Sunnyvale, CA); Sean Lie (Los Altos, CA); Gary R. Lauterbach (Los Altos, CA); Stanislav Funiak (St Lucia, AU)
Assignee: Cerebras Systems Inc.
G06F9/54G06F9/5027G06F18/214G06N3/04G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,625,745
App. No.
17/771,606
Granted
May 12, 2026
Kind
B2
Abstract

Techniques in optimized placement for efficiency for accelerated deep learning provide improvements in one or more of accuracy, performance, and energy efficiency. An array of processing elements comprising a portion of a neural network accelerator performs flow-based computations on wavelets of data. Each processing element comprises a compute element to execute programmed instructions using the data and a router to route the wavelets. The routing is in accordance with virtual channel specifiers of the wavelets and controlled by routing configuration information of the router. A software stack determines optimized placement based on a description of a neural network. The determined placement is used to configure the routers including usage of the respective colors. The determined placement is used to configure the compute elements including the respective programmed instructions each is configured to execute.

Claims (29)

1 . A method comprising:

extracting a model from a neural network description;

computing delays based on convergent nodes of the extracted model;

determining routing to implement data communication based on arcs of the extracted model;

determining accelerator configuration information usable to configure a deep learning accelerator to provide a trained model, wherein the accelerator configuration information indicates delay buffer placement based on the delays and on the routing, and wherein the deep learning accelerator comprises a fabric and a plurality of processing elements enabled to communicate packets with each other via the fabric in accordance with a plurality of communication pathways identifiable by respective virtual channel identifiers; and

configuring the deep learning accelerator based on the accelerator configuration information.

2 . The method of claim 1 , wherein the determining the routing ignores interactions between routes.

3 . The method of claim 2 , further comprising scanning results based on the determining the routing to produce hotspot information to repeat the determining the routing in accordance therewith.

4 . The method of claim 1 , wherein the determining the routing ignores coloring and bandwidth interactions with other routes.

5 . A non-transitory computer-readable medium comprising one or more instructions encoded thereon that, when executed by one or more processors, cause the one or more processors to perform actions comprising:

extracting a model from a neural network description;

computing delays based on convergent nodes of the extracted model;

determining routing to implement data communication based on arcs of the extracted model;

determining accelerator configuration information usable to configure a deep learning accelerator to provide a trained model, wherein the accelerator configuration information indicates delay buffer placement based on the delays and on the routing, and wherein the deep learning accelerator comprises a fabric and a plurality of processing elements enabled to communicate packets with each other via the fabric in accordance with a plurality of communication pathways identifiable by respective virtual channel identifiers; and

configuring the deep learning accelerator based on the accelerator configuration information.

6 . The non-transitory computer-readable medium of claim 5 , wherein the determining the routing ignores interactions between routes.

7 . The non-transitory computer-readable medium of claim 6 , further comprising scanning results based on the determining the routing to produce hotspot information to repeat the determining the routing in accordance therewith.

8 . The non-transitory computer-readable medium of claim 5 , wherein the determining the routing ignores coloring and bandwidth interactions with other routes.

9 . A deep learning accelerator comprising:

a fabric; and

circuitry configured as a plurality of processing elements that is enabled to communicate packets with each other via the fabric in accordance with a plurality of communication pathways identifiable by respective virtual channel identifiers, wherein the circuitry is configured to:

extract a model from a neural network description;

compute delays based on convergent nodes of the extracted model;

determine routing to implement data communication based on arcs of the extracted model;

determine accelerator configuration information usable to configure the deep learning accelerator to provide a trained model, wherein the accelerator configuration information indicates delay buffer placement based on the delays and on the routing; and

configure the deep learning accelerator based on the accelerator configuration information.

10 . The system of claim 9 , wherein the circuitry, when determining the routing, is configured to ignore interactions between routes.

11 . The system of claim 10 , wherein the circuitry is further configured to scan results based on the determining the routing to produce hotspot information to repeat the determining the routing in accordance therewith.

12 . The system of claim 9 , wherein the circuitry, when determining the routing, is configured to ignore coloring and bandwidth interactions with other routes.

Assignments (2)
SECURITY INTEREST Recorded Jun 18, 2026
From: CEREBRAS SYSTEMS INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS THE COLLATERAL AGENT
Reel/Frame 075845/0844 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2022
From: KIBARDIN, VLADIMIR; JAMES, MICHAEL EDWIN; MORRISON, MICHAEL; LIE, SEAN; LAUTERBACH, GARY R.; FUNIAK, STANISLAV
To: CEREBRAS SYSTEMS INC.
Reel/Frame 061044/0881 →
Continuity (3)
Provisional Application 62929055 · Oct 31, 2019
Provisional Application 62928198 · Oct 30, 2019
Related Publication 20230125522A1 · Apr 27, 2023
References Cited (34)
US 7986629B1 · Ferguson et al. · 2011 [cited by applicant]
US 11256985B2 · Regev · 2022 [cited by applicant]
US 11580376B2 · Hwang et al. · 2023 [cited by applicant]
US 11934945B2 · Lie et al. · 2024 [cited by applicant]
US 20050131660A1 · Yadegar et al. · 2005 [cited by applicant]
US 20050220094A1 · Parker et al. · 2005 [cited by applicant]
US 20150302295A1 · Rivera et al. · 2015 [cited by applicant]
US 20150304797A1 · Rhoads et al. · 2015 [cited by applicant]
US 20150324684A1 · Alvarez-Icaza Rivera et al. · 2015 [cited by applicant]
US 20170295061A1 · Wittenschlaeger · 2017 [cited by applicant]
US 20180189642A1 · Boesch et al. · 2018 [cited by applicant]
US 20180314941A1 · Lie · 2018 [cited by examiner]
US 20190102338A1 · Tang et al. · 2019 [cited by applicant]
US 20190258919A1 · Lie et al. · 2019 [cited by applicant]
US 20230071424A1 · Kibardin et al. · 2023 [cited by applicant]
WO 2012044432A1 · 2012 [cited by applicant]
WO 2021074795A1 · 2021 [cited by applicant]
WO 2021074865A1 · 2021 [cited by applicant]
WO 2021074867A1 · 2021 [cited by applicant]
WO 2021084485A1 · 2021 [cited by applicant]
WO 2021084506A1 · 2021 [cited by applicant]
WO 2021084505A1 · 2021 [cited by applicant]
International Search Report in PCT/IB2020/060231 (the international stage of the instant case), Feb. 22, 2021, 4 pages. [cited by applicant]
Written Opinion of the International Searching Authority in PCT/IB2020/060231 (the international stage of the instant case), Feb. 22, 2021, 5 pages. [cited by applicant]
International Preliminary Report On Patentability (Chapter II) in PCT/IB2020/060231 (the International stage of the instant case), Jan. 27, 2022, 5 pages. [cited by applicant]
Mohammed Amine Meghabber et al., ‘A Flexible Network on-Chip Router for Data-Flow Monitoring’, The 5th International Conference on Electrical Engineering—Boumerdes (ICEE-B), Oct. 31, 2017, 6 pages. [cited by applicant]
Cotter F. et al., Deep Learning In The Wavelet Domain, arXiv: 1811.06115v1 [cs.CV]. pp. 1-5. Nov. 14, 2018. [cited by applicant]
International Preliminary Report On Patentability (Ch II) in PCT/IB2020/060188, Jan. 26, 2022, 5 pages. [cited by applicant]
International Search Report in the related case PCT/IB2020/060188, Feb. 22, 2021, 4 pages. [cited by applicant]
Narayanamurthy N et al: “Evolving Bio Plausible Design with Heterogeneous Noc”, The 15th International Conference on Advanced Communications Technology-ICACT2013, Jan. 27, 2013, (pp. 451-456), 6 pages. [cited by applicant]
Shawahna A. et al., ‘FPGA-Based Accelerators of Deep Learning Networks for Learning and Classification: A Review’, IEEE Access, vol. 7, Dec. 28, 2018, pp. 7825-7828. [cited by applicant]
Written Opinion of the International Searching Authority in PCT/IB2020/060188 (the international stage of the instant case), Feb. 22, 2021, 4 pages. [cited by applicant]
Xinyu You et al., “Toward Packet Routing with Fully-distributed Multi-agent Deep Reinforcement Learning”, arXiv:1905.03494v1, May 9, 2019 [retrieved on Jan. 27, 2021]. Retrieved from <https:// arxiv.org/pdf/1905.03494v1… [cited by applicant]
Yiping Dong et al: “Network on Chip Architecture for BP neural network”, Communications, Circuits and Systems, 2008. ICCCAS 2008. International Conference On, IEEE, Piscataway, NJ, USA, May 25, 2008 (May 25, 2008), pp. … [cited by applicant]