IP Library Granted Patent US 12,462,350
Granted Patent B2
US 12,462,350 · App. 18/405,440 · Granted Nov 4, 2025

Circuit for executing stateful neural network

Inventors: Andrew C. Mihal (San Jose, CA); Steven L. Teig (Menlo Park, CA); Eric A. Sather (Palo Alto, CA)
Assignee: Amazon Technologies, Inc.
G06T5/70G06N3/04G06N3/044G06N3/049G06N3/063G06N3/084G06N5/046G06N20/10G06T2207/10016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,350
App. No.
18/405,440
Granted
Nov 4, 2025
Kind
B2
Abstract

Some embodiments provide a neural network inference circuit for executing a neural network that includes multiple nodes that use state data from previous executions of the neural network. The neural network inference circuit includes (i) a set of computation circuits configured to execute the nodes of the neural network and (ii) a set of memories configured to implement a set of one or more registers to store, while executing the neural network for a particular input, state data generated during at least two executions of the network for previous inputs. The state data is for use by the set of computation circuits when executing a set of the nodes of the neural network for the particular input.

Claims (33)

1 . A neural network inference circuit that executes a neural network comprising a plurality of layers, the neural network inference circuit comprising:

a set of computation circuits configured to execute the plurality of layers of the neural network; and

a set of non-transitory machine-readable memories configured to store state data generated by a first set of layers of the neural network while executing the neural network for a first input, the state data for retrieval and use by a second set of layers of the neural network while executing the neural network for each of a plurality of inputs after the first input,

wherein pointer data indicating that the state data was generated for the first input, is modified during subsequent executions of the neural network to indicate a number of executions that have occurred since the state data was generated.

2 . The neural network inference circuit of claim 1 , wherein a plurality of nodes are arranged in the first set of layers, and wherein the plurality of nodes partially overlaps with nodes arranged in the second set of layers.

3 . The neural network inference circuit of claim 1 , wherein a plurality of nodes are arranged in a first set of layers, and wherein the plurality of nodes overlap with nodes arranged in the second set of layers.

4 . The neural network inference circuit of claim 1 , wherein the first set of layers is completely separate from the second set of layers.

5 . The neural network inference circuit of claim 1 , wherein the set of non-transitory machine-readable memories are configured as a set of shift registers.

6 . The neural network inference circuit of claim 5 , wherein state data generated by the first set of layers during execution of the neural network for the first input is stored in a first set of memory locations of the set of shift registers.

7 . The neural network inference circuit of claim 6 , wherein the state data generated by the first set of layers during execution of the neural network for the first input is shifted to different sets of memory locations after execution of the neural network for each of a set of subsequent inputs.

8 . The neural network inference circuit of claim 6 , wherein the state data generated by the first set of layers during execution of the neural network for the first input is shifted to different sets of memory locations prior to execution of the neural network for each of a set of subsequent inputs.

9 . The neural network inference circuit of claim 6 , wherein the state data generated by the first set of layers during execution of the neural network for the first input is shifted to different sets of memory locations during execution of the neural network for each of a set of subsequent inputs.

10 . The neural network inference circuit of claim 6 , wherein:

the state data generated by the first set of layers during execution of the neural network for the first input is not moved from the first set of memory locations.

11 . The neural network inference circuit of claim 1 , wherein the set of non-transitory machine-readable memories are further configured to store (i) weight data for executing the neural network and (ii) intermediate activation values while executing the neural network.

12 . The neural network inference circuit of claim 1 , wherein:

the first input is a first frame of video from a video stream; and

the plurality of inputs after the first input are subsequent frames of video from the video stream.

13 . The neural network inference circuit of claim 1 further comprising a plurality of cores, wherein:

each core comprises (i) a subset of the set of computation circuits and (ii) a subset of the set of non-transitory machine-readable memories; and

each respective layer of the neural network is executed by the subset of the set of computation circuits belonging to a respective subset of the plurality of cores of neural network inference circuit.

14 . The neural network inference circuit of claim 13 , wherein the subset of the set of non-transitory machine-readable memories belonging to the respective subset of the plurality of cores that execute the second set of layers store the state data generated by the first set of layers.

15 . The neural network inference circuit of claim 14 , wherein the subset of the set of computation circuits belonging to the respective subset of the plurality of cores that execute the first set of layers generate the state data.

16 . The neural network inference circuit of claim 15 , wherein the respective subset of the plurality of cores that execute the first set of layers and the respective subset of the plurality of cores that execute the second set of layers are different groups of cores.

17 . The neural network inference circuit of claim 1 , wherein a particular layer of the first set of layers generates state data based on input from a plurality of previous layers of the neural network.

18 . The neural network inference circuit of claim 17 , wherein the plurality of previous layers that provide input to the particular layer comprises at least one layer of the second set of layers.

19 . The neural network inference circuit of claim 1 , wherein the set of non-transitory machine-readable memories store state data from a particular number of previous executions, and wherein the set of computation circuits use a subset of the state data for a particular execution of the neural network.

20 . A method for using a neural network inference circuit to execute a neural network comprising a plurality of layers, the method comprising:

executing, using a set of computation circuits, the plurality of layers of the neural network, wherein executing the plurality of layers of the neural network comprises:

generating, for a first input by a first set of layers of the neural network, state data;

storing the state data in a set of non-transitory machine-readable memories for retrieval and use by a second set of layers of the neural network while executing the neural network for each of a plurality of inputs after the first input,

generating pointer data that indicates that the state data was generated for the first input, and

modifying, during subsequent executions of the neural network, the pointer data to indicate a number of executions that have occurred since the state data was generated.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 16, 2025
From: MIHAL, ANDREW C.; TEIG, STEVEN L.; SATHER, ERIC A.
To: PERCEIVE CORPORATION
Reel/Frame 071730/0325 →
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
Continuity (4)
Continuation 16584891 · Sep 26, 2019
Provisional Application 62901740 · Sep 17, 2019
Provisional Application 62888413 · Aug 16, 2019
Related Publication 20240153044A1 · May 9, 2024
References Cited (43)
US 5956703A · Turner et al. · 1999 [cited by applicant]
US 10572979B2 · Vogels et al. · 2020 [cited by applicant]
US 10740434B1 · Duong · 2020 [cited by examiner]
US 11151695B1 · Mihal · 2021 [cited by examiner]
US 11170289B1 · Duong · 2021 [cited by examiner]
US 11615322B1 · Thomas · 2023 [cited by examiner]
US 11620495B1 · Mihal et al. · 2023 [cited by applicant]
US 11868871B1 · Mihal et al. · 2024 [cited by applicant]
US 20040078403A1 · Scheuermann et al. · 2004 [cited by applicant]
US 20140132786A1 · Saitwal et al. · 2014 [cited by applicant]
US 20140180987A1 · Arthur et al. · 2014 [cited by applicant]
US 20180032846A1 · Yang et al. · 2018 [cited by applicant]
US 20180046458A1 · Kuramoto · 2018 [cited by examiner]
US 20180181406A1 · Kuramoto · 2018 [cited by examiner]
US 20180293713A1 · Vogels et al. · 2018 [cited by applicant]
US 20190114499A1 · Delaye et al. · 2019 [cited by applicant]
US 20190114544A1 · Sundaram et al. · 2019 [cited by applicant]
US 20190114804A1 · Sundaresan et al. · 2019 [cited by applicant]
US 20190303750A1 · Kumar · 2019 [cited by examiner]
US 20190327486A1 · Liao · 2019 [cited by examiner]
US 20190354868A1 · Wierstra et al. · 2019 [cited by applicant]
US 20200034971A1 · Xu · 2020 [cited by examiner]
US 20200082540A1 · Chowdhury et al. · 2020 [cited by applicant]
US 20200160150A1 · Hashemi · 2020 [cited by examiner]
US 20220383068A1 · Matveev · 2022 [cited by examiner]
CN 110390384A · 2019 [cited by examiner]
WO WO2017151757A1 · 2017 [cited by examiner]
Joy Bose, “Engineering A Sequence Machine Through Spiking Neurons Employing Rank-Order Codes”, 2007, University of Manchester, pp. 1-216 (Year: 2007). [cited by examiner]
Bose, Joy, “Engineering a Sequence Machine Using Spike Neurons Employing Rank Order Codes,” Doctoral Thesis, Jan. 2007, 65 pages, University of Manchester. [cited by applicant]
Chen, Chen, et al., “Learning to See in the Dark,” May 4, 2018, 10 pages, arXiv:1805.01934v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Chen, Tianqi, et al., “Training Deep Nets with Sublinear Memory Cost,” Apr. 22, 2016, 12 pages, arXiv:1604.06174v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
Guo, Jia, et al., “Stacked Dense U-Nets with Dual Transformers for Robust Face Alignment,” Dec. 5, 2018, 13 pages, arXiv:1812.01936, arXiv.org. [cited by applicant]
He, Kaiming, et al., “Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification,” Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Dec. 7-13, 2015, 9 pag… [cited by applicant]
Huang, Gao, et al., “Multi-Scale Dense Networks for Resource Efficient Image Classification,” Proceedings of the 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 14 pages, ICLR,… [cited by applicant]
Jain, Anil K., et al., “Artificial Neural Networks: A Tutorial,” Computer, Mar. 1996, 14 pages, vol. 29, Issue 3, IEEE. [cited by applicant]
Lin, Huangxing, et al., “A∧2Net: Adjacent Aggregation Networks for Image Raindrop Removal,” Nov. 24, 2018, 9 pages, arXiv:1811.09780v1, arXiv.org. [cited by applicant]
Nair, Vinod, et al., “Rectified Linear Units Improve Restricted Boltzmann Machines,” Proceedings of the 27th International Conference on Machine Learning, Jun. 21-24, 2010, 8 pages, Omnipress, Haifa, Israel. [cited by applicant]
Srivastava, Rupesh Kumar, et al., “Highway Networks,” Nov. 3, 2015, 6 pages, arXiv:1505.00387v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Wang, Jung-Hua, et al., “Learning Temporal Sequences Using Dual-Weight Neurons,” Journal of the Chinese Institute of Engineers, Month Unknown 2001, 16 pages, vol. 24, No. 3, Taylor & Francis. [cited by applicant]
Yu, Fisher, et al., “Deep Layer Aggregation,” Jan. 4, 2019, 10 pages, arXiv:1707.06484v3, arXiv.org. [cited by applicant]
Zilly, Julian Georg, et al., “Recurrent Highway Networks,” Jul. 4, 2017, 12 pages, arXiv:1607.03474v5, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]