IP Library › Granted Patent US 12,406,174
Granted Patent B2
US 12,406,174 · App. 16/161,867 · Granted Sep 2, 2025

Multi-agent instruction execution engine for neural inference processing

Inventors: Andrew S. Cassidy (San Jose, CA); Simon J. Hollis (San Jose, CA); Hartmut Penner (San Jose, CA); Jun Sawada (Austin, TX); Pallab Datta (San Jose, CA); John V. Arthur (Mountain View, CA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/063G06F9/30036G06F9/3887G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,174
App. No.
16/161,867
Granted
Sep 2, 2025
Kind
B2
Abstract

Multi-agent instruction execution engines for neural inference processing are provided. In various embodiments, a neural core is provided. The neural core includes an instruction memory. The instruction memory comprises a plurality of instruction streams, each instruction stream associated with one of a plurality of agents. The instruction memory further comprises a plurality of shared functional units. The neural core is adapted to concurrently execute the plurality of instruction streams on the plurality of associated agents. The execution includes maintaining a separate program counter for each of the plurality of agents, determining a plurality of operations from the instructions of each instruction stream, and directing the operations to the shared functional units. The instructions of each instruction stream are statically scheduled prior to runtime to ensure their execution is conflict free.

Claims (38)

1. A neural core comprising:

an instruction memory, the instruction memory comprising a plurality of instruction streams, each instruction stream associated with one of a plurality of agents, each agent of the plurality of agents comprising a control engine and a program counter such that a first agent comprises a first control engine and a first program counter, each control engine of the plurality of control engines performing independent control operations, each program counter comprising a register; and

a plurality of shared functional units,

wherein the neural core is adapted to concurrently execute the plurality of instruction streams on the plurality of associated agents, wherein the execution comprises:

modifying, by the first control engine, the first program counter,

determining a plurality of operations from the instructions of each instruction stream, and

directing the operations and at least one no-operation instruction to the shared functional units according to a prior modeling of access to the shared functional units and an offline state of each of the separate program counters, the operations being directed from the instruction memory and the at least one no-operation instruction delaying one or more of the plurality of agents to avoid simultaneous agent access to the shared functional units.

2. The neural core of claim 1 , wherein the shared functional units comprise arithmetic, communication, address, and/or computation units.

3. The neural core of claim 1 , wherein the each of the plurality of operations control one of the shared functional units.

4. The neural core of claim 1 , wherein each of the plurality of instruction streams is statically scheduled.

5. The neural core of claim 4 , wherein the static schedule is conflict free.

6. The neural core of claim 5 , wherein the static schedule requires that no two operations access the same shared functional unit simultaneously.

7. The neural core of claim 1 , wherein the plurality of operations are directed to the shared functional units at runtime.

8. The neural core of claim 7 , wherein the plurality of operations are directed to the shared functional units within a sequence of time windows.

9. The neural core of claim 7 , wherein directing the plurality of operations to the shared functional units comprises merging operations from each of the plurality of instruction streams.

10. The neural core of claim 9 , wherein merging operations comprises detecting conflicts between operations directed to the same shared functional unit.

11. The neural core of claim 1 , wherein determining the plurality of operations comprises decoding instructions of each instruction stream.

12. The neural core of claim 1 , adapted to map the plurality of operations to any of the shared functional units.

13. The neural core of claim 1 , wherein the instruction memory is logically segmented.

14. The neural core of claim 1 , wherein the execution is divided into a plurality of cycles.

15. The neural core of claim 1 , further comprising a plurality of parallel data paths, each comprising a subset of the plurality of shared functional units.

16. The neural core of claim 1 , wherein the plurality of agents execute synchronously.

17. The neural core of claim 16 , wherein synchronous execution is provided via a synchronization signal.

18. The neural core of claim 1 , wherein the independent control operations comprise updating one or more loop counter and/or sequence counter.

19. A method comprising:

reading a plurality of instruction streams from an instruction memory of a neural core, each instruction stream associated with one of a plurality of agents, each agent of the plurality of agents comprising a control engine and a program counter such that a first agent comprises a first control engine and a first program counter, each control engine of the plurality of control engines performing independent control operations, each program counter comprising a register;

concurrently executing the plurality of agents by the neural core;

modifying, by the first control engine, the first program counter;

determining a plurality of operations from the instructions of each instruction stream; and

directing the operations and at least one no-operation instruction to shared functional units of the neural core according to a prior modeling of access to the shared functional units and to an offline state of each of the separate program counters, the operations being directed from the instruction memory and the at least one no-operation instruction delaying one or more of the plurality of agents to avoid simultaneous agent access to the shared functional units.

20. The method of claim 19 , further comprising:

computing by the neural core a portion of a neural network layer.

21. A method comprising:

executing a plurality of instruction streams, each by one of a plurality of agents, each agent of the plurality of agents comprising a control engine and a program counter such that a first agent comprises a first control engine and a first program counter, each control engine of the plurality of control engines performing independent control operations, each program counter comprising a register; and

modifying, by the first control engine, the first program counter,

wherein

a plurality of shared functional units is controlled by the plurality of instruction streams, the plurality of the shared functional units performing an inference operation, and wherein a prior modeling of access to the shared functional units and offline states of the plurality of program counters are used to avoid simultaneous agent access to the shared functional units by delaying one or more of the plurality of agents using at least one no-operation instruction.

22. The method of claim 21 , wherein the inference operations comprise computation, communication, or memory addressing operations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2018
From: CASSIDY, ANDREW S.; HOLLIS, SIMON J.; PENNER, HARTMUT; SAWADA, JUN; DATTA, PALLAB; ARTHUR, JOHN V.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 047198/0693 →
Continuity (1)
Related Publication 20200117465A1 · Apr 16, 2020
References Cited (21)
US 7970970B2 · Subramanian et al. · 2011 [cited by applicant]
US 9063783B2 · Drepper · 2015 [cited by applicant]
US 9577911B1 · Castleman · 2017 [cited by applicant]
US 9843628B2 · Brueckner · 2017 [cited by applicant]
US 9961012B2 · Dutta et al. · 2018 [cited by applicant]
US 20140173122A1 · Sunderrajan · 2014 [cited by applicant]
US 20140229421A1 · Thomas · 2014 [cited by applicant]
US 20170161638A1 · Garagic et al. · 2017 [cited by applicant]
US 20180024870A1 · Grebnov et al. · 2018 [cited by applicant]
US 20180096283A1 · Wang et al. · 2018 [cited by applicant]
US 20180096284A1 · Stets et al. · 2018 [cited by applicant]
US 20180225116A1 · Henry · 2018 [cited by examiner]
US 20180260691A1 · Nagaraja · 2018 [cited by examiner]
US 20180260700A1 · Nagaraja · 2018 [cited by examiner]
WO 2018071392A1 · 2018 [cited by applicant]
Byrd, Gregory T., and Mark A. Holliday. “Multithreaded processor architectures.” IEEE Spectrum 32.8 (1995): 38-46. (Year: 1995). [cited by examiner]
Hirata, Hiroaki, et al. “An elementary processor architecture with simultaneous instruction issuing from multiple threads.” Proceedings of the 19th annual international symposium on Computer architecture. 1992. (Year: 1… [cited by examiner]
Lopes et al., “Negotiation Strategies For Autonomous Computational Agents,” (2004). [cited by applicant]
Neruda et al., “Computational Intelligence Agent-Oriented Modelling,” 4th WSEAS Int. Conf. on Computational Intelligence, Man-Machine Systems and Cybernetics, Miami, Florida, USA, pp. 238-241 (Nov. 17-19, 2005). [cited by applicant]
Smith et al., “Computational Inference of Neural Information Flow Networks,” PLoS Computational Biology, 2(11): e161, pp. 1436-1449 (2006). [cited by applicant]
Neruda et al., “Toward Dynamic Generation of Computational Agents by Means of Logical Descriptions.” International Transactions on Systems Science and Applications: 6 pages (2008). [cited by applicant]