IP Library › Granted Patent US 12,210,959
Granted Patent B2
US 12,210,959 · App. 17/155,896 · Granted Jan 28, 2025

Branching operation for neural processor circuit

Inventors: Kenneth W. Waters (San Jose, CA); Christopher L. Mills (Saratoga, CA)
Assignee: APPLE INC.
G06N3/063G06F9/3005G06F9/4843G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,210,959
App. No.
17/155,896
Granted
Jan 28, 2025
Kind
B2
Abstract

A neural processor includes neural engines for performing convolution operations on input data corresponding to one or more tasks to generate output data. The neural processor circuit also includes a data processor circuit that is coupled to one or more neural engine. The data processor circuit receives the output data from the neural engine and generates a branching command from the output data. The neural processor circuit further includes a task manager that is coupled to the data processor circuit. The task manager receives the branching command from the data processor circuit. The task manager enqueues one of two or more segment branches according to the received branching command. The two or more segment branches are subsequent to a pre-branch task segment that includes the pre-branch task. The task manager transmits a task from the selected one of the segment branches to data processor circuit to perform the task.

Claims (53)

1. A neural processor circuit, comprising:

one or more neural engine circuits configured to perform convolution operations on input data corresponding to a first task to generate output data;

a data processor circuit coupled to the one or more neural engine circuits, the data processor circuit configured to:

receive the output data from the one or more neural engine circuits; and

generate a branching command from the output data; and

a task manager circuit coupled to the data processor circuit, the task manager circuit configured to:

receive the branching command from the data processor circuit;

enqueue one of two or more segment branches according to the received branching command, the two or more segment branches subsequent to a task segment comprising the first task, wherein the branching command and the one of the two or more segment branches are executable by the neural processor circuit; and

transmit a second task from a selected one of the two or more segment branches to the data processor circuit to perform the second task.

2. The neural processor circuit of claim 1 , wherein the task segment is assigned to a first neural network and the selected one of the two or more segment branches is assigned to a second neural network different from the first neural network.

3. The neural processor circuit of claim 2 , wherein the one or more neural engine circuits are configured to perform the convolution operations on the input data corresponding to the first task via the first neural network, and wherein the task manager circuit is further configured to:

enqueue a context switch task that causes the second neural network to perform convolution operations on input data corresponding to the second task.

4. The neural processor circuit of claim 1 , wherein the branching command is generated based on comparing one or more values in the output data to one or more reference values.

5. The neural processor circuit of claim 1 , wherein the selected one of the two or more segment branches is enqueued to a task queue in which the task segment is enqueued.

6. The neural processor circuit of claim 1 , wherein addresses of the two or more segment branches are stored in the task segment.

7. The neural processor circuit of claim 1 , the first task that determines the branching command is identified by a branching task identifier stored in the task segment.

8. The neural processor circuit of claim 1 , wherein the first task that determines the branching command is a last task in the task segment.

9. The neural processor circuit of claim 1 , wherein the task manager circuit is further configured to:

determine, prior to receiving the branching command, that tasks in the task segment have been transmitted for execution; and

pause the task manager circuit until the branching command is received.

10. The neural processor circuit of claim 1 , wherein the task manager circuit is further configured to:

determine, prior to receiving the branching command, that tasks in the task segment have been transmitted for execution; and

process, prior to receiving the branching command, a separate task segment that is different from the two or more segment branches.

11. A method of performing neural processing operations by a neural processor circuit, comprising:

performing, by one or more neural engine circuits, convolution operations on input data corresponding to a first task to generate output data;

receiving, by a data processor circuit, the output data from the one or more neural engine circuits;

generating a branching command from the output data;

receiving, by a task manager circuit, the branching command from the data processor circuit;

enqueuing one of two or more segment branches according to the branching command, the two or more segment branches subsequent to a task segment comprising the first task, wherein the branching command and the one of the two or more segment branches are executable by the neural processor circuit; and

transmitting a second task from a selected one of the two or more segment branches to the data processor circuit to perform the second task.

12. The method of claim 11 , wherein the pre branch task segment is assigned to a first neural network and the selected one of the two or more segment branches is assigned to a second neural network different from the first neural network.

13. The method of claim 12 , performing the convolution operations on the input data corresponding to the first task comprises performing the convolution operations on the input data corresponding to the first task via the first neural network, and wherein the method further comprises:

enqueuing a context switch task that causes the second neural network to perform convolution operations on input data corresponding to the second task.

14. The method of claim 11 , wherein the branching command is generated based on comparing one or more values in the output data to one or more reference values.

15. The method of claim 11 , wherein the selected one of the two or more segment branches is enqueued to a task queue in which the task segment is enqueued.

16. The method of claim 11 , further comprising:

determining, prior to receiving the branching command, that tasks in the task segment have been transmitted for execution; and

processing, prior to receiving the branching command, a separate task segment that is different from the two or more segment branches.

17. An electronic device, comprising:

a system memory storing one or more machine learning models; and

a neural processor, comprising:

one or more neural engine circuits configured to perform convolution operations on input data corresponding to a first task to generate output data;

a data processor circuit coupled to the one or more neural engine circuits, the data processor circuit configured to:

receive the output data from the one or more neural engine circuits; and

generate a branching command from the output data; and

a task manager circuit coupled to the data processor circuit, the task manager circuit configured to:

receive the branching command from the data processor circuit;

enqueue one of two or more segment branches according to the received branching command, the two or more segment branches subsequent to a task segment comprising the first task, wherein the branching command and the one of the two or more segment branches are executable by the neural processor; and

transmit a second task from a selected one of the two or more segment branches to the data processor circuit to perform the second task.

18. The electronic device of claim 17 , wherein the task segment is assigned to a first neural network and the selected one of the two or more segment branches is assigned to a second neural network different from the first neural network.

19. The electronic device of claim 18 , wherein the one or more neural engine circuits are configured to perform the convolution operations on the input data corresponding to the first task via the first neural network, and wherein the task manager circuit is further configured to:

enqueue a context switch task that causes the second neural network to perform convolution operations on input data corresponding to the second task.

20. The electronic device of claim 17 , wherein the branching command correspond to a prediction result of one of the machine learning models.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2021
From: WATERS, KENNETH W; MILLS, CHRISTOPHER L.
To: APPLE INC.
Reel/Frame 055003/0081 →
Continuity (1)
Related Publication 20220237439A1 · Jul 28, 2022
References Cited (33)
US 2A · Goulding · 1836 [cited by examiner]
US 5608662A · Large · 1997 [cited by examiner]
US 9360276B1 · Meek · 2016 [cited by examiner]
US 10422139B1 · Warmerdam · 2019 [cited by examiner]
US 10661065B1 · Jackson · 2020 [cited by examiner]
US 10719760B2 · Ma · 2020 [cited by examiner]
US 11173338B1 · Johnson · 2021 [cited by examiner]
US 20090254572A1 · Redlich · 2009 [cited by examiner]
US 20100250497A1 · Redlich · 2010 [cited by examiner]
US 20180095653A1 · Hasek · 2018 [cited by examiner]
US 20180104644A1 · Cox, Jr. · 2018 [cited by examiner]
US 20180301222A1 · Dew, Sr. · 2018 [cited by examiner]
US 20190340014A1 · Fishel · 2019 [cited by examiner]
US 20200073702A1 · Han · 2020 [cited by examiner]
US 20200104167A1 · Chen · 2020 [cited by examiner]
US 20210271958A1 · Mills · 2021 [cited by examiner]
US 20210279557A1 · Febbo · 2021 [cited by examiner]
US 20220036158A1 · Kuo · 2022 [cited by examiner]
US 20220131896A1 · Schmugar · 2022 [cited by examiner]
US 20220237438A1 · Mills · 2022 [cited by examiner]
US 20220237439A1 · Waters · 2022 [cited by examiner]
US 20220317978A1 · Barik · 2022 [cited by examiner]
US 20220317979A1 · Araujo Soares · 2022 [cited by examiner]
US 20230141807A1 · Groenewegen · 2023 [cited by examiner]
US 20230306275A1 · Briceno · 2023 [cited by examiner]
Baek et al., “A Multi-Neural Network Acceleration Architecture”, 2020 ACM/IEEE, 47th Annual International Symposium on Computer Architecture(ISCA), May 30, 2020, pp. 940-953. (Year: 2020). [cited by examiner]
Yi, et al., “FPGA Based Accelerator for Neural Networks Computation with Flexible Pipelining,” arXiv:2112.15443, Dec. 28, 2021, pp. 1-6. (Year: 2021). [cited by examiner]
Baek, E. et al., “A Multi-Neural Network Acceleration Architecture,” 2020 ACM/IEEE 47 [cited by applicant]
Gong, L. et al., “MALOC: A Fully Pipelined FPGA Accelerator for Convolutional Neural Networks With All Layers Mapped on Chip,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 37, No. … [cited by applicant]
Kara, K. et al., “PipeArch: Generic and Context-Switch Capable Data Processing on FPGAs,” ACM Transactions on Reconfigurable Technology and Systems, vol. 14, No. 1, Article 3, Nov. 2020, pp. 1-28. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Application No. PCT/US2022/011880, Apr. 25, 2022, 18 pages. [cited by applicant]
Shomron, G. et al., “Non-Blocking Simultaneous Multithreading: Embracing the Resiliency of Deep Neural Networks,” 2020 53 [cited by applicant]
Yi, Q. et al., “FPGA Based Accelerator for Neural Networks Computation with Flexible Pipelining,” arXiv:2112.15443, Dec. 28, 2021, pp. 1-6. [cited by applicant]