IP Library › Granted Patent US 12,443,831
Granted Patent B1
US 12,443,831 · App. 16/751,038 · Granted Oct 14, 2025

Neural network execution streams

Inventors: Bin Fan (San Jose, CA); Yuan Lin (Cupertino, CA)
Assignee: NVIDIA Corporation
G06N3/063G06F9/3869G06F9/54
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,443,831
App. No.
16/751,038
Granted
Oct 14, 2025
Kind
B1
Abstract

Apparatuses, systems, and techniques to perform a neural network. In at least one embodiment, an application programming interface schedules two or more graph nodes to be performed by two or more parallel processing pipelines based, at least in part, on an order of layers in the neural network.

Claims (36)

1. A non-transitory machine-readable medium having stored thereon an application programming interface (API), which if performed by one or more processors, cause the one or more processors to at least:

cause two or more neural network graph nodes to be synchronized and performed in parallel by two or more parallel processing pipelines of the one or more processors based, at least in part, on whether two or more neural network layers corresponding to the two or more neural network graph nodes are dependent on each other, wherein each node of the two or more neural network graph nodes comprises one or more operations to be performed by one or more execution streams associated with the two or more parallel processing pipelines.

2. The non-transitory machine-readable medium of claim 1 , wherein the one or operations correspond to one or more basic blocks.

3. The non-transitory machine-readable medium of claim 1 , wherein the one or more operations are to be performed in series.

4. The non-transitory machine-readable medium of claim 1 , wherein the two or more neural network graph nodes are performed based at least in part on dividing a graph into portions and executing each portion on a respective execution stream of the one or more execution streams.

5. The non-transitory machine-readable medium of claim 4 , wherein a graph node is assigned to a portion of the graph based at least in part on a distance of the graph node to an exit of the graph.

6. The non-transitory machine-readable medium of claim 5 , wherein the distance is based at least in part on an estimated cost of performing an operation associated with the graph node.

7. The non-transitory machine-readable medium of claim 5 , wherein the graph node is determined to be a candidate for assignment to the portion when an in-degree of the graph node is zero.

8. The non-transitory machine-readable medium of claim 1 , wherein execution of a first basic block by a first of the two or more neural network graph nodes to be performed in parallel is synchronized with execution of a second basic block by a second of the two or more neural network graph nodes to be performed in parallel.

9. One or more processors, comprising:

one or more arithmetic logic units (ALUs) to be configured to cause two or more neural network graph nodes to be synchronized and performed in parallel by two or more parallel processing pipelines of the one or more processors based, at least in part, on whether two or more neural network layers corresponding to the two or more neural network graph nodes are dependent on each other, wherein each node of the two or more neural network graph nodes comprises one or more operations to be performed by one or more execution streams associated with the two or more parallel processing pipelines.

10. The one or more processors of claim 9 , wherein a graph is divided into portions, and wherein each portion is executed by a respective execution stream of the one or more execution streams associated with each one of the two or more neural network graph nodes to be performed in parallel.

11. The one or more processors of claim 10 , wherein the one or more execution streams are associated with a plurality of operations to execute in series.

12. The one or more processors of claim 10 , wherein a graph node is assigned to a portion of the graph based at least in part on a distance of the graph node to an exit of the graph.

13. The one or more processors of claim 10 , wherein the graph is divided into the portions based at least in part on assignment of graph nodes associated with a basic block to a first execution stream of the one or more execution streams.

14. The one or more processors of claim 9 , wherein a graph node corresponds to an operation of a neural network comprising the two or more neural network graph nodes, and wherein an edge of a graph corresponds to data flow between operations.

15. The one or more processors of claim 9 , wherein each of the two or more neural network graph nodes to be performed in parallel is to execute operations associated with the one or more execution streams.

16. The one or more processors of claim 15 , wherein a node of a graph is added to a list of candidates for assigning to the one or more execution streams when an in-degree of the node is zero.

17. The one or more processors of claim 15 , wherein operations associated with a basic block are executed by the one or more execution streams.

18. A system, comprising:

one or more processors to be configured to at least cause two or more neural network graph nodes to be synchronized and performed in parallel by two or more parallel processing pipelines of the one or more processors based, at least in part, on whether two or more neural network layers corresponding to the two or more neural network graph nodes are dependent on each other, wherein each node of the two or more neural network graph nodes comprises one or more operations to be performed by one or more execution streams associated with the two or more parallel processing pipelines.

19. The system of claim 18 , wherein the one or operations correspond to one or more basic blocks.

20. The system of claim 18 , wherein the one or more operations are to be performed in series.

21. The system of claim 18 , wherein the two or more neural network graph nodes are to be performed based at least in part on division of a graph into portions and execution of the portions by the one or more execution streams.

22. The system of claim 21 , wherein division of the graph into the portions comprises assignment of a graph node to a portion of the graph based at least in part on a distance of the graph node to an exit of the graph.

23. The system of claim 22 , wherein the distance is based at least in part on an estimated cost of performing an operation associated with the graph node.

24. The system of claim 21 , wherein a graph node is determined to be a candidate for assignment to the portion of the graph when an in-degree of a graph node is zero.

25. The system of claim 18 , wherein execution of a first basic block by a first of the two or more neural network graph nodes to be performed in parallel is synchronized with execution of a second basic block by a second of the two or more neural network graph nodes to be performed in parallel.

26. A system, comprising:

one or more processors to be configured to perform a neural network by at least receiving a graph definition and causing two or more neural network graph nodes of a graph to be synchronized and performed in parallel by two or more parallel processing pipelines of the one or more processors based, at least in part, on whether two or more neural network layers corresponding to the two or more neural network graph nodes are dependent on each other, wherein each node of the two or more neural network graph nodes comprises one or more operations to be performed by one or more execution streams associated with the two or more parallel processing pipelines.

27. The system of claim 26 , wherein the one or operations correspond to one or more basic blocks.

28. The system of claim 26 , wherein the neural network is performed based at least in part on dividing the graph into portions and executing each portion on a respective execution stream of the one or more execution streams.

29. The system of claim 28 , wherein a graph node is assigned to a portion of the graph based at least in part on a distance of the graph node to an exit of the graph.

30. The system of claim 29 , wherein the distance is based at least in part on an estimated cost of performing an operation associated with the graph node.

31. The system of claim 28 , wherein a graph node is determined to be a candidate for assignment to a portion of the graph when an in-degree of the graph node is zero.

32. The system of claim 28 , wherein execution of a first basic block by a first of the two or more neural network graph nodes to be performed in parallel is synchronized with execution of a second basic block by a second of the two or more neural network graph nodes to be performed in parallel.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2020
From: FAN, BIN; LIN, YUAN
To: NVIDIA CORPORATION
Reel/Frame 051630/0747 →
References Cited (36)
US 10789402B1 · Vemuri · 2020 [cited by examiner]
US 11531902B2 · Horesh · 2022 [cited by examiner]
US 11748622B1 · Borkovic · 2023 [cited by examiner]
US 20180136912A1 · Venkataramani · 2018 [cited by examiner]
US 20180204314A1 · Kaplanyan · 2018 [cited by examiner]
US 20180314931A1 · Sarel · 2018 [cited by examiner]
US 20190073590A1 · Wu · 2019 [cited by examiner]
US 20190114534A1 · Teng · 2019 [cited by examiner]
US 20190205737A1 · Bleiweiss · 2019 [cited by examiner]
US 20190205747A1 · Srivastava · 2019 [cited by examiner]
US 20190213775A1 · Dimitrov · 2019 [cited by examiner]
US 20190302883A1 · Greer · 2019 [cited by examiner]
US 20190362227A1 · Seshadri · 2019 [cited by examiner]
US 20190392296A1 · Brady · 2019 [cited by examiner]
US 20200167654A1 · Guo · 2020 [cited by examiner]
US 20200242734A1 · Wang · 2020 [cited by examiner]
US 20210081691A1 · Chen · 2021 [cited by examiner]
US 20220245454A1 · Sridharan · 2022 [cited by examiner]
Gaunt, Alexander L., et al. “AMPNet: Asynchronous model-parallel training for dynamic neural networks.” arXiv preprint arXiv:1705.09786 (2017): 1-18 (Year: 2017). [cited by examiner]
Wang, Siqi, et al. “High-throughput CNN inference on embedded ARM Big. LITTLE multicore processors.” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 39.10 (2019): 2254-2267. (Year: 2019). [cited by examiner]
Zhan, Jun, and Jinghui Zhang. “Pipe-torch: Pipeline-based distributed deep learning in a gpu cluster with heterogeneous networking.” 2019 Seventh International Conference on Advanced Cloud and Big Data (CBD). IEEE, 2019… [cited by examiner]
Cavicchioli, Roberto, et al. “Novel methodologies for predictable CPU-to-GPU command offloading.” 31st Euromicro Conference on Real-Time Systems (ECRTS 2019). Schloss Dagstuhl-Leibniz-Zentrum fuer Informatik, 2019: 22:1… [cited by examiner]
Xiang, Yecheng, and Hyoseung Kim. “Pipelined data-parallel CPU/GPU scheduling for multi-DNN real-time inference.” 2019 IEEE Real-Time Systems Symposium (RTSS). IEEE, 2019: 392-405 (Year: 2019). [cited by examiner]
Jiang, Wenbin, et al. “Exploiting potential of deep neural networks by layer-wise fine-grained parallelism.” Future Generation Computer Systems 102 (Jan. 1, 2020): 210-221. (Year: 2020). [cited by examiner]
Vokorokos, Liberios, and Norbert Ádám. “An introduction to the Neural DF architecture.” 2011 IEEE 9th International Symposium on Applied Machine Intelligence and Informatics (SAMI). IEEE, 2011. (Year: 2011). [cited by examiner]
Mohammadi, Mahnaz, et al. “A flexible scalable hardware architecture for radial basis function neural networks.” 2015 28th International Conference on VLSI Design. IEEE, 2015. (Year: 2015). [cited by examiner]
Elsken, Thomas, Jan Hendrik Metzen, and Frank Hutter. “Neural architecture search: A survey.” The Journal of Machine Learning Research 20.1 (2019): 1997-2017. (Year: 2017). [cited by examiner]
Lin, Yu-Shiang, et al. “qCUDA: GPGPU virtualization for high bandwidth efficiency.” 2019 IEEE International Conference on Cloud Computing Technology and Science (CloudCom). IEEE, 2019. (Year: 2019). [cited by examiner]
Poli, Gustavo, et al. “Processing neocognitron of face recognition on high performance environment based on GPU with CUDA architecture.” 2008 20th International Symposium on Computer Architecture and High Performance Co… [cited by examiner]
Tagliavini, Giuseppe, et al. “Enabling OpenVX support in mW-scale parallel accelerators.” Proceedings of the International Conference on Compilers, Architectures and Synthesis for Embedded Systems. 2016. (Year: 2016). [cited by examiner]
Huang, Jiayi, et al. “Active-routing: Compute on the way for near-data processing.” 2019 IEEE International symposium on high performance computer architecture (HPCA). IEEE, 2019. (Year: 2019). [cited by examiner]
Ma, Lingxiao, et al. “{NeuGraph}: Parallel deep neural network computation on large graphs.” 2019 USENIX Annual Technical Conference (USENIX ATC 19). 2019. (Year: 2019). [cited by examiner]
Hajewski, Jeff, and Suely Oliveira. “A scalable system for neural architecture search.” 2020 10th Annual Computing and Communication Workshop and Conference (CCWC). IEEE, Jan. 6, 2020. (Year: 2020). [cited by examiner]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, Sep. 30, 2… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Cited By (2)
US 12,626,093 US 12,684,914