IP Library Granted Patent US 12,353,921
Granted Patent B2
US 12,353,921 · App. 18/535,810 · Granted Jul 8, 2025

Massively parallel in-network compute

Inventors: William Brad Matthews (Los Gatos, CA); Puneet Agarwal (Santa Clara, CA); Bruce Hui Kwan (Santa Clara, CA)
Assignee: Innovium, Inc.
G06F9/5072H04L67/10G06N20/00H04L49/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,353,921
App. No.
18/535,810
Granted
Jul 8, 2025
Kind
B2
Abstract

Efficient scaling of in-network compute operations to large numbers of compute nodes is disclosed. Each compute node is connected to a same plurality of network compute nodes, such as compute-enabled network switches. Compute processes at the compute nodes generate local gradients or other vectors by, for instance, performing a forward pass on a neural network. Each vector comprises values for a same set of vector elements. Each network compute node is assigned to, based on the local vectors, reduce vector data for a different a subset of the vector elements. Each network compute node returns a result chunk for the elements it processed back to each of the compute nodes, whereby each compute node receives the full result vector. This configuration may, in some embodiments, reduce buffering, processing, and/or other resource requirements for the network compute node or network at large.

Claims (29)

1. A method comprising:

generating, by a plurality of compute nodes implemented at least in part with one or more hardware processors, vector chunks having values for a common set of vector elements, the common set of vector elements being common to all vectors generated for a common distributed application, wherein the common set of vector elements includes a plurality of subsets of vector elements, wherein each vector chunk of the vector chunks has a subset of the values for a respective subset in the plurality subsets of vector elements;

sending, to a network switch, the vector chunks over a plurality of switch ports, wherein each vector chunk of the vector chunks is sent to a respective switch port of the plurality of switch ports by a respective compute node in the plurality of compute nodes; and

receiving, from the network switch by each compute node of the plurality of compute nodes, a single result chunk, the single result chunk being formed at the network switch from a subset of vector chunks that were generated by the plurality of compute nodes and sent to the network switch.

2. The method of claim 1 , further comprising: sharing the single result chunk by a plurality of compute processes located at each compute node of the plurality of compute nodes.

3. The method of claim 2 , wherein each compute node of the plurality of compute nodes comprises one or more compute entities; wherein each compute entity of the one or more compute entities hosts one or more respective processes in the plurality of compute processes.

4. The method of claim 1 , wherein the vector chunks are reduced to the single result chunk through one or more of: summation, averaging, multiplying, selecting a minimum value, or selecting a maximum value.

5. The method of claim 1 , wherein the network switch and the plurality of compute nodes collectively implement the common distributed application.

6. The method of claim 5 , wherein the common distributed application includes one or more artificial neural networks.

7. The method of claim 5 , wherein the common distributed application represents one or more of: deep learning applications, machine learning applications or artificial intelligence applications.

8. The method of claim 1 , wherein an error occurs in processing the vector chunks received from the plurality of compute nodes; wherein a message is sent to each compute nodes in the plurality of compute nodes to inform the error.

9. The method of claim 1 , wherein each compute node in the plurality of compute nodes is implemented at least in part with one or more of: central processing units, graphics processing units, tensor processing units, floating point units, hardware accelerators, or other computing processor.

10. The method of claim 1 , wherein a plurality of network switches operates with the plurality of compute nodes to reduce sets of vector chunks sent by the plurality of compute nodes; wherein the plurality of network switches includes the network switch.

11. A system comprising:

a plurality of switch ports;

a plurality of compute nodes implemented by one or more hardware processors;

wherein the system performs:

generating, by the plurality of compute nodes implemented at least in part with one or more hardware processors, vector chunks having values for a common set of vector elements, the common set of vector elements being common to all vectors generated for a common distributed application, wherein the common set of vector elements includes a plurality of subsets of vector elements, wherein each vector chunk of the vector chunks has a subset of the values for a respective subset in the plurality subsets of vector elements;

sending, to a network switch, the vector chunks over the plurality of switch ports, wherein each vector chunk of the vector chunks is sent to a respective switch port of the plurality of switch ports by a respective compute node in the plurality of compute nodes; and

receiving, from the network switch by each compute node of the plurality of compute nodes, a single result chunk, the single result chunk being formed at the network switch from a subset of vector chunks that were generated by the plurality of compute nodes and sent to the network switch.

12. The system of claim 11 , wherein the single result chunk is shared by a plurality of compute processes located at each compute node of the plurality of compute nodes.

13. The system of claim 12 , wherein each compute node of the plurality of compute nodes comprises one or more compute entities; wherein each compute entity of the one or more compute entities hosts one or more respective processes in the plurality of compute processes.

14. The system of claim 11 , wherein the vector chunks are reduced to the single result chunk through one or more of: summation, averaging, multiplying, selecting a minimum value, or selecting a maximum value.

15. The system of claim 11 , wherein the network switch and the plurality of compute nodes collectively implement the common distributed application.

16. The system of claim 15 , wherein the common distributed application includes one or more artificial neural networks.

17. The system of claim 15 , wherein the common distributed application represents one or more of: deep learning applications, machine learning applications or artificial intelligence applications.

18. The system of claim 11 , wherein an error occurs in processing the vector chunks received from the plurality of compute nodes; wherein a message is sent to each compute nodes in the plurality of compute nodes to inform the error.

19. The system of claim 11 , wherein each compute node in the plurality of compute nodes is implemented at least in part with one or more of: central processing units, graphics processing units, tensor processing units, floating point units, hardware accelerators, or other computing processor.

20. The system of claim 11 , wherein a plurality of network switches operates with the plurality of compute nodes to reduce sets of vector chunks sent by the plurality of compute nodes; wherein the plurality of network switches includes the network switch.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 11, 2024
From: MATTHEWS, WILLIAM BRAD; KWAN, BRUCE HUI; AGARWAL, PUNEET
To: INNOVIUM, INC.
Reel/Frame 066097/0687 →
Continuity (3)
Continuation 17742354 · May 11, 2022
Continuation 17200463 · Mar 12, 2021
Related Publication 20250190273A1 · Jun 12, 2025
References Cited (32)
US 10167800B1 · Chung · 2019 [cited by examiner]
US 10931588B1 · Matthews et al. · 2021 [cited by applicant]
US 10931602B1 · Matthews et al. · 2021 [cited by applicant]
US 11048661B2 · Sankaralingam et al. · 2021 [cited by applicant]
US 11425195B1 · Matthews · 2022 [cited by examiner]
US 11888931B1 · Matthews · 2024 [cited by examiner]
US 20110060891A1 · Jia · 2011 [cited by applicant]
US 20110173413A1 · Chen · 2011 [cited by examiner]
US 20190188239A1 · Serrano et al. · 2019 [cited by applicant]
US 20190303387A1 · Smarda et al. · 2019 [cited by applicant]
Sapio, Amedeo, et al. “Scaling distributed machine learning with {In-Network} aggregation.” 18th USENIX Symposium on Networked Systems Design and Implementation. (Year: 2019). [cited by examiner]
Graham, Richard L., et al. “Scalable hierarchical aggregation and reduction protocol (sharp) tm streaming-aggregation hardware design and evaluation.” High Performance Computing: 35th International Conference, ISC High … [cited by examiner]
Todorov, Ilian T., et al. “DL_POLY_3: new dimensions in molecular dynamics simulations via massive parallelism.” Journal of Materials Chemistry 16.20. (Year: 2006). [cited by examiner]
Thakur, Rajeev, Rolf Rabenseifner, and William Gropp. “Optimization of collective communication operations in MPICH.” The International Journal of High Performance Computing Applications 19.1. (Year: 2005). [cited by examiner]
Abadi et al., “Tensorflow: Large-Scale Machine Learning on Heterogeneous Distributed Systems.” arXiv Preprint arXiv: 1603.04467. (Year: 2016). [cited by applicant]
Chahal et al., “A Hitchhiker's Guide on Distributed Training of Deep Neural Networks,” Journal of Parallel and Distributed Computing 137: 65-76. (Year: 2020). [cited by applicant]
Dean et al., “MapReduce: Simplified Data Processing on Large Clusters.” Communications of the ACM 51.1: 107-113. (Year: 2008). [cited by applicant]
Devendar, “Sharp: In-Network Scalable Streaming Hierarchical Aggregation and Reduction Protocol”, Aug. 31, 2020 (Aug. 31, 2020), pp. 1-36, Retrieved From the Internet: URL: https://mug.mvapich.cse.ohio-state.edu/static/… [cited by applicant]
He et al., “High-Performance Support Vector Machines and Its Applications”, ArXiv abs/1905.00331 (Year: 2019). [cited by applicant]
Liu et al., “The Parallelization of Back Propagation Neural Network in Mapreduce and Spark.” International Journal of Parallel Programming 45: 760-779. (Year: 2017). [cited by applicant]
Luo et al., “Parameter Hub: A Rack-Scale Parameter Server for Distributed Deep Neural Network Training.” Proceedings of the ACM Symposium on Cloud Computer. (Year: 2018). [cited by applicant]
Sapio et al., “Scaling Distributed Machine Learning With In-Network Aggregation”, ArXiV abs/1903.06701 (Year: 2020). [cited by applicant]
Sapio et al., “Scaling Distributed Machine Learning With In-Network Aggregation”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University, Ithaca, NY 14853, Sep. 30, 2020 (Sep. 30, 2020). [cited by applicant]
United States Patent and Trademark Office, U.S. Appl. No. 17/200,463, Non-Final Office Action dated Jan. 3, 2022. [cited by applicant]
United States Patent and Trademark Office, U.S. Appl. No. 17/200,463, Notice of Allowance dated Apr. 7, 2022. [cited by applicant]
United States Patent and Trademark Office, U.S. Appl. No. 17/742,354, Advisory Action dated Aug. 23, 2023. [cited by applicant]
United States Patent and Trademark Office, U.S. Appl. No. 17/742,354, Final Office Action dated Jun. 20, 2023. [cited by applicant]
United States Patent and Trademark Office, U.S. Appl. No. 17/742,354, Non-Final Office Action dated Mar. 17, 2023. [cited by applicant]
United States Patent and Trademark Office, U.S. Appl. No. 17/742,354, Notice of Allowance dated Sep. 18, 2023. [cited by applicant]
Wickramasinghe et al., “A Survey of Methods for Collective Communication Optimization and Tuning”, ArXiV abs/1611.06334 (Year: 2016). [cited by applicant]
World Intellectual Property Organization, Application No. PCT/US22/20087, International Search Report dated Jul. 19, 2022. [cited by applicant]
Zhang et al., “A Comparison of Distributed Machine Learning Platforms.” 2017 26th International Conference on Computer Communication and Networks (ICCCN). IEEE. (Year: 2017). [cited by applicant]