IP Library Granted Patent US 11,658,923
Granted Patent B2
US 11,658,923 · App. 17/221,538 · Granted May 23, 2023

Forwarding element data plane performing floating point computations

Inventors: Masoud Moshref Javadi (San Jose, CA); Changhoon Kim (Palo Alto, CA); Patrick W. Bosshart (Plano, TX); Anurag Agrawal (Santa Clara, CA)
Assignee: Barefoot Networks, Inc.
H04L49/3063G06F9/45558G06N3/08G06F2009/45595
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,658,923
App. No.
17/221,538
Granted
May 23, 2023
Kind
B2
Abstract

Some embodiments provide a network forwarding element with a data-plane forwarding circuit that has a parameter collecting circuit to store and distribute parameter values computed by several machines in a network. In some embodiments, the machines perform distributed computing operations, and the parameter values that compute are parameter values associated with the distributed computing operations. The parameter collecting circuit of the data-plane forwarding circuit (data plane) in some embodiments (1) stores a set of parameter values computed and sent by a first set of machines, and (2) distributes the collected parameter values to a second set of machines once it has collected the set of parameter values from all the machines in the first set. The first and second sets of machines are the same set of machines in some embodiments, while they are different sets of machines (e.g., one set has at least one machine that is not in the other set) in other embodiments. In some embodiments, the parameter collecting circuit performs computations on the parameter values that it collects and distributes the result of the computations once it has processed all the parameter values distributed by the first set of machines. The computations are aggregating operations (e.g., adding, averaging, etc.) that combine corresponding subset of parameter values distributed by the first set of machines.

Claims (48)

1. An apparatus comprising:

packet processing circuitry to:

perform computations based on floating-point parameter values extracted from one or more received packets; and

provide results of the computations in one or more packets for transmission to at least one destination, wherein the perform computations or provide results comprise performance of one or more match action operations;

wherein:

the results of the computations are to be provided as one or more key-value pairs in one or more protocol headers of the one or more packets for the transmission to the at least one destination; and

the results of the computations are for use in one or more distributed machine-learning operations.

2. The apparatus of claim 1 , wherein the packet processing circuitry comprises:

at least one programmable packet processing stage to perform packet

forwarding

operations on received packets and

at least one programmable packet processing stage to perform computations and provide results.

3. The apparatus of claim 1 , wherein the packet processing circuitry comprises

at least one programmable packet processing stage that is programmed to perform the computations.

4. The apparatus of claim 1 , wherein the computations comprise one or more of: addition or averaging.

5. The apparatus of claim 1 , wherein the computations are associated with at least one machine learning training operation.

6. The apparatus of claim 2 , wherein the at least one machine learning training operation comprises an all reduce, or all gather operation.

7. The apparatus of claim 1 , wherein the apparatus further comprises a switch that comprises at least one ingress port to receive one or more packets.

8. The apparatus of claim 1 , wherein the apparatus further comprises a switch that comprises at least one egress port to receive one or more packets.

9. A non-transitory computer readable medium storing instructions, that when executed, configure a packet processing circuitry of a forwarding element to perform operations comprising:

perform computations based on floating-point parameter values in one or more received packets received by the forwarding element; and

provide results of the computations in one or more packets for transmission by the forwarding element to at least one destination, wherein the perform computations or provide results comprise performance of one or more match action operations by the packet processing circuitry;

wherein:

the results of the computations are to be provided as one or more key-value pairs in one or more protocol headers of the one or more packets for the transmission to the at least one destination; and

the results of the computations are for use in one or more distributed machine-learning operations.

10. The non-transitory computer readable medium of claim 9 , wherein the packet processing circuitry comprises:

at least one programmable packet processing stage to perform packet forwarding operations on received packets and

at least one programmable packet processing stage to perform the computations and provide results.

11. The non-transitory computer readable medium of claim 9 , wherein the packet processing circuitry comprises:

at least one programmable packet processing stage that is programmed to perform the computations.

12. The non-transitory computer readable medium of claim 9 , wherein the computations comprise one or more of: addition or averaging.

13. The non-transitory computer readable medium of claim 9 , wherein the computations are associated with at least one machine learning training operation.

14. The non-transitory computer readable medium of claim 9 , wherein the at least one machine learning training operation comprises an all reduce or all gather operation.

15. A method comprising:

in a packet processing circuitry of a forwarding element:

perform computations based on floating-point parameter values in one or more received packets received by the forwarding element; and

provide results of the computations in one or more packets for transmission by the forwarding element to at least one destination, wherein the perform computations or provide results comprise performance of one or more match action operations by the packet processing circuitry;

wherein:

the results of the computations are to be provided as one or more key-value pairs in one or more protocol headers of the one or more packets for the transmission to the at least one destination; and

the results of the computations are for use in one or more distributed machine-learning operations.

16. The method of claim 15 , comprising:

in at least one programmable packet processing stage of the packet processing circuitry, performing packet forwarding operations on received packets and

in at least one programmable packet processing stage of the packet processing circuitry, performing the computations and providing the results.

17. The method of claim 15 , wherein the computations comprise one or more of: addition or averaging.

18. The method of claim 15 , wherein the computations are associated with at least one machine learning training operation.

19. The method of claim 18 , wherein the at least one machine learning training operation comprises an all reduce or all gather operation.

20. The method of claim 15 , comprising:

receiving a configuration to program the packet processing circuitry of the forwarding element to perform the computations and provide the results of the computations in the one or more packets for the transmission by the forwarding element to the at least one destination.

Continuity (4)
Continuation 16147755 · Sep 30, 2018
Provisional Application 62733441 · Sep 19, 2018
Provisional Application 62718373 · Aug 13, 2018
Related Publication 20210399997A1 · Dec 23, 2021