IP Library › Granted Patent US 10,812,388
Granted Patent B2
US 10,812,388 · App. 16/269,723 · Granted Oct 20, 2020

Preventing damage to flows in an SDN fabric by predicting failures using machine learning

Inventors: Pascal Thubert (La Colle sur Loup, FR); Jean-Philippe Vasseur (Saint Martin d'uriage, FR); Eric Levy-Abegnoli (Valbonne, FR); Patrick Wetterwald (Mouans Sartoux, FR)
Assignee: Cisco Technology, Inc.
H04L47/122H04L41/16H04L45/64H04L47/2483
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,812,388
App. No.
16/269,723
Filed
Feb 7, 2019
Granted
Oct 20, 2020
Kind
B2
Art Unit
2474
USPC
370/216
Abstract

In one embodiment, a supervisory device for a software defined networking (SDN) fabric predicts a failure in the SDN fabric using a machine learning-based failure prediction model. The supervisory device identifies a plurality of traffic flows having associated leaves in the SDN fabric that would be affected by the predicted failure. The supervisory device selects a subset of the identified plurality of traffic flows and their associated leaves. The supervisory device disaggregates routes for the selected subset of traffic flows and their associated leaves, to avoid the predicted failure.

Claims (49)

1. A method comprising:

predicting, by a supervisory device for a software defined networking (SDN) fabric, a failure in the SDN fabric using a machine learning-based failure prediction model;

identifying, by the supervisory device, a plurality of traffic flows having associated leaves in the SDN fabric that would be affected by the predicted failure;

selecting, by the supervisory device, a subset of the identified plurality of traffic flows and associated leaves of the subset of the identified traffic flows;

identifying, by the supervisory device, critical traffic flows which require continuous service among the selected subset of traffic flows; and

disaggregating, by the supervisory device, routes only for the critical traffic flows and associated leaves of the critical traffic flows, to avoid the predicted failure.

2. The method as in claim 1 , wherein predicting the failure in the SDN fabric comprises:

using link telemetry data regarding links in the SDN fabric as input to the failure prediction model.

3. The method as in claim 1 , wherein selecting the subset of the identified plurality of traffic flows and associated leaves of the subset of the identified traffic flows comprises:

computing a size of the subset based on a number of the affected traffic flows that can be rerouted based on a current state of the SDN fabric; and

populating the subset with traffic flows from the plurality based on their associated measures of criticality.

4. The method as in claim 1 , wherein the subset of the identified plurality of traffic flows is selected based in part on a false positive rate of the failure prediction model.

5. The method as in claim 1 , wherein disaggregating routes only for the critical traffic flows and associated leaves of the critical traffic flows, to avoid the predicted failure, comprises:

sending a unicast routing message to a leaf associated with one or more of traffic flows of the critical traffic flows, to cause that leaf to reroute the one or more traffic flows to a different plane partition of the SDN fabric.

6. The method as in claim 5 , wherein the supervisory device is a top of fabric device in the SDN fabric.

7. The method as in claim 5 , wherein the unicast routing message identifies the one or more traffic flows, and wherein the leaf receiving the unicast routing message is configured to reroute those flows on a per-flow basis.

8. The method as in claim 1 , wherein the failure prediction model is configured to predict a timeframe associated with the predicted failure, and wherein the subset of traffic flows is selected based in part on the predicted timeframe associated with the predicted failure.

9. An apparatus, comprising:

one or more network interfaces to communicate with a software defined networking (SDN) fabric;

a processor coupled to the network interfaces and configured to execute one or more processes; and

a memory configured to store a process executable by the processor, the process when executed configured to:

predict a failure in the SDN fabric using a machine learning-based failure prediction model;

identify a plurality of traffic flows having associated leaves in the SDN fabric that would be affected by the predicted failure;

select a subset of the identified plurality of traffic flows and associated leaves of the subset of the identified traffic flows;

identifying, by the supervisory device, critical traffic flows which require continuous service among the selected subset of traffic flows; and

disaggregate routes only for the critical traffic flows and associated leaves of the critical traffic flows, to avoid the predicted failure.

10. The apparatus as in claim 9 , wherein the apparatus predicts the failure in the SDN fabric comprises:

using link telemetry data regarding links in the SDN fabric as input to the failure prediction model.

11. The apparatus as in claim 9 , wherein the apparatus selects the subset of the identified plurality of traffic flows and associated leaves of the subset of the identified traffic flows by:

computing a size of the subset based on a number of the affected traffic flows that can be rerouted based on a current state of the SDN fabric; and

populating the subset with traffic flows from the plurality based on their associated measures of criticality.

12. The apparatus as in claim 1 , wherein the subset of the identified plurality of traffic flows is selected based in part on a false positive rate of the failure prediction model.

13. The apparatus as in claim 9 , wherein the apparatus disaggregates routes only for the critical traffic flows and associated leaves of the critical traffic flows, to avoid the predicted failure, by:

sending a unicast routing message to a leaf associated with one or more of traffic flows of the critical traffic flows, to cause that leaf to reroute the one or more traffic flows to a different plane partition of the SDN fabric.

14. The apparatus as in claim 13 , wherein the apparatus disaggregates routes for the selected subset of traffic flows and associated leaves, to avoid the predicted failure, by:

sending a unicast routing message to a node in the fabric associated with one or more of the traffic flows in the selected subset, to cause that node to reroute the one or more traffic flows to a different parent of the node.

15. The apparatus as in claim 13 , wherein the unicast routing message identifies the one or more traffic flows, and wherein the leaf receiving the unicast routing message is configured to reroute those flows on a per-flow basis.

16. The apparatus as in claim 9 , wherein the failure prediction model is configured to predict a timeframe associated with the predicted failure, and wherein the subset of traffic flows is selected based in part on the predicted timeframe associated with the predicted failure.

17. A tangible, non-transitory, computer-readable medium storing program instructions that cause a supervisory device for a software defined networking (SDN) fabric to execute a process comprising:

predicting, by the supervisory device for the SDN fabric, a failure in the SDN fabric using a machine learning-based failure prediction model;

identifying, by the supervisory device, a plurality of traffic flows having associated leaves in the SDN fabric that would be affected by the predicted failure;

selecting, by the supervisory device, a subset of the identified plurality of traffic flows and associated leaves of the subset of the identified traffic flows;

identifying, by the supervisory device, critical traffic flows which require continuous service among the selected subset of traffic flows; and

disaggregating, by the supervisory device, routes only for the critical traffic flows and associated leaves of the critical traffic flows, to avoid the predicted failure.

18. The computer-readable medium as in claim 17 , wherein predicting the failure in the SDN fabric comprises:

using link telemetry data regarding links in the SDN fabric as input to the failure prediction model.

19. The computer-readable medium as in claim 17 , wherein disaggregating routes for the selected subset of traffic flows and associated leaves of the subset of the identified traffic flows, to avoid the predicted failure, comprises:

sending a unicast routing message to a leaf associated with one or more of the traffic flows in the selected subset.

20. The computer-readable medium as in claim 17 , wherein the supervisory device is a top of fabric device in the SDN fabric.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2019
From: THUBERT, PASCAL; VASSEUR, JEAN-PHILIPPE; LEVY-ABEGNOLI, ERIC; WETTERWALD, PATRICK
To: CISCO TECHNOLOGY, INC.
Reel/Frame 048264/0223 →
Continuity (1)
Related Publication 20200259746A1 · Aug 13, 2020
Cited By (1)
US 12,301,690