IP Library › Granted Patent US 12,483,462
Granted Patent B2
US 12,483,462 · App. 17/989,998 · Granted Nov 25, 2025

Cloud network failure auto-correlator

Inventors: Tianqiong Luo (Santa Clara, CA); Hui Liu (San Ramon, CA); Hongkun Yang (San Jose, CA); Gargi Adhav (San Jose, CA); Anantanarayanan Govindarajan Iyengar (Saratoga, CA); Yihan Zhang (Evanston, IL)
Assignee: Google LLC
H04L41/0645H04L41/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,483,462
App. No.
17/989,998
Granted
Nov 25, 2025
Kind
B2
Abstract

Analysis of a root cause of errors within a cloud network is manually complex and computationally intensive. Methods and systems are provided to determine a subset of elements of the cloud network to analyze, and to identify a subset of analyzers for analyzing the subset of elements to determine the root cause for the error. Thus, when configuring a network, a user may be provided with an identification of the root cause of error, enabling the user to quickly identify and correct the error.

Claims (45)

1 . A method of evaluating a root cause of a cloud network failure, the method comprising:

receiving, at a computing system interface, one or more triggers for analysis, the one or more triggers corresponding to a failure within the cloud network;

providing access to configuration data in response to receiving the one or more triggers, the configuration data being indicative of configuration of the cloud network and being subject to update, and determining a set of configuration changes based on the configuration data;

comparing, using one or more processors, the one or more triggers against the set of configuration changes to generate a subset of configuration changes that could have caused the failure;

determining, using the one or more processors, a scope of analysis within the cloud network based on the subset of configuration changes and the one or more triggers, the scope of analysis corresponding to a portion of the cloud network to be analyzed and including a set of network resources which are related to a set of network resources directly affected by the failure, wherein knowledge based dependency graphs are used to find the set of network resources which are related;

selecting, using the one or more processors, a set of analyzers from an analyzer module based on at least the scope of the analysis, the set of analyzers being associated with respective ones of the knowledge based dependency graphs, and the knowledge based dependency graphs evolving over time based on events of a scheduler queue and through the evolving over time of expert rules on which the knowledge based dependency graphs are built, the expert rules evolving over time according to user feedback with the extent of the evolving determined by machine learning;

performing, using the one or more processors, an analysis of the determined scope of the cloud network by using the selected set of analyzers to analyze the failure, wherein the selected set of analyzers determine potential causes of the failure;

determining, using the one or more processors, the root cause of the failure based on correlating the potential causes;

providing, via a user interface, an indication of the root cause to a user;

wherein a scheduler module is to receive the one or more triggers for analysis; and

wherein the scheduler queue contains a list of triggers or events, and wherein the list of triggers or events are used in scheduling analysis of the cloud network.

2 . The method of claim 1 wherein the scope of analysis is determined using a software module trained using a machine learning model.

3 . The method of claim 2 wherein the machine learning model is trained using a set of expert rules.

4 . The method of claim 1 wherein the scope of analysis is determined based on a software module trained using expert rules.

5 . The method of claim 1 wherein the selection of the set of analyzers is based on a trained machine learning model.

6 . The method of claim 1 wherein a scheduler initiates the evaluation of the root cause at predetermined intervals.

7 . The method of claim 1 wherein a correlator builds the knowledge based dependency graphs.

8 . A system comprising one or more processors coupled to a non-transitory memory, the non-transitory memory comprising instructions which when executed by the processors perform the steps of:

receiving, at an interface of the system, one or more triggers for analysis, the one or more triggers corresponding to a failure within the cloud network;

providing access to configuration data in response to receiving the one or more triggers, the configuration data being indicative of configuration of the cloud network and being subject to update, and determining a set of configuration changes based on the configuration data;

comparing, using the one or more processors, the one or more triggers against the set of configuration changes to generate a subset of configuration changes that could have caused the failure;

determining, using the one or more processors, the scope of analysis within the cloud network based on the subset of configuration changes and the one or more triggers, the scope of analysis corresponding to a portion of the cloud network to be analyzed and including a set of network resources which are related to a set of network resources directly affected by the failure, wherein knowledge based dependency graphs are used to find related resources;

selecting, using the one or more processors, a set of analyzers from an analyzer module based on at least the scope of the analysis, the set of analyzers being associated with respective ones of the knowledge based dependency graphs, and the knowledge based dependency graphs evolving over time based on events of a scheduler queue and through the evolving over time of expert rules on which the knowledge based dependency graphs are built, the expert rules evolving over time according to user feedback with the extent of the evolving determined by machine learning;

performing, using the one or more processors, an analysis of the determined scope of the cloud network by using the selected set of analyzers to analyze the failure, wherein the selected set of analyzers determine potential causes of the failure;

determining, using the one or more processors, the root cause of the failure based on correlating the potential causes;

providing, via a user interface, an indication of the root cause to a user;

wherein a scheduler module is to receive the one or more triggers for analysis; and

wherein the scheduler queue contains a list of triggers or events, and wherein the list of triggers or events are used in scheduling analysis of the cloud network.

9 . The system of claim 8 wherein analyzers within the analyzer module have a plurality of hierarchies.

10 . The system of claim 9 wherein the plurality of hierarchies correspond to logical levels within the cloud network.

11 . The system of claim 8 further comprising a model module, the model module comprising a plurality of models, wherein each model of the plurality of models is to select a scope of the cloud network or resources within the cloud network for analysis.

12 . The system of claim 10 wherein one or more models are selected based on the one or more triggers.

13 . A non-transient computer readable medium containing program instructions, the instructions when executed perform the steps of:

receiving, at a computing system interface, one or more triggers for analysis, the one or more triggers corresponding to a failure within the cloud network;

providing access to configuration data in response to receiving the one or more triggers, the configuration data being indicative of configuration of the cloud network and being subject to update, and determining a set of configuration changes based on the configuration data;

comparing, using one or more processors, the one or more triggers against the set of configuration changes to generate a subset of configuration changes that could have caused the failure;

determining, using the one or more processors, the scope of analysis within the cloud network based on the subset of configuration changes and the one or more triggers, the scope of analysis corresponding to a portion of the cloud network to be analyzed and including a set of network resources which are related to a set of network resources directly affected by the failure, wherein knowledge based dependency graphs are used to find related resources;

selecting, using the one or more processors, a set of analyzers from an analyzer module based on at least the scope of the analysis, the set of analyzers being associated with respective ones of the knowledge based dependency graphs, and the knowledge based dependency graphs evolving over time based on events of a scheduler queue and through the evolving over time of expert rules on which the knowledge based dependency graphs are built, the expert rules evolving over time according to user feedback with the extent of the evolving determined by machine learning;

performing, using the one or more processors, an analysis of the determined scope of the cloud network by using the selected set of analyzers to analyze the failure, wherein the selected set of analyzers determine potential causes of the failure;

determining, using the one or more processors, the root cause of the failure based on correlating the potential causes;

providing, via a user interface, an indication of the root cause to a user;

wherein a scheduler module is to receive the one or more triggers for analysis; and

wherein the scheduler queue contains a list of triggers or events, and wherein the list of triggers or events are used in scheduling analysis of the cloud network.

14 . The non-transient computer readable medium of claim 13 wherein the scope of analysis is done based on a software module trained using a machine learning model.

15 . The non-transient computer readable medium of claim 14 wherein the machine learning model is trained using a set of expert rules.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 30, 2022
From: LUO, TIANQIONG; IYENGAR, ANANTANARAYANAN GOVINDARAJAN; YANG, HONGKUN; ADHAV, GARGI; ZHANG, YIHAN; LIU, HUI
To: GOOGLE LLC
Reel/Frame 061926/0007 →
Continuity (2)
Provisional Application 63281990 · Nov 22, 2021
Related Publication 20230164022A1 · May 25, 2023
References Cited (15)
US 10855536B1 · Blackburn et al. · 2020 [cited by applicant]
US 20040187048A1 · Angamuthu · 2004 [cited by examiner]
US 20090222697A1 · Thakkar · 2009 [cited by examiner]
US 20170019315A1 · Tapia · 2017 [cited by examiner]
US 20170034010A1 · Fong · 2017 [cited by examiner]
US 20170372212A1 · Zasadzinski · 2017 [cited by examiner]
US 20180248905A1 · Côté et al. · 2018 [cited by applicant]
US 20190306023A1 · Vasseur et al. · 2019 [cited by applicant]
US 20200110761A1 · Cooper · 2020 [cited by examiner]
US 20200379875A1 · Krishnaswamy · 2020 [cited by examiner]
US 20210152416A1 · A · 2021 [cited by examiner]
US 20210350253A1 · Wang · 2021 [cited by examiner]
US 20220004546A1 · Rogers · 2022 [cited by examiner]
LU 101632B1 · 2021 [cited by applicant]
Extended European Search Report for European Patent Application No. 22208840.3 dated Mar. 22, 2023. 8 pages. [cited by applicant]