IP Library › Granted Patent US 12,294,483
Granted Patent B2
US 12,294,483 · App. 18/302,629 · Granted May 6, 2025

Fault recovery plan determining method, apparatus, and system, and computer storage medium

Inventors: Yuming Xie (Nanjing, CN); Ye Li (Shenzhen, CN); Juchang Li (Nanjing, CN); Yunpeng Gao (Nanjing, CN); Yanxiang Hou (Beijing, CN); Yancheng Yang (Shenzhen, CN)
Assignee: Huawei Technologies Co., Ltd.
H04L41/0631H04L41/0627H04L41/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,294,483
App. No.
18/302,629
Granted
May 6, 2025
Kind
B2
Abstract

This application discloses a fault recovery plan determining method, apparatus, and system, and a computer storage medium, and belongs to the field of network technologies. First, a control device obtains a similar known fault that is in a plurality of known faults and whose fault root cause and a fault root cause of a target fault in a network meet a similarity condition. Then, the control device obtains a fault recovery plan corresponding to the similar known fault. The control device determines, based on the fault recovery plan corresponding to the similar known fault, a fault recovery plan corresponding to the target fault.

Claims (58)

1. A fault recovery plan determining method, comprising:

obtaining, by a control device, a similar known fault that is in a plurality of known faults and whose fault root cause and a fault root cause of a target fault in a network meet a similarity condition, wherein the fault root cause is represented by a fault root cause feature, the fault root cause feature comprises a fault root cause object, a fault root cause event, the fault root cause event is an abnormal event causing a fault, the fault root cause object indicates a type of a fault root cause network entity, wherein the type of a fault root cause network entity includes an interface type, a protocol type, or a service type, and the fault root cause network entity is a network entity to which the fault root cause event belongs;

obtaining, by the control device, a fault recovery plan corresponding to the similar known fault; and

determining, by the control device based on the fault recovery plan corresponding to the similar known fault, a fault recovery plan corresponding to the target fault.

2. The method according to claim 1 , wherein

the fault root cause network entity is a physical interface, and the fault root cause feature further comprises one or more of an intermittent interface disconnection indication of the fault root cause network entity, an interface suspension indication of the fault root cause network entity, a packet receiving/sending state of the fault root cause network entity, an interface protocol status of the fault root cause network entity, or a physical interface status of a device in which the fault root cause network entity is located;

the fault root cause network entity is a border gateway protocol (BGP) peer, and the fault root cause feature further comprises a BGP route flapping indication of the fault root cause network entity and/or a physical interface status of a device in which the fault root cause network entity is located; or

the fault root cause feature further comprises a physical interface status of a device in which the fault root cause network entity is located.

3. The method according to claim 1 , wherein the obtaining, by a control device, a similar known fault that is in a plurality of known faults and whose fault root cause and a fault root cause of a target fault in a network meet a similarity condition comprises:

obtaining, by the control device, fault root cause features of the plurality of known faults;

for each of the plurality of known faults, calculating, by the control device, a similarity between the fault root cause of the target fault and a fault root cause of the known fault based on a fault root cause feature of the target fault and the fault root cause feature of the known fault; and

determining, by the control device based on the similarities between the fault root cause of the target fault and the fault root causes of the plurality of known faults, the similar known fault that is in the plurality of known faults and whose fault root cause and the fault root cause of the target fault meet the similarity condition.

4. The method according to claim 3 , wherein the calculating, by the control device, a similarity between the fault root cause of the target fault and a fault root cause of the known fault based on a fault root cause feature of the target fault and the fault root cause feature of the known fault comprises:

inputting, by the control device, the fault root cause feature of the target fault and the fault root cause feature of the known fault into a similarity model, to obtain the similarity that is between the fault root cause of the target fault and the fault root cause of the known fault and that is output by the similarity model, wherein the similarity model is obtained through training based on fault root cause features of a plurality of sample faults, the sample fault is marked with a category label, and sample faults marked with a same category label correspond to a same fault recovery plan.

5. The method according to claim 4 , wherein the method further comprises:

inputting, by the control device, fault root cause features of a plurality of sample fault pairs into the similarity model in batches, to obtain a similarity that is between fault root causes of each sample fault pair and that is output by the similarity model, wherein the plurality of sample fault pairs comprise a first-category sample fault pair and a second-category sample fault pair, the first-category sample fault pair comprises two sample faults marked with a same category label, and the second-category sample fault pair comprises two sample faults marked with different category labels; and

determining, by the control device, a similarity threshold based on the similarities between the fault root causes of the plurality of sample fault pairs.

6. The method according to claim 1 , wherein the determining, by the control device based on the fault recovery plan corresponding to the similar known fault, a fault recovery plan corresponding to the target fault comprises:

evaluating, by the control device based on a network configuration of the network, feasibility of the fault recovery plan corresponding to the similar known fault, wherein the network configuration comprises a networking topology and/or device data, and the device data comprises one or more of management plane data, data plane data, or control plane data; and

determining, by the control device, one or more of feasible fault recovery plans as the fault recovery plan corresponding to the target fault.

7. The method according to claim 6 , wherein the determining, by the control device, one or more of feasible fault recovery plans as the fault recovery plan corresponding to the target fault comprises:

in response to a case in which a plurality of fault recovery plans are feasible, separately evaluating, by the control device based on the network configuration of the network, degrees of impact of the plurality of fault recovery plans on a service running in the network; and

determining, by the control device, a fault recovery plan that is in the plurality of fault recovery plans and that has a minimum degree of impact on the service running in the network as the fault recovery plan corresponding to the target fault.

8. The method according to claim 1 , wherein the method further comprises:

determining, by the control device based on the target fault and the fault recovery plan corresponding to the target fault, a target network device of a to-be-executed plan in the network; and

sending, by the control device, a plan execution instruction to the target network device, wherein the plan execution instruction is used to instruct the target network device to execute the fault recovery plan corresponding to the target fault, and the plan execution instruction comprises the fault recovery plan corresponding to the target fault.

9. A device for determining fault recovery plan, wherein the device comprises:

at least one processor; and

at least one memory, coupled to the at least one processor and configured to store instructions that when executed by the at least one processor cause the device to:

obtain a similar known fault that is in a plurality of known faults and whose fault root cause and a fault root cause of a target fault in a network meet a similarity condition, wherein the fault root cause is represented by a fault root cause feature, the fault root cause feature comprises a fault root cause object, a fault root cause event, the fault root cause event is an abnormal event causing a fault, the fault root cause object indicates a type of a fault root cause network entity, wherein the type of a fault root cause network entity includes an interface type, a protocol type, or a service type, and the fault root cause network entity is a network entity to which the fault root cause event belongs;

obtain a fault recovery plan corresponding to the similar known fault; and

determine, based on the fault recovery plan corresponding to the similar known fault, a fault recovery plan corresponding to the target fault.

10. The device according to claim 9 , wherein

the fault root cause network entity is a physical interface, and the fault root cause feature further comprises one or more of an intermittent interface disconnection indication of the fault root cause network entity, an interface suspension indication of the fault root cause network entity, a packet receiving/sending state of the fault root cause network entity, an interface protocol status of the fault root cause network entity, or a physical interface status of a device in which the fault root cause network entity is located;

the fault root cause network entity is a border gateway protocol (BGP) peer, and the fault root cause feature further comprises a BGP route flapping indication of the fault root cause network entity and/or a physical interface status of a device in which the fault root cause network entity is located; or

the fault root cause feature further comprises a physical interface status of a device in which the fault root cause network entity is located.

11. The device according to claim 9 , wherein when executed by the at least one processor, the instructions further cause the device to:

obtain fault root cause features of the plurality of known faults;

for each of the plurality of known faults, calculate a similarity between the fault root cause of the target fault and a fault root cause of the known fault based on a fault root cause feature of the target fault and the fault root cause feature of the known fault; and

determine, based on the similarities between the fault root cause of the target fault and the fault root causes of the plurality of known faults, the similar known fault that is in the plurality of known faults and whose fault root cause and the fault root cause of the target fault meet the similarity condition.

12. The device according to claim 11 , wherein when executed by the at least one processor, the instructions further cause the device to:

input the fault root cause feature of the target fault and the fault root cause feature of the known fault into a similarity model, to obtain the similarity that is between the fault root cause of the target fault and the fault root cause of the known fault and that is output by the similarity model, wherein the similarity model is obtained through training based on fault root cause features of a plurality of sample faults, the sample fault is marked with a category label, and sample faults marked with a same category label correspond to a same fault recovery plan.

13. The device according to claim 12 , wherein when executed by the at least one processor, the instructions further cause the device to:

input fault root cause features of a plurality of sample fault pairs into the similarity model in batches, to obtain a similarity that is between fault root causes of each sample fault pair and that is output by the similarity model, wherein the plurality of sample fault pairs comprise a first-category sample fault pair and a second-category sample fault pair, the first-category sample fault pair comprises two sample faults marked with a same category label, and the second-category sample fault pair comprises two sample faults marked with different category labels; and

determine a similarity threshold based on the similarities between the fault root causes of the plurality of sample fault pairs.

14. The device according to claim 9 , wherein when executed by the at least one processor, the instructions further cause the device to:

evaluate, based on a network configuration of the network, feasibility of the fault recovery plan corresponding to the similar known fault, wherein the network configuration comprises a networking topology and/or device data, and the device data comprises one or more of management plane data, data plane data, or control plane data; and

determine one or more of feasible fault recovery plans as the fault recovery plan corresponding to the target fault.

15. The device according to claim 14 , wherein when executed by the at least one processor, the instructions further cause the device to:

in response to a case in which a plurality of fault recovery plans are feasible, separately evaluate, based on the network configuration of the network, degrees of impact of the plurality of fault recovery plans on a service running in the network; and

determine, a fault recovery plan that is in the plurality of fault recovery plans and that has a minimum degree of impact on the service running in the network as the fault recovery plan corresponding to the target fault.

16. The device according to claim 9 , wherein when executed by the at least one processor, the instructions further cause the device to:

determine, based on the target fault and the fault recovery plan corresponding to the target fault, a target network device of a to-be-executed plan in the network; and

send a plan execution instruction to the target network device, wherein the plan execution instruction is used to instruct the target network device to execute the fault recovery plan corresponding to the target fault, and the plan execution instruction comprises the fault recovery plan corresponding to the target fault.

17. One or more non-transitory computer-readable media storing computer instructions, that when executed by one or more processors, cause a computing device to perform operations comprising:

obtaining a similar known fault that is in a plurality of known faults and whose fault root cause and a fault root cause of a target fault in a network meet a similarity condition, wherein the fault root cause is represented by a fault root cause feature, the fault root cause feature comprises a fault root cause object, a fault root cause event, the fault root cause event is an abnormal event causing a fault, the fault root cause object indicates a type of a fault root cause network entity, wherein the type of a fault root cause network entity includes an interface type, a protocol type, or a service type, and the fault root cause network entity is a network entity to which the fault root cause event belongs;

obtaining a fault recovery plan corresponding to the similar known fault; and

determining, based on the fault recovery plan corresponding to the similar known fault, a fault recovery plan corresponding to the target fault.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2025
From: XIE, YUMING; LI, YE; LI, JUCHANG; GAO, YUNPENG; HOU, YANXIANG; YANG, YANCHENG
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 070617/0659 →
Priority Claims (2)
CN 202011123661.5 · Oct 20, 2020 · national
CN 202011622270.8 · Dec 31, 2020 · national
Continuity (2)
Continuation PCTCN2021124377 · Oct 18, 2021
Related Publication 20230318906A1 · Oct 5, 2023
References Cited (11)
US 11442652B1 · Dailey · 2022 [cited by examiner]
US 20190294490A1 · Zhang · 2019 [cited by examiner]
US 20230161637A1 · Soualhia · 2023 [cited by examiner]
CN 103473409B · 2016 [cited by applicant]
CN 108090567A · 2018 [cited by applicant]
CN 105467975B · 2018 [cited by applicant]
CN 111082401A · 2020 [cited by applicant]
CN 111209472A · 2020 [cited by applicant]
Pernici et al., “Automatic Learning of Repair Strategies for Web Services,” Proceedings of the Fifth European Conference on Web Services, Nov. 1, 2007, pp. 119-128. [cited by applicant]
Extended European Search Report in European Appln No. 21881956.3, dated Jan. 24, 2024, 10 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/CN2021/124377, mailed on Jan. 14, 2022, 17 pages (with English translation). [cited by applicant]