IP Library Granted Patent US 11,675,654
Granted Patent B2
US 11,675,654 · App. 17/553,785 · Granted Jun 13, 2023

Systems and methods for error recovery

Inventors: Bharadwaj Pudipeddi (San Jose, CA); Maral Mesmakhosroshahi (Sunnyvale, CA); Jinwen Xi (Sunnyvale, CA); Saurabh M. Kulkarni (Redmond, WA); Marc Tremblay (Bellevue, WA); Matthias Baenninger (Seattle, WA); Nuno Claudino Pereira Lopes (Cambridge, GB)
Assignee: Microsoft Technology Licensing, LLC
G06F11/0793G06F11/0724G06F11/0751
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,675,654
App. No.
17/553,785
Granted
Jun 13, 2023
Kind
B2
Abstract

Embodiments of the present disclosure include an error recovery method comprising detecting a computing error, restarting a first artificial intelligence processor of a plurality of artificial intelligence processors processing a data set, and loading a model in the artificial intelligence processor, wherein the model corresponds to a same model processed by the plurality of artificial intelligence processors during a previous processing iteration by the plurality of artificial intelligence processors on data from the data set.

Claims (50)

1. A method comprising:

detecting a computing error in a first artificial intelligence processor of a plurality of artificial intelligence processors during a first processing iteration of data from a data set;

eliminating the error from the first artificial intelligence processor;

waiting, by a second artificial intelligence processor of the plurality of artificial intelligence processors, while the error is eliminated from the first artificial intelligence processor; and

loading a model in one or more artificial intelligence processors of the plurality of artificial intelligence processors, wherein the model corresponds to a same model processed by the plurality of artificial intelligence processors during the first processing iteration of the data from the data set.

2. The method of claim 1 comprising:

generating a first result by the first artificial intelligence processor based on the model;

generating a second result by the second artificial intelligence processor based on the model; and

combining the first result and the second result to produce an updated model.

3. The method of claim 1 wherein the error is detected during a result aggregation phase of the first processing iteration, and wherein the second artificial intelligence processor waits for the first artificial intelligence processor to produce a valid result during the result aggregation phase.

4. The method of claim 3 wherein the first artificial intelligence processor sends an invalid result indicator to the second artificial intelligence processor.

5. The method of claim 3 wherein the result aggregation phase includes an All-Reduce operation.

6. The method of claim 1 comprising:

loading different portions of the model in the one or more of the artificial intelligence processors; and

processing a first portion of the data in the one or more artificial intelligence processors, the first portion received by the first artificial intelligence processor in connection with the first processing iteration.

7. The method of claim 1 comprising:

loading the model in the first artificial intelligence processor; and

processing a first portion of the data in the one or more artificial intelligence processors, the first portion received by the first artificial intelligence processor in connection with the first processing iteration.

8. The method of claim 1 wherein the first artificial intelligence processor receives the model from a controller.

9. The method of claim 1 wherein the first artificial intelligence processor receives the model from one or more other processors of the plurality of artificial intelligence processors.

10. The method of claim 1 wherein the first artificial intelligence processor receives the model from a local memory of the first artificial intelligence processor.

11. The method of claim 1 wherein the model includes a set of artificial intelligence parameters.

12. The method of claim 1 wherein the model includes a set of neural network weights.

13. The method of claim 1 wherein the data set includes a training data set.

14. A non-transitory computer readable storage medium having stored thereon program code executable by one or more processors, execution of the program code causing the one or more processors to:

detect an error in a first artificial intelligence processor of a plurality of artificial intelligence processors during a first processing iteration of data from a data set;

eliminate the error from the first artificial intelligence processor;

wait, by a second artificial intelligence processor of the plurality of artificial intelligence processors, while the error is eliminated from the first artificial intelligence processor; and

load a model in one or more artificial intelligence processors of the plurality of artificial intelligence processors, wherein the model corresponds to a same model processed by the plurality of artificial intelligence processors during the first processing iteration of the data from the data set.

15. The non-transitory computer readable storage medium of claim 14 wherein execution of the program code causes the one or more processors to:

generate a first result by the first artificial intelligence processor based on the model;

generate a second result by the second artificial intelligence processor based on the model; and

combine the first result and the second result to produce an updated model.

16. The non-transitory computer readable storage medium of claim 14 wherein execution of the program code causes the one or more processors to:

detect the error during a result aggregation phase of the first processing iteration, wherein the second artificial intelligence processor waits for the first artificial intelligence processor to produce a valid result during the result aggregation phase.

17. The non-transitory computer readable storage medium of claim 14 wherein execution of the program code causes the one or more processors to:

send an invalid result indicator from the first artificial intelligence processor to the second artificial intelligence processor.

18. A system comprising:

a plurality of artificial intelligence processors; and

memory having stored thereon program code, execution of the program code causing the system to:

detect an error in a first artificial intelligence processor of a plurality of artificial intelligence processors during a first processing iteration of data from a data set;

eliminate the error from the first artificial intelligence processor;

wait, by a second artificial intelligence processor of the plurality of artificial intelligence processors, while the error is eliminated from the first artificial intelligence processor; and

load a model in one or more artificial intelligence processors of the plurality of artificial intelligence processors, wherein the model corresponds to a same model processed by the plurality of artificial intelligence processors during the first processing iteration of the data from the data set.

19. The system of claim 18 wherein execution of the program code causes the system to:

generate a first result by the first artificial intelligence processor based on the model;

generate a second result by the second artificial intelligence processor based on the model; and

combine the first result and the second result to produce an updated model.

20. The system of claim 18 wherein execution of the program code causes the system to:

detect the error during a result aggregation phase of the first processing iteration, and wherein the second artificial intelligence processor waits for the first artificial intelligence processor to produce a valid result during the result aggregation phase.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2021
From: PUDIPEDDI, BHARADWAJ; MESMAKHOSROSHAHI, MARAL; XI, JINWEN; KULKARNI, SAURABH M; TREMBLAY, MARC; BAENNINGER, MATTHIAS; CLAUDINO PEREIRA LOPES, NUNO
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 058411/0500 →
Continuity (3)
Continuation 16833191 · Mar 27, 2020
Provisional Application 62966019 · Jan 26, 2020
Related Publication 20220107864A1 · Apr 7, 2022
Cited By (6)
US 12,399,687 US 12,585,435 US 12,625,680 US 12,645,429 US 12,650,836 US 12,699,556