IP Library › Granted Patent US 12,511,112
Granted Patent B2
US 12,511,112 · App. 18/191,031 · Granted Dec 30, 2025

Automated update management for cloud services

Inventors: Nidhi Verma (Bellevue, WA); Rahul Nigam (Bothell, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F8/65G06F11/0751H04W28/0268
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,112
App. No.
18/191,031
Granted
Dec 30, 2025
Kind
B2
Abstract

An orchestration system implements a rollout service that deploys a series of updates to a cloud service while minimizing an impact of a regression caused in the cloud service by one of the updates. The system includes an orchestrator host computer hosting the rollout service; a network interface with a network on which the cloud service is provided; and a database of deployment policy information and records of previous updates to the cloud service. The rollout service automatically determines a deployment policy for an update using the database, implements a deployment of the update according to the deployment policy, monitors for evidence of a regression caused by the update, and identifies occurrence of the regression caused by the update to the cloud service to enable mitigation of an impact of the regression.

Claims (50)

1 . An orchestration system for implementing a rollout service that deploys a series of updates to a cloud service while minimizing an impact of a regression caused in the cloud service by one of the updates, the system comprising:

a network interface with a network on which the cloud service is provided;

a database of deployment policy information and records of previous updates to the cloud service; and

an orchestrator host computer hosting the rollout service, the orchestrator host computer comprising a processor and memory storing instructions that, when executed by the processor, cause the orchestrator host computer to perform operations comprising:

determining a deployment policy for a first update using the database,

determining, based on a risk of the first update, a frequency for checking a health metric of a first state of the cloud service, wherein the first state of the cloud service corresponds to the first update,

deploying the first update according to the deployment policy,

checking, at the determined frequency for checking the health metric, whether or not the health metric of the first state of the cloud service indicates that the regression in the cloud service has occurred,

in response to determining that the health metric indicates the regression in the cloud service has not occurred, tagging the first state as a known good state, and

in response to identifying an occurrence of a regression caused by a second update to the cloud service subsequent to the first update, returning the cloud service to the first state tagged as the known good state.

2 . The orchestration system of claim 1 , further comprising a service health engine to automatically detect a regression based on Quality of Service (QOS) signals from components supporting the cloud service.

3 . The orchestration system of claim 2 , wherein the rollout service receives a health signal from the service health engine and, in response to the health signal indicating the regression caused by the second update with a magnitude exceeding a threshold, halts rollout of the second update.

4 . The orchestration system of claim 2 , wherein the service health engine focuses on different QoS signals based on a type of update being deployed, where past regressions caused by different types of updates have been correlated to different QoS signals.

5 . The orchestration system of claim 2 , wherein the service health engine compares current QOS signals while rolling out a current build with QoS signals recorded during rollout of a previous build, the rollout service detecting a regression based on a change in the current QoS signals corresponding to the current build as compared with the QoS signals corresponding to the previous build.

6 . The orchestration system of claim 1 , wherein the rollout service quantifies the impact of the regression caused by the second update and, when the impact exceeds a threshold, takes action to mitigate the regression caused by the second update.

7 . The orchestration system of claim 6 , further comprising:

a state machine that stores a succession of states of the cloud service, wherein the rollout service tags a number of states in the succession of states as “good” based on telemetry received,

the rollout service to return the cloud service to a last known good state in response to the regression caused by the second update by implementing the last known good state from data in the state machine.

8 . The orchestration system of claim 7 , wherein the rollout service increases a rate at which states are checked and tagged as “good” in response to an increasing level of risk of a current update.

9 . The orchestration system of claim 1 , further comprising a policy service storing indications of which segments of a userbase are sensitive to the regression caused by the second update;

wherein the rollout service places the userbase segments that are sensitive to the regression caused by the second update in later stages of the update and userbase segments that are not sensitive to the regression caused by the second update in earlier stages of the second update.

10 . The orchestration system of claim 1 , further comprising a data-driven temperature monitor for different segments of a userbase, the rollout service to implement a data-driven temperature-based rollout by determining a staged deployment policy for the second update based on a temperature of the different segments of the userbase.

11 . The orchestration system of claim 1 , wherein, in response to detection of the regression caused by the second update, the rollout service is to provide information about the regression caused by the second update and a matching change in the second update that caused the regression caused by the second update to a number of administrators of client systems of the cloud service.

12 . The orchestration system of claim 1 , wherein a planned fault that will cause a regression is periodically placed in an update being implemented by the rollout service to test response of the rollout service to a resulting planned regression.

13 . The orchestration system of claim 1 , wherein

the first update has a corresponding update type indicating which attributes of the cloud service are affected by the first update,

the frequency for checking the health metric of the first state of the cloud service is determined based on a risk of the update type.

14 . The orchestration system of claim 1 , wherein

the first update has a corresponding update type indicating which attributes of the cloud service are affected by the first update,

the deployment policy includes a selection of Quality of Service (QOS) signals from components of the cloud service, the QoS signals are included in the health metric of the first state of the cloud service, where the QoS signals are selected based on the update type.

15 . A method for implementing a rollout service that deploys a series of updates to a cloud service while minimizing an impact of a regression caused in the cloud service by one of the updates, the method comprising:

determining a deployment policy for a first update using a database of deployment policy information and records of previous updates to the cloud service;

determining, based on a risk of the first update, a frequency for checking a health metric of a first state of the cloud service, wherein the first state of the cloud service corresponds to the first update;

deploying the first update according to the deployment policy,

checking, at the determined frequency for checking the health metric, whether or not the health metric of the first state of the cloud service indicates that the regression in the cloud service has occurred;

in response to determining that the health metric indicates the regression in the cloud service has not occurred, tagging the first state as a known good state; and

in respond to identifying an occurrence of a regression caused by a second update to the cloud service subsequent to the first update, returning the cloud service to the first state tagged as the known good state.

16 . The method of claim 15 , wherein

the first update has a corresponding update type indicating which attributes of the cloud service are affected by the first update, and

the frequency for checking the health metric of the first state of the cloud service is determined based on a risk of the update type.

17 . The method of claim 15 , wherein

the first update has a corresponding update type indicating which attributes of the cloud service are affected by the first update, and

the deployment policy includes a selection of Quality of Service (QOS) signals from components of the cloud service, the QoS signals are included in the health metric of the first state of the cloud service, where the QoS signals are selected based on the update type.

18 . A machine-readable medium storing instructions for implementing a rollout service that deploys a series of updates to a cloud service while minimizing an impact of a regression caused in the cloud service by one of the updates, where the instructions, when executed by a processor, cause:

determining a deployment policy for a first update using a database of deployment policy information and records of previous updates to the cloud service;

determining, based on a risk of the first update, a frequency for checking a health metric of a first state of the cloud service, wherein the first state of the cloud service corresponds to the first update;

deploying the first update according to the deployment policy,

checking, at the determined frequency for checking the health metric, whether or not the health metric of the first state of the cloud service indicates that the regression in the cloud service has occurred;

in response to determining that the health metric indicates the regression in the cloud service has not occurred, tagging the first state as a known good state; and

in respond to identifying an occurrence of a regression caused by a second update to the cloud service subsequent to the first update, returning the cloud service to the first state tagged as the known good state.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: VERMA, NIDHI; NIGAM, RAHUL
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 063129/0439 →
Continuity (1)
Related Publication 20240303062A1 · Sep 12, 2024
References Cited (11)
US 11422790B1 · Zahn · 2022 [cited by examiner]
US 20180373521A1 · Huang et al. · 2018 [cited by applicant]
US 20220050674A1 · Liljeback · 2022 [cited by applicant]
US 20220350509A1 · Verma et al. · 2022 [cited by applicant]
US 20240069886A1 · Verma · 2024 [cited by examiner]
DE 102018113625A1 · 2019 [cited by examiner]
Translated DE 102018113625 A1 (Year: 2019). [cited by examiner]
Euna Kim et al.; Automatic State Saving and Rollback in ns-3; ACM; pp. 263-266; retrieved on Jul. 11, 2025 (Year: 2017). [cited by examiner]
Demarne, et al., “Reliability Analytics for Cloud Based Distributed Databases”, Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data, Jun. 11, 2020, pp. 1479-1492. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/019330, Jun. 19, 2024, 20 pages. [cited by applicant]
International Preliminary Report on Patentability (Chapter I) for PCT Application No. PCT/US2024/019330, Oct. 9, 2025, 14 pages. [cited by applicant]