IP Library Granted Patent US 12688478
Granted Patent B2
US 12688478 · App. 17/753,477 · Granted Jul 21, 2026

Method and system for identification and analysis of regime shift

Inventors: Rajan Kumar (Pune, IN); Vivek Kumar (Pune, IN); Manendra Singh Parihar (Pune, IN); Venkataramana Runkana (Pune, IN)
Assignee: TATA CONSULTANCY SERVICES LIMITED
G06Q10/06393
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12688478
App. No.
17/753,477
Granted
Jul 21, 2026
Kind
B2
Abstract

This disclosure relates generally to identification and analysis of regime shift. The identification and analysis of the regime shift includes regime shift identification (RSI), root cause analysis of the identified regime shift and a recommendation unit to rectify the identified regime shift. The disclosure proposes to monitor a system continuously to identify a regime shift at real-time as presence of regime shifts in any system decreases quality of process and products and makes the system less efficient. The regime shift is identified at real-time based on key performance indicators (KPIs), a set of relevant features and real time input data using machine learning techniques. Further the disclosure also proposes techniques for detecting at least one root cause for the identified regime shift and also recommends a rectification action to rectify the identified regime shift based on optimization techniques.

Claims (309)

1 . A processor-implemented method for identification and analysis of a regime shift, the method comprising:

receiving and pre-processing a plurality of input data, associated with an industrial plant, from one or more sources, wherein the one or more sources refers to an industry plant unit,

wherein the step of pre-processing includes performing iterations for pre-processing the plurality of input data associated with the industrial plant, wherein each iteration comprises removing outliers from the input data using a multi-level outlier model to obtain a filtered data and the filtered data is categorized into multiple categories to identify missing data based on a frequency of occurrence of a plurality of parameters, wherein the missing data is imputed based on the multiple categories to obtain imputed data which is clustered into plurality of data clusters based on a predefined criteria, wherein after every iteration, determining whether the imputed data associated with a current iteration is clustered into the same data clusters as associated with a previous iteration and a plurality of iterations are performed until the data clusters in the previous iteration and the current iteration are similar to obtain a pre-processed plurality of input data;

identifying a plurality of key performance indicators (KPIs) and a set of relevant features from the pre-processed plurality of input data using an exhaustive domain knowledge and feature selection techniques, wherein the feature selection techniques include correlation techniques, statistics and machine learning techniques followed by ranking and consolidation, wherein the correlation techniques are performed by calculating a correlation co-efficient and ranking the set of relevant features based on a calculated correlation value, wherein the correlation technique is expressed as:

r

xy

=

(

x

i

-

x

_

)

(

y

i

-

y

_

)

(

x

i

-

x

_

)

2

(

y

i

-

y

_

)

2

where

r xy is the correlation co-efficient,

wherein x i is a feature and y i is a KPI at a time interval t,

wherein the machine learning techniques are performed by building a model and ranking features based on the built model, wherein ranking the feature based on a correlation of the model and a rank of the feature is consolidated based on a dynamic average ranking parameter computed based on the rank and a frequency of the feature, which is expressed as:

score

j

=

k

r

j

,

k

f

j

where r j,k is the rank of the feature j using the machine learning technique k, and f j is the frequency of the feature to be selected amongst k machine learning techniques;

creating a regime shift identification (RSI) model using the identified KPIs and the identified set of features based on machine learning techniques;

receiving and pre-processing a plurality of real-time input data from one or more sources;

continuously monitoring the industrial plant for identifying the regime shift at real-time using the created RSI model and the plurality of real time input data based on a domain knowledge, wherein the regime shift is identified based on a statistical or a machine learning technique based on a nature of a source determined by the domain knowledge, wherein the statistical techniques include hypothesis techniques and the machine learning techniques include auto encoder techniques, wherein steps for identifying regime shift based on hypothesis technique include identifying full historic data or window of data and short series of regime shift length based on the exhaustive domain knowledge, wherein a hypothesis test is performed between the identified historic data and the short series of regime shift length to identify the regime shift,

wherein a responsible parameter is identified at real-time based on the model using techniques including reconstruction error and mean hypothesis techniques, wherein the responsible parameter is identified while performing the reconstruction error and mean hypothesis techniques, wherein the KPI that is active during identification of RSI is identified as the responsible parameter, wherein active refers to the KPI that is shifting respective regime beyond the threshold computed with a mean and a standard deviation of the reconstruction error,

wherein the threshold for the KPI and the responsible parameter is dynamically defined based on the reconstruction error of an auto-encoder,

wherein the regime shift is identified based on a comparison of the real-time input data with the threshold;

detecting a plurality of root cause for the identified regime shift using the identified responsible parameter; and

performing a rectification action to rectify the identified regime shift, by performing optimization techniques or a pre-defined rule engine based technique, the optimization technique is performed by dynamically defining an objective function and a constraint function based on the domain knowledge and identified root cause analysis (RCA), or the pre-defined rule engine based technique is performed by computing a mean value of the identified responsible parameter, and generating an ideal value based on the domain knowledge and the computed mean value to change a value of the responsible parameter based on the ideal value to mitigate the identified regime shift as the rectification action, wherein the root cause analysis (RCA) is performed individually for measured KPI and derived KPI, wherein the RCA for the measured KPI is identified based on identification of the responsible parameter and the derived KPI is identified based on the responsible parameter and RCA consolidation techniques,

wherein the objective function is defined based on the identified KPI, which is expressed as:

O=(KPI−KPI′) where O is the objective function and KPI is the key performance indicators, KPI′ is the base value of key performance indicators,

wherein the constraint function is defined based on a pre-defined upper bound and lower bound of manipulated responsible parameters and process responsible parameters to maintain a pre-defined operating point, which is expressed as follows:

Cm

nL

<

Cm

n

<

Cm

nU

where

Cm nL is the lower bound of the manipulated responsible parameter,

Cm n is the constraint function,

Cm nU is the upper bound of the manipulated responsible parameter,

Cp

nL

<

Cp

n

<

Cp

nU

where

Cp nL is the lower bound of the process responsible parameter,

Cp n is the constraint function,

Cp nU is the upper bound of the process responsible parameter;

dynamically updating a domain knowledge database with the exhaustive domain knowledge of the one or more sources including the industry plant unit comprising a plurality of operating units including Enterprise Resource Planning (ERP), Distributed Control System (DCS), Laboratory information management system (LIMS) which sends multivariate data that comprises raw material quality, composition and feed rate, process parameters, condition of equipment, environmental emission parameters, product quality and production amount, wherein the domain knowledge is received from the domain knowledge database, and wherein the domain knowledge database configured for sharing the dynamically updated domain knowledge with a system.

2 . The method of claim 1 , wherein identification and analysis of the regime shift includes regime shift identification (RSI), root cause analysis of the identified regime shift and rectifies the identified regime shift.

3 . The method of claim 1 , wherein the identified regime shift, the root cause and the rectification action is displayed on a display module.

4 . The method of claim 1 , the plurality of KPIs identified from the pre-processed input data include measured KPIs and derived KPI, wherein the measured KPI and derived KPI is identified based on the domain knowledge received from the domain knowledge database.

5 . The method of claim 1 , wherein the RCA consolidation techniques for identifying RCA in derived KPI includes assigning weightage and ranking based on the plurality of domain knowledge, wherein assigning the weightage includes assigning a pre-defined weightage to each KPI and ranking the responsible parameter based on a weighted score, represented as:

WS

j

=

i

w

i

s

ji

i

s

ji

where

WS j is the weighted score

W i is the weightage of i th KPI and

S ji is an importance score of the responsible parameter j in KPI i.

6 . A system for identification and RCA of a regime shift comprising:

a memory for storing instructions;

one or more communication interfaces;

one or more hardware processors communicatively coupled to the memory using the one or more communication interfaces, wherein the one or more hardware processors are configured by the instructions for identification and root cause analysis of a regime shift, the system is configured for:

receiving a plurality of input data and a plurality of real-time input data, associated with an industrial plant, from one or more sources, wherein the one or more sources refers to an industry plant unit;

pre-processing the received plurality of input data and the plurality of real-time input data, wherein the pre-processing includes performing iterations for pre-processing the plurality of input data associated with the industrial plant, wherein each iteration comprises removing outliers from the input data using a multi-level outlier model to obtain a filtered data and the filtered data is categorized into multiple categories to identify missing data based on a frequency of occurrence of a plurality of parameters, wherein the missing data is imputed based on the multiple categories to obtain imputed data which is clustered into plurality of data clusters based on a predefined criteria, wherein after every iteration, determining whether the imputed data associated with a current iteration is clustered into the same data clusters as associated with a previous iteration and a plurality of iterations are performed until the data clusters in the previous iteration and the current iteration are similar to obtain a pre-processed plurality of input data;

a domain knowledge database configured for sharing dynamically updated domain knowledge with the system, wherein the domain knowledge database is dynamically updated with the exhaustive domain knowledge of the one or more sources including the industry plant unit comprising a plurality of operating units including Enterprise Resource Planning (ERP), Distributed Control System (DCS), Laboratory information management system (LIMS) which sends multivariate data that comprises raw material quality, composition and feed rate, process parameters, condition of equipment, environmental emission parameters, product quality and production amount, wherein the domain knowledge is received from the domain knowledge database;

continuously monitoring the industrial plant for identifying a plurality of key performance indicators (KPI) from the pre-processed plurality of input data using an exhaustive domain knowledge;

selecting a set of features from the identified plurality of KPIs based on feature selection techniques and domain knowledge, wherein the feature selection techniques include correlation techniques, statistics and machine learning techniques followed by ranking and consolidation, wherein the correlation techniques are performed by calculating a correlation co-efficient and ranking the set of relevant features based on a calculated correlation value, wherein the correlation technique is expressed as:

r

xy

=

(

x

i

-

x

_

)

(

y

i

-

y

_

)

(

x

i

-

x

_

)

2

(

y

i

-

y

_

)

2

where

r xy is the correlation co-efficient,

wherein x i is a feature and y i is a KPI at a time interval t,

wherein the machine learning techniques are performed by building a model and ranking features based on the built model, wherein ranking the feature based on a correlation of the model and a rank of the feature is consolidated based on a dynamic average ranking parameter computed based on the rank and a frequency of the feature, which is expressed as:

score

j

=

k

r

j

,

k

f

j

where r j,k is the rank of the feature j using the machine learning technique k, and f j is the frequency of the feature to be selected amongst k machine learning techniques; and

creating a model using the selected set of features using machine learning techniques;

identifying the regime shift at real-time using the created model and the plurality of real time input data based on a domain knowledge, wherein the regime shift is identified based on a statistical or a machine learning technique based on a nature of a source determined by the domain knowledge, wherein the statistical techniques include hypothesis techniques and the machine learning techniques include auto encoder techniques, wherein steps for identifying regime shift based on hypothesis technique include identifying full historic data or window of data and short series of regime shift length based on the exhaustive domain knowledge,

wherein a hypothesis test is performed between the identified historic data and the short series of regime shift length to identify the regime shift,

wherein a responsible parameter is identified at real-time based on the model using techniques including reconstruction error and mean hypothesis techniques, wherein the responsible parameter is identified while performing the reconstruction error and mean hypothesis techniques, wherein the KPI that is active during identification of RSI is identified as the responsible parameter, wherein active refers to the KPI that is shifting respective regime beyond the threshold computed with a mean and a standard deviation of the reconstruction error,

wherein the threshold for the KPI and the responsible parameter is dynamically defined based on the reconstruction error of an auto-encoder,

wherein the regime shift is identified based on a comparison of the real-time input data with the threshold;

detecting a plurality of root cause for the identified regime shift using the identified responsible parameter; and

performing a rectification action to rectify the identified regime shift by performing optimization techniques or a pre-defined rule engine based technique, the optimization technique is performed by dynamically defining an objective function and a constraint function based on the domain knowledge and identified root cause analysis (RCA), or the pre-defined rule engine based technique is performed by computing a mean value of the identified responsible parameter, and generating an ideal value based on the domain knowledge and the computed mean value to change a value of the responsible parameter based on the ideal value to mitigate the identified regime shift as the rectification action, wherein the root cause analysis (RCA) is performed individually for measured KPI and derived KPI, wherein the RCA for the measured KPI is identified based on identification of the responsible parameter and the derived KPI is identified based on the responsible parameter and RCA consolidation techniques,

wherein the objective function is defined based on the identified KPI, which is expressed as:

O=(KPI−KPI′) where O is the objective function and KPI is the key performance indicators, KPI′ is the base value of key performance indicators,

wherein the constraint function is defined based on a pre-defined upper bound and lower bound of manipulated responsible parameters and process responsible parameters to maintain a pre-defined operating point, which is expressed as follows:

Cm

nL

<

Cm

n

<

Cm

nU

where

Cm nL is the lower bound of the manipulated responsible parameter,

Cm n is the constraint function,

Cm nU is the upper bound of the manipulated responsible parameter,

Cp

nL

<

Cp

n

<

Cp

nU

where

Cp nL is the lower bound of the process responsible parameter,

Cp n is the constraint function,

Cp nU is the upper bound of the process responsible parameter; and

a display module configured for displaying the identified regime shift, the root cause and the rectification action.

7 . A non-transitory computer-readable medium having embodied thereon a computer readable program for identification and root cause analysis of a regime shift wherein the computer readable program, when executed by one or more hardware processors, cause:

receiving and pre-processing a plurality of input data, associated with an industrial plant, from one or more sources, wherein the one or more sources refers to an industry plant unit,

wherein the step of pre-processing includes performing iterations for pre-processing the plurality of input data associated with the industrial plant, wherein each iteration comprises removing outliers from the input data using a multi-level outlier model to obtain a filtered data and the filtered data is categorized into multiple categories to identify missing data based on a frequency of occurrence of a plurality of parameters, wherein the missing data is imputed based on the multiple categories to obtain imputed data which is clustered into plurality of data clusters based on a predefined criteria, wherein after every iteration, determining whether the imputed data associated with a current iteration is clustered into the same data clusters as associated with a previous iteration and a plurality of iterations are performed until the data clusters in the previous iteration and the current iteration are similar to obtain a pre-processed plurality of input data;

identifying a plurality of key performance indicators (KPIs) and a set of relevant features from the pre-processed plurality of input data using an exhaustive domain knowledge and feature selection techniques, wherein the feature selection techniques include correlation techniques, statistics and machine learning techniques followed by ranking and consolidation, wherein the correlation techniques are performed by calculating a correlation co-efficient and ranking the set of relevant features based on a calculated correlation value, wherein the correlation technique is expressed as:

r

xy

=

(

x

i

-

x

_

)

(

y

i

-

y

_

)

(

x

i

-

x

_

)

2

(

y

i

-

y

_

)

2

where

r xy is the correlation co-efficient,

wherein x i is a feature and y i is a KPI at a time interval t,

wherein the machine learning techniques are performed by building a model and ranking features based on the built model, wherein ranking the feature based on a correlation of the model and a rank of the feature is consolidated based on a dynamic average ranking parameter computed based on the rank and a frequency of the feature, which is expressed as:

score

j

=

k

r

j

,

k

f

j

where r j,k is the rank of the feature j using the machine learning technique k, and f j is the frequency of the feature to be selected amongst k machine learning techniques;

creating a regime shift identification (RSI) model using the identified KPIs and the identified set of features based on machine learning techniques;

receiving and pre-processing a plurality of real-time input data from one or more sources;

continuously monitoring the industrial plant for identifying the regime shift at real-time using the created RSI model and the plurality of real time input data based on a domain knowledge, wherein the regime shift is identified based on a statistical or a machine learning technique based on a nature of a source determined by the domain knowledge, wherein the statistical techniques include hypothesis techniques and the machine learning techniques include auto encoder techniques, wherein steps for identifying regime shift based on hypothesis technique include identifying full historic data or window of data and short series of regime shift length based on the exhaustive domain knowledge, wherein a hypothesis test is performed between the identified historic data and the short series of regime shift length to identify the regime shift,

wherein a responsible parameter is identified at real-time based on the model using a variety of techniques that include reconstruction error and mean hypothesis techniques, wherein the responsible parameter is identified while performing the reconstruction error and mean hypothesis techniques, wherein the KPI that is active during identification of RSI is identified as the responsible parameter, wherein active refers to the KPI that is shifting its regime beyond the threshold computed with mean and standard deviation of the reconstruction error,

wherein the threshold for the KPI and the responsible parameter is dynamically defined based on the reconstruction error of an auto-encoder,

wherein the regime shift is identified based on a comparison of the real-time input data with the threshold;

detecting a plurality of root cause for the identified regime shift using the identified responsible parameter; and

performing a rectification action to rectify the identified regime shift by performing optimization techniques or a pre-defined rule engine based technique, the optimization technique is performed by dynamically defining an objective function and a constraint function based on the domain knowledge and identified root cause analysis (RCA), or the pre-defined rule engine based technique is performed by computing a mean value of the identified responsible parameter, and generating an ideal value based on the domain knowledge and the computed mean value to change a value of the responsible parameter based on the ideal value to mitigate the identified regime shift as the rectification action, wherein the root cause analysis (RCA) is performed individually for measured KPI and derived KPI, wherein the RCA for the measured KPI is identified based on identification of the responsible parameter and the derived KPI is identified based on the responsible parameter and RCA consolidation techniques,

wherein the objective function is defined based on the identified KPI, which is expressed as:

O=(KPI−KPI′) where O is the objective function and KPI is the key performance indicators, KPI′ is the base value of key performance indicators,

wherein the constraint function is defined based on a pre-defined upper bound and lower bound of manipulated responsible parameters and process responsible parameters to maintain a pre-defined operating point, which is expressed as follows:

Cm

nL

<

Cm

n

<

Cm

nU

where

Cm nL is the lower bound of the manipulated responsible parameter,

Cm n is the constraint function,

Cm nU is the upper bound of the manipulated responsible parameter,

Cp

nL

<

Cp

n

<

Cp

nU

where

Cp nL is the lower bound of the process responsible parameter,

Cp n is the constraint function,

Cp nU is the upper bound of the process responsible parameter;

dynamically updating a domain knowledge database with the exhaustive domain knowledge of the one or more sources including the industry plant unit comprising a plurality of operating units including Enterprise Resource Planning (ERP), Distributed Control System (DCS), Laboratory information management system (LIMS) which sends multivariate data that comprises raw material quality, composition and feed rate, process parameters, condition of equipment, environmental emission parameters, product quality and production amount, wherein the domain knowledge is received from the domain knowledge database, and wherein the domain knowledge database configured for sharing the dynamically updated domain knowledge with a system.