Method and system for identifying relevant variables
The invention relates to a method for identifying variables relevant to a dataset, said variables being derived from a plurality of variables involved in processing the dataset, said method comprising: a step of generating a subset of variables from the plurality of variables, a step of assigning a quantization value to each variable of the generated subset of variables, a step of selecting a relevant variable, a further step of generating a new subset of variables when the quantitative value of the selected variable is below a predetermined threshold value.
1 . A method for identifying relevant variables in a dataset including a plurality of data subsets implemented by a computing device comprising a variable selection module, each of the plurality of data subsets corresponding to a variable so as to form a set of variables, said method comprising:
issuing a request, by said variable selection module of said computing device, to a high-performance computing infrastructure to retrieve said dataset,
wherein said computing device is a computing server integrated into a high-performance computer system comprising a plurality of resources including at least one of
a network disk characterized by inputs/outputs and reading/writing on said network disk,
a memory characterized by usage rate,
a network characterized by a bandwidth,
a processor characterized by a usage percentage or by an occupancy rate of caches thereof, a random access memory characterized by a quantity allocated,
each associated with an industrial production sensor or computing probe,
wherein the dataset is implemented within a learning model used to monitor an industrial process that produces consumer goods, said industrial process selected from an agri-food production process, a manufacturing production process, a chemical synthesis process, a packaging process, or an IT infrastructure monitoring process,
wherein all of each variable of the set of variables of said each of the plurality of data subsets describe a current operating state of the high-performance computing infrastructure used to detect a future anomaly that leads to a disruption of services provided by the high-performance computing infrastructure,
wherein said each variable is a statistical unit that is observed and for which a numerical value is assigned,
wherein said each variable is associated with a resource of said plurality of resources of said high-performance computing infrastructure,
wherein said each variable allows a behavior of the computing device to be interpreted,
wherein each resource of said plurality of resources is characterized by a performance indicator,
wherein said plurality of resources comprise one or more of elements and functions characterized by parameters or capacities that allow operation of said high-performance computing infrastructure or of an application process within said high-performance computing infrastructure,
continuously receiving, in real time, said dataset and the set of variables from within said dataset, forming a representation space, as a function of said request,
generating a subset of variables from the set of variables as identified variables of said plurality of resources, by said variable selection module, by implementing a filter-type variable selection algorithm to train said learning model to monitor said industrial process, wherein variable selection is performed as a search problem over states in the representation space, wherein each state of said states specifies a possible subset of variables and wherein transitions are schematized by a partially ordered graph;
verifying relevance of the identified variables by assigning performance indicators to said identified variables;
comparing the performance indicators that are assigned with defined performance indicators;
removing variables from the subset of variables that is generated when the variables are redundant, to reduce a size of the representation space and thereby optimize consumption of hardware and software resources and reduce time to implement the method;
continuously calculating a performance value for said each variable of the subset of variables that is generated of said each resource, by the variable selection module,
classifying the variables based on the performance value that is associated therewith, wherein said performance value of said each variable
corresponds to a quantitative measurement of a technical or functional property of said plurality of resources of said high-performance computing infrastructure including one or more of CPU usage, memory usage, network bandwidth consumption, disk input/output rates, and
represents the current operating state of said high-performance computing infrastructure,
selecting at least one relevant variable from said identified variables of said subset of variables by applying a genetic algorithm, or by applying selection techniques, using said variable selection module via a finite sequence of operations or instructions,
wherein the selecting comprises
calculating a correlation between the variables of the subset,
encoding the variables according to the correlation that is calculated, and
recursively eliminating variables based on said encoding;
training the learning model using said at least one relevant variable to detect a future appearance of a breakdown or anomaly of the high-performance computing infrastructure,
wherein said selecting said at least one relevant variable allows for dynamic modification of said at least one relevant variable that is selected to generate said learning model,
such that said at least one relevant variable is constantly evaluated, adapted or replaced with a variable from said identified variables relevant to a particular industrial context to monitor said industrial process, including predictive maintenance, fraud detection, or cyber-attack detection, and continuously evaluating the performance value of the at least one relevant variable that is selected during monitoring of the industrial process to determine whether the performance value associated with the at least one relevant variable satisfies a predetermined threshold, and,
generating a new subset of variables, and
adjusting the learning model and operating the learning model, during continued monitoring of the industrial process, using only the new subset of variables as input variables when the performance value deteriorates, wherein the at least one relevant variable previously used is replaced by variables of the new subset of variables as input,
using results of the learning model operating with the new subset of variables to anticipate a future breakdown or resource requirement of the high-performance computing infrastructure; and
preventing a slowdown or shutdown of services provided by the high-performance computing infrastructure based on the future breakdown or resource requirement that is anticipated
to avoid obsolescence of the learning model as operating conditions of the high-performance computing infrastructure evolve,
to avoid failure of services provided by the high-performance computing infrastructure, and
to avoid degradation of the learning model;
wherein preventing the slowdown or shutdown comprises preventing a technical incident corresponding to a slowdown or shutdown of at least part of the high-performance computing infrastructure and its applications.
2 . The method according to claim 1 , wherein the selecting the at least one relevant variable by applying a genetic algorithm comprises
calculating a fit value for each variable of the subset of variables that is generated, generating a new variable comprising selecting a variable from a predefined selection function,
transforming the at least one relevant variable, that is selected, according to a predefined transformation function,
calculating a fit value for the new variable,
comparing the fit value of the new variable with all fit values of each variable of the subset of variables that is generated, comprising
generating a second new subset of variables when a quantitative value of the at least one relevant variable that is selected is less than a predetermined threshold value, said quantitative value corresponding to the fit value of the new variable and the predetermined threshold value corresponding to a lowest fit value of the variables in the subset that is generated.
3 . The method according to claim 1 , wherein the selecting the at least one relevant variable by said applying said selection techniques comprises
selecting a univariate variable, or
selecting a multivariate variable, or
selecting a variable by recursive elimination.
4 . The method according to claim 1 , wherein the industrial production sensor includes one or more of connected objects, machinesensors, environmental sensors, computing probes.
5 . A system that identifies variables relevant to a dataset, implemented by a non-transitory computing device, said variables being derived from a plurality of variables involved in processing the dataset, said system comprising:
a variable selection module within the non-transitory computing device,
wherein said non-transitory computing device is a computing server integrated into a high-performance computer system comprising a plurality of resources including at least one of
a network disk characterized by inputs/outputs and reading/writing on said network disk,
a memory characterized by usage rate,
a network characterized by a bandwidth,
a processor characterized by a usage percentage or by an occupancy rate of caches thereof,
a random access memory characterized by a quantity allocated,
each associated with an industrial production sensor or computing probe,
wherein the variable selection module is implemented by said non-transitory computing device that executes a finite sequence of operations configured to
issue a request to a high-performance computing infrastructure to retrieve said dataset,
continuously receive, in real time, said dataset and a set of variables from within said dataset, forming a representation space, as a function of said request issued by said non-transitory computing device,
wherein the dataset is implemented within a learning model used to monitor an industrial process that produces consumer goods,
said industrial process selected from an agri-food production process, a manufacturing production process, a chemical synthesis process, a packaging process, or an IT infrastructure monitoring process,
wherein all of each variable of the set of variables of said plurality of variables of said dataset describe a current operating state of the high-performance computing infrastructure used to detect a future anomaly that leads to a disruption of services provided by the high-performance computing infrastructure,
wherein said each variable is a statistical unit that is observed and for which a numerical value is assigned,
wherein said each variable is associated with a resource of said plurality of resources of said high-performance computing infrastructure,
wherein said each variable allows a behavior of the non-transitory computing device to be interpreted,
wherein each resource of said plurality of resources is characterized by a performance indicator,
wherein said plurality of resources comprise one or more of elements and functions characterized by parameters or capacities that allow operation of said high-performance computing infrastructure or of an application process within said high-performance computing infrastructure,
generate a subset of variables from the plurality of variables involved in processing the dataset as identified variables of said plurality of resources, by implementing a filter-type variable selection algorithm, to train said learning model to monitor said industrial process, wherein variable selection is performed as a search problem over states in the representation space, wherein each state of said states specifies a possible subset of variables and wherein transitions are schematized by a partially ordered graph,
verify relevance of the identified variables by assigning performance indicators to said identified variable,
compare the performance indicators that are assigned with defined performance indicators,
remove variables from the subset of variables that is generated when the variables are redundant, to reduce a size of the representation space and thereby optimize consumption of hardware and software resources and reduce time to implement the system,
continuously calculate and assign a quantitative value to said each variable of the subset of variables that is generated of said each resource,
classify the variables based on the quantitative value that is associated therewith, wherein said quantitative value
corresponds to a quantitative measurement of a technical or functional property of said plurality of resources of said high-performance computing infrastructure including one or more of CPU usage, memory usage, network bandwidth consumption, disk input/output rates, and
represents the current operating state of said high-performance computing infrastructure,
select a relevant variable from said identified variables of said subset of variables by applying a genetic algorithm, or by applying selection techniques, via said finite sequence of operations or instructions,
wherein the select the relevant variable comprises calculate a correlation between the variables of the subset,
encode the variables according to the correlation that is calculated, and
recursively eliminate variables based on said encode,
train said learning model using said relevant variable to detect a future appearance of a breakdown or anomaly of the high-performance computing infrastructure,
wherein said select said relevant variable allows for dynamic modification of said relevant variable that is selected to generate said learning model,
such that said relevant variable is constantly evaluated, adapted or replaced with a variable from said identified variables relevant to a particular industrial context to monitor said industrial process, including predictive maintenance, fraud detection, or cyber-attack detection, and
continuously evaluate the quantitative value of the at least one relevant variable that is selected during monitoring of the industrial process to determine whether the quantitative value associated with the at least one relevant variable satisfies a predetermined threshold, and,
generate a new subset of variables, and
adjust the learning model and operate the learning model, during continued monitoring of the industrial process, using only the new subset of variables as input variables when the quantitative value deteriorates, wherein the at least one relevant variable previously used is replaced by variables of the new subset of variables as input,
use results of the learning model operating with the new subset of variables to anticipate a future breakdown or resource requirement of the high-performance computing infrastructure; and
prevent a slowdown or shutdown of services provided by the high-performance computing infrastructure based on the future breakdown or resource requirement that is anticipated
to avoid obsolescence of the learning model as operating conditions of the high-performance computing infrastructure evolve,
to avoid failure of services provided by the high-performance computing infrastructure, and
to avoid degradation of the learning model.
6 . A variable selection module that identifies variables relevant to a dataset, implemented by a computing device comprising a computing server, said variables being derived from a plurality of variables involved in processing the dataset, said variable selection module comprising:
a processor configured to be integrated into a high performance computer system comprising a plurality of resources that executes a finite sequence of operations within said computing server configured to issue a request to a high-performance computing infrastructure to retrieve said dataset,
continuously receive, in real time, said dataset and the variables from within said dataset, forming a representation space, as a function of said request issued by said computing device,
wherein the dataset is implemented within a learning model used to monitor an industrial process that produces consumer goods, said industrial process selected from an agri-food production process, a manufacturing production process, a chemical synthesis process, a packaging process, or an IT infrastructure monitoring process,
wherein all of each variable of the variables of said plurality of variables of the dataset describe a current operating state of the high-performance computing infrastructure used to detect a future anomaly that leads to a disruption of services provided by the high-performance computing infrastructure,
wherein said each variable is a statistical unit that is observed and for which a numerical value is assigned,
wherein said each variable is associated with a resource of said plurality of resources of said high-performance computing infrastructure,
wherein said each variable allows a behavior of the computing device to be interpreted,
wherein each resource of said plurality of resources is characterized by a performance indicator,
wherein said plurality of resources comprise one or more of elements and functions characterized by parameters or capacities that allow operation of said high-performance computing infrastructure or of an application process within said high-performance computing infrastructure,
wherein said plurality of resources comprise at least one of
a network disk characterized by inputs/outputs and reading/writing on said network disk,
a memory characterized by usage rate,
a network characterized by a bandwidth,
a processor characterized by a usage percentage or by an occupancy rate of caches thereof,
a random access memory characterized by a quantity allocated,
each associated with an industrial production sensor or computing probe,
generate a subset of variables from the plurality of variables involved in processing the dataset as identified variables of said plurality of resources, by implementing a filter-type variable selection algorithm to train said learning model to monitor said industrial process, wherein variable selection is performed as a search problem over states in the representation space, wherein each state of said states specifies a possible subset of variables and wherein transitions are schematized by a partially ordered graph,
verify relevance of the identified variables by assigning performance indicators to said identified variables,
compare the performance indicators that are assigned with defined performance indicators,
remove variables from the subset of variables that is generated when the variables are redundant, to reduce a size of the representation space and thereby optimize consumption of hardware and software resources and reduce time to implement the variable selection module,
continuously calculate and assign a quantitative value to said each variable of the subset of variables that is generated of said each resource,
classify the variables based on the quantitative value that is associated therewith, wherein said quantitative value
corresponds to a measurement of a technical or functional property of said plurality of resources of said high-performance computing infrastructure including one or more of CPU usage, memory usage, network bandwidth consumption, disk input/output rates, and
represents the current operating state of said high-performance computing infrastructure,
select a relevant variable from said identified variables of said subset of variables by applying a genetic algorithm, or by applying selection techniques, via said finite sequence of operations or instructions,
wherein the select the relevant variable comprises calculate a correlation between the variables of the subset,
encode the variables according to the correlation that is calculated, and
recursively eliminate variables based on said encode,
train said learning model using said relevant variable to detect a future appearance of a breakdown or anomaly of the high-performance computing infrastructure,
wherein said select said relevant variable allows for dynamic modification of said relevant variable that is selected to generate said learning model,
such that said relevant variable is constantly evaluated, adapted or replaced with a variable from said identified variables relevant to a particular industrial context to monitor said industrial process, including predictive maintenance, fraud detection, or cyber-attack detection, and
continuously evaluate the quantitative value of the at least one relevant variable that is selected during monitoring of the industrial process to determine whether the quantitative value associated with the at least one relevant variable satisfies a predetermined threshold, and,
generate a new subset of variables, and
adjust the learning model and operate the learning model, during continued monitoring of the industrial process, using only the new subset of variables when the quantitative value deteriorates, wherein the at least one relevant variable previously used is replaced by variables of the new subset of variables as input,
use results of the learning model operating with the new subset of variables to anticipate a future breakdown or resource requirement of the high-performance computing infrastructure; and
prevent a slowdown or shutdown of services provided by the high-performance computing infrastructure based on the future breakdown or resource requirement that is anticipated to avoid obsolescence of the learning model as operating conditions of the high-performance computing infrastructure evolve,
to avoid failure of services provided by the high-performance computing infrastructure, and
to avoid degradation of the learning model.
7 . The method according to claim 1 , wherein the performance value comprises a performance indicator value corresponding to a measurement or calculation value of the technical or functional property of one or more elements of an IT infrastructure representing an operating state of the IT infrastructure.
8 . The method of claim 1 , wherein the defined performance indicators correspond to threshold values below which one or more variables are determined to be unsuitable for generating the learning model, and wherein generating the subset of variables is repeated when the one or more variables are determined to be unsuitable.
9 . The method of claim 1 , wherein verifying relevance of the identified variables includes evaluating a regression of the learning model using at least one performance indicator selected from a mean absolute error (MAE), a root mean square error (RMSE), and a coefficient of determination (R 2 ).
10 . The method of claim 1 , wherein calculating the correlation between the variables comprises calculating correlation using at least one of a Pearson correlation coefficient and a Spearman correlation coefficient.
11 . The method of claim 1 , wherein the partially ordered graph schematizes transitions between the states such that each transition corresponds to adding at least one variable to a current state to form a child state specifying a different subset of variables, wherein each child state has one more attribute than its parent state.