Distributed dataset distillation for efficient bootstrapping of operational states classification models
One example method includes, at a node, installing a default parametrization configuration that facilitates performance of a domain task, obtaining, by the node, a distilled dataset, and obtaining the distilled dataset is either: obtaining the distilled dataset from another node; or leveraging a synthetic state assembled in the node to select the distilled dataset from another node based on state similarity of the node to the another node. The example method further includes training a model at the node, and the training is performed using the distilled dataset, and the trained model is operable to leverage information received by the node to propose changes to the parametrization configuration so as to optimize execution of a task by the node.
1 . A method, comprising:
at a node, installing a default parametrization configuration that facilitates performance of a domain task;
obtaining, by the node, a distilled dataset, and obtaining the distilled dataset comprises either:
obtaining the distilled dataset, which has been distilled from an original dataset by another node, has been used to train a model of the another node, and is much smaller than the original dataset, from the another node; or
leveraging a synthetic state assembled in the node to select the distilled dataset from the another node based on state similarity of the node to the another node;
training a model at the node by using the distilled dataset,
wherein the model is operable to leverage information received by the node to propose changes to the parametrization configuration so as to optimize execution of a task by the node.
2 . The method as recited in claim 1 , wherein the information received by the node comprise telemetry and provenance data associated with the node.
3 . The method as recited in claim 1 , wherein the synthetic state comprises input data received by the node, and baseline provenance and telemetry data.
4 . The method as recited in claim 1 , wherein the distilled dataset is received at the node in response to a broadcast from the node.
5 . The method as recited in claim 1 , wherein the distilled dataset is selected from another node based on state similarity information received from the another node by the node.
6 . The method as recited in claim 1 , wherein the node is a newly deployed node.
7 . A non-transitory computer readable storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations comprising:
at a node, installing a default parametrization configuration that facilitates performance of a domain task;
obtaining, by the node, a distilled dataset, which has been distilled from an original dataset by another node, has been used to train a model of the another node, and is much smaller than the original dataset, from the another node, wherein obtaining the distilled dataset comprises either:
obtaining the distilled dataset from the another node; or
leveraging a synthetic state assembled in the node to select the distilled dataset from another node based on state similarity of the node to the another node; and
training a model at the node by using the distilled dataset,
wherein the model is operable to leverage information received by the node to propose changes to the parametrization configuration so as to optimize execution of a task by the node.
8 . The non-transitory computer readable storage medium as recited in claim 7 , wherein the information received by the node comprise telemetry and provenance data associated with the node.
9 . The non-transitory computer readable storage medium as recited in claim 7 , wherein the synthetic state comprises input data received by the node, and baseline provenance and telemetry data.
10 . The non-transitory computer readable storage medium as recited in claim 7 , wherein the distilled dataset is received at the node in response to a broadcast from the node.
11 . The non-transitory computer readable storage medium as recited in claim 7 , wherein the distilled dataset is selected from another node based on state similarity information received from the another node by the node.
12 . The non-transitory computer readable storage medium as recited in claim 7 , wherein the node is a newly deployed node.
13 . A method, comprising:
generating, by a node, a distilled dataset, which has been distilled from an original dataset by the node, has been used to train a model of the node, and is much smaller than the original dataset;
determining, by the node, that the node has adequate resources to perform a task;
training a model at the node so as to produce a trained model by using the distilled dataset; and
transferring the trained model to a newly deployed node,
wherein the trained model is operable to leverage information received by the newly deployed node to propose changes to a default parametrization configuration so as to optimize execution of a task by the newly deployed node.
14 . The method as recited in claim 13 , wherein when the newly deployed node has adequate local data, the newly deployed node is able to operate the trained model.
15 . The method as recited in claim 14 , wherein the local data comprises input data, telemetry, and task output.
16 . The method as recited in claim 13 , wherein the model is trained by the node independently of a setup process performed at the newly deployed node.
17 . The method as recited in claim 13 , wherein the trained model is in a compressed state when it is sent from the node to the newly deployed node.
18 . The method as recited in claim 13 , wherein the model is trained before deployment of the newly deployed node.
19 . The method as recited in claim 13 , wherein the distilled dataset comprises telemetry and provenance information about the node.
20 . The method as recited in claim 13 , wherein the trained model is usable at the newly deployed node for transfer learning, data classification, and tuning of data received by the newly deployed node.