IP Library Granted Patent US 12,645,985
Granted Patent B2
US 12,645,985 · App. 17/447,980 · Granted Jun 2, 2026

Distributed dataset distillation for efficient bootstrapping of operational states classification models

Inventors: Paulo Abelha Ferreira (Rio de Janeiro, BR); Vinicius Michel Gottin (Rio de Janeiro, BR)
Assignee: EMC IP Holding Company LLC
G06N20/00G06F9/4401
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,985
App. No.
17/447,980
Granted
Jun 2, 2026
Kind
B2
Abstract

One example method includes, at a node, installing a default parametrization configuration that facilitates performance of a domain task, obtaining, by the node, a distilled dataset, and obtaining the distilled dataset is either: obtaining the distilled dataset from another node; or leveraging a synthetic state assembled in the node to select the distilled dataset from another node based on state similarity of the node to the another node. The example method further includes training a model at the node, and the training is performed using the distilled dataset, and the trained model is operable to leverage information received by the node to propose changes to the parametrization configuration so as to optimize execution of a task by the node.

Claims (37)

1 . A method, comprising:

at a node, installing a default parametrization configuration that facilitates performance of a domain task;

obtaining, by the node, a distilled dataset, and obtaining the distilled dataset comprises either:

obtaining the distilled dataset, which has been distilled from an original dataset by another node, has been used to train a model of the another node, and is much smaller than the original dataset, from the another node; or

leveraging a synthetic state assembled in the node to select the distilled dataset from the another node based on state similarity of the node to the another node;

training a model at the node by using the distilled dataset,

wherein the model is operable to leverage information received by the node to propose changes to the parametrization configuration so as to optimize execution of a task by the node.

2 . The method as recited in claim 1 , wherein the information received by the node comprise telemetry and provenance data associated with the node.

3 . The method as recited in claim 1 , wherein the synthetic state comprises input data received by the node, and baseline provenance and telemetry data.

4 . The method as recited in claim 1 , wherein the distilled dataset is received at the node in response to a broadcast from the node.

5 . The method as recited in claim 1 , wherein the distilled dataset is selected from another node based on state similarity information received from the another node by the node.

6 . The method as recited in claim 1 , wherein the node is a newly deployed node.

7 . A non-transitory computer readable storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations comprising:

at a node, installing a default parametrization configuration that facilitates performance of a domain task;

obtaining, by the node, a distilled dataset, which has been distilled from an original dataset by another node, has been used to train a model of the another node, and is much smaller than the original dataset, from the another node, wherein obtaining the distilled dataset comprises either:

obtaining the distilled dataset from the another node; or

leveraging a synthetic state assembled in the node to select the distilled dataset from another node based on state similarity of the node to the another node; and

training a model at the node by using the distilled dataset,

wherein the model is operable to leverage information received by the node to propose changes to the parametrization configuration so as to optimize execution of a task by the node.

8 . The non-transitory computer readable storage medium as recited in claim 7 , wherein the information received by the node comprise telemetry and provenance data associated with the node.

9 . The non-transitory computer readable storage medium as recited in claim 7 , wherein the synthetic state comprises input data received by the node, and baseline provenance and telemetry data.

10 . The non-transitory computer readable storage medium as recited in claim 7 , wherein the distilled dataset is received at the node in response to a broadcast from the node.

11 . The non-transitory computer readable storage medium as recited in claim 7 , wherein the distilled dataset is selected from another node based on state similarity information received from the another node by the node.

12 . The non-transitory computer readable storage medium as recited in claim 7 , wherein the node is a newly deployed node.

13 . A method, comprising:

generating, by a node, a distilled dataset, which has been distilled from an original dataset by the node, has been used to train a model of the node, and is much smaller than the original dataset;

determining, by the node, that the node has adequate resources to perform a task;

training a model at the node so as to produce a trained model by using the distilled dataset; and

transferring the trained model to a newly deployed node,

wherein the trained model is operable to leverage information received by the newly deployed node to propose changes to a default parametrization configuration so as to optimize execution of a task by the newly deployed node.

14 . The method as recited in claim 13 , wherein when the newly deployed node has adequate local data, the newly deployed node is able to operate the trained model.

15 . The method as recited in claim 14 , wherein the local data comprises input data, telemetry, and task output.

16 . The method as recited in claim 13 , wherein the model is trained by the node independently of a setup process performed at the newly deployed node.

17 . The method as recited in claim 13 , wherein the trained model is in a compressed state when it is sent from the node to the newly deployed node.

18 . The method as recited in claim 13 , wherein the model is trained before deployment of the newly deployed node.

19 . The method as recited in claim 13 , wherein the distilled dataset comprises telemetry and provenance information about the node.

20 . The method as recited in claim 13 , wherein the trained model is usable at the newly deployed node for transfer learning, data classification, and tuning of data received by the newly deployed node.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2023
From: FERREIRA, PAULO ABELHA; GOTTIN, VINICIUS MICHEL
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 063288/0051 →
Continuity (1)
Related Publication 20230103817A1 · Apr 6, 2023
References Cited (18)
US 11004025B1 · Gottin et al. · 2021 [cited by applicant]
US 11182321B2 · Pinho et al. · 2021 [cited by applicant]
US 11218500B2 · Soeder et al. · 2022 [cited by applicant]
US 11275987B1 · Gottin et al. · 2022 [cited by applicant]
US 11403525B2 · Gottin et al. · 2022 [cited by applicant]
US 11562223B2 · Gottin et al. · 2023 [cited by applicant]
US 11586919B2 · Zhu · 2023 [cited by examiner]
US 12061991B2 · Chen · 2024 [cited by examiner]
US 20150356461A1 · Vinyals · 2015 [cited by examiner]
US 20180293146A1 · Ushiki · 2018 [cited by examiner]
US 20200272526A1 · Bhole · 2020 [cited by examiner]
US 20220067527A1 · Xu · 2022 [cited by examiner]
US 20230004854A1 · Calmon et al. · 2023 [cited by applicant]
US 20230066249A1 · Ferreira et al. · 2023 [cited by applicant]
US 20230229919A1 · Kar · 2023 [cited by examiner]
AU 2020353380A1 · 2022 [cited by examiner]
T. Wang, J. Zhu, A. Torralba and A. Efros, “Dataset distillation,” arXiv, vol. preprint arXiv:1811.10959., 2018. [cited by applicant]
F. Zhuang, Z. Qi, K. Duan, D. Xi, Y. Zhu, H. Zhu, H. Xiong and Q. He, “A comprehensive survey on transfer learning,” Proceedings of IEEE, vol. 109, No. 1, pp. pp. 43-76., 2020. [cited by applicant]