IP Library Granted Patent US 12700877
Granted Patent B2
US 12700877 · App. 18/897,124 · Granted Aug 4, 2026

Optimizing lossy compression for nonhomogeneous multivariate black-box classification models

Inventors: Rômulo Teixeira De Abreu Pinho (Niterói, BR); Vinicius Michel Gottin (Rio de Janeiro, BR); Paulo de Figueiredo Pires (Niterói, BR); Alex Laier Bordignon (Niterói, BR); João Victor Daher Daibes (Niterói, BR); Franklin Jordan Ventura Quico (Niterói, BR)
Assignee: Dell Products L.P.
H03M7/6041H03M7/6005H03M7/6011
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12700877
App. No.
18/897,124
Granted
Aug 4, 2026
Kind
B2
Abstract

A service compresses a data set in accordance with a data compression rate, resulting in generation of compressed data having a first data quality. The service decompresses the compressed data, resulting in generation of decompressed data. The service filters, from the decompressed data, data that is identified as belonging to a selected class of data. The service causes an ML classifier to perform a classification operation on the data that is identified as belonging to the selected class of data. The service determines a classification accuracy of the classification operation performed by the ML classifier. The service determines whether the classification accuracy at least meets a threshold accuracy requirement. If the threshold is met, the data compression rate is increased; otherwise, it is decreased.

Claims (61)

1 . A method comprising:

compressing a data set in accordance with a data compression rate, resulting in generation of compressed data having a first data quality;

decompressing the compressed data, resulting in generation of decompressed data;

filtering, from the decompressed data, data that is identified as belonging to a selected class of data;

causing a machine learning (ML) classifier to perform a classification operation on the data that is identified as belonging to the selected class of data;

determining a classification accuracy of the classification operation performed by the ML classifier;

determining whether the classification accuracy at least meets a threshold accuracy requirement;

in response to determining that the classification accuracy at least meets the threshold accuracy requirement, increasing the data compression rate, resulting in generation of an increased data compression rate, wherein use of the increased data compression rate causes a subsequent reduction in data quality for data that is compressed at the increased data compression rate as compared to the first data quality; and

in response to determining that the classification accuracy does not at least meet the threshold accuracy requirement, decreasing the data compression rate, resulting in generation of a decreased data compression rate, wherein use of the decreased data compression rate causes a subsequent increase in data quality for data that is compressed at the decreased data compression rate as compared to the first data quality.

2 . The method of claim 1 , wherein compressing the data set and subsequently decompressing the compressed data is performed to mimic a data transmission between an edge device and a cloud node.

3 . The method of claim 1 , wherein determining the classification accuracy is performed by:

accessing a first result of the ML classifier, the first result being generated in response to the ML classifier performing the classification operation on the data that is identified as belonging to the selected class of data;

filtering, from the data set, second data that is identified as belonging to the selected class of data;

causing the ML classifier to perform the classification operation on the second data, which has not been subjected to compression or decompression, wherein the ML classifier generates a second result in response to performing the classification operation on the second data; and

comparing the first result with the second result to identify a difference; and

determining the classification accuracy based on the difference.

4 . The method of claim 1 , wherein said method is iteratively performed until a convergence data compression rate is identified.

5 . The method of claim 1 , wherein random variations are injected into the data set.

6 . The method of claim 1 , wherein the data set is included in a larger data set that has been divided to produce a validation data set and a training data set, and wherein said data set is the training data set.

7 . The method of claim 1 , wherein a stability evaluation is performed prior to compressing the data set, and wherein the stability evaluation involves identification of an instability pattern of a data compressor, which performs said compressing, and the ML classifier.

8 . The method of claim 1 , wherein determining the classification accuracy of the classification operation is based on a relation between (i) a first accuracy obtained when the ML classifier operates on the data that is identified as belonging to the selected class of data and (ii) a second accuracy obtained when the ML classifier operates on uncompressed data obtained from the data set, where the uncompressed data is also data that is identified as belonging to the selected class of data.

9 . A computer system comprising:

one or more processors; and

one or more hardware storage devices that store instructions that are executable by the one or more processors to cause the computer system to:

compress a data set in accordance with a data compression rate, resulting in generation of compressed data having a first data quality;

decompress the compressed data, resulting in generation of decompressed data;

filter, from the decompressed data, data that is identified as belonging to a selected class of data;

cause a machine learning (ML) classifier to perform a classification operation on the data that is identified as belonging to the selected class of data;

determine a classification accuracy of the classification operation performed by the ML classifier;

determine whether the classification accuracy at least meets a threshold accuracy requirement;

in response to determining that the classification accuracy at least meets the threshold accuracy requirement, increase the data compression rate, resulting in generation of an increased data compression rate, wherein use of the increased data compression rate causes a subsequent reduction in data quality for data that is compressed at the increased data compression rate as compared to the first data quality; and

in response to determining that the classification accuracy does not at least meet the threshold accuracy requirement, decrease the data compression rate, resulting in generation of a decreased data compression rate, wherein use of the decreased data compression rate causes a subsequent increase in data quality for data that is compressed at the decreased data compression rate as compared to the first data quality.

10 . The computer system of claim 9 , wherein compressing the data set and subsequently decompressing the compressed data is performed to mimic a data transmission between an edge device and a cloud node.

11 . The computer system of claim 9 , wherein determining the classification accuracy is performed by:

accessing a first result of the ML classifier, the first result being generated in response to the ML classifier performing the classification operation on the data that is identified as belonging to the selected class of data;

filtering, from the data set, second data that is identified as belonging to the selected class of data;

causing the ML classifier to perform the classification operation on the second data, which has not been subjected to compression or decompression, wherein the ML classifier generates a second result in response to performing the classification operation on the second data; and

comparing the first result with the second result to identify a difference; and

determining the classification accuracy based on the difference.

12 . The computer system of claim 9 , wherein random variations are injected into the data set.

13 . The computer system of claim 9 , wherein determining the classification accuracy of the classification operation is based on a relation between (i) a first accuracy obtained when the ML classifier operates on the data that is identified as belonging to the selected class of data and (ii) a second accuracy obtained when the ML classifier operates on uncompressed data obtained from the data set, where the uncompressed data is also data that is identified as belonging to the selected class of data.

14 . One or more hardware storage devices that store instructions that are executable by one or more processors to cause the one or more processors to perform operations comprising:

compressing a data set in accordance with a data compression rate, resulting in generation of compressed data having a first data quality;

decompressing the compressed data, resulting in generation of decompressed data;

filtering, from the decompressed data, data that is identified as belonging to a selected class of data;

causing a machine learning (ML) classifier to perform a classification operation on the data that is identified as belonging to the selected class of data;

determining a classification accuracy of the classification operation performed by the ML classifier;

determining whether the classification accuracy at least meets a threshold accuracy requirement;

in response to determining that the classification accuracy at least meets the threshold accuracy requirement, increasing the data compression rate, resulting in generation of an increased data compression rate, wherein use of the increased data compression rate causes a subsequent reduction in data quality for data that is compressed at the increased data compression rate as compared to the first data quality; and

in response to determining that the classification accuracy does not at least meet the threshold accuracy requirement, decreasing the data compression rate, resulting in generation of a decreased data compression rate, wherein use of the decreased data compression rate causes a subsequent increase in data quality for data that is compressed at the decreased data compression rate as compared to the first data quality.

15 . The one or more hardware storage devices of claim 14 , wherein determining the classification accuracy is performed by:

accessing a first result of the ML classifier, the first result being generated in response to the ML classifier performing the classification operation on the data that is identified as belonging to the selected class of data;

filtering, from the data set, second data that is identified as belonging to the selected class of data;

causing the ML classifier to perform the classification operation on the second data, which has not been subjected to compression or decompression, wherein the ML classifier generates a second result in response to performing the classification operation on the second data; and

comparing the first result with the second result to identify a difference; and

determining the classification accuracy based on the difference.

16 . The one or more hardware storage devices of claim 14 , wherein random variations are injected into the data set.

17 . The one or more hardware storage devices of claim 14 , wherein the data set is included in a larger data set that has been divided to produce a validation data set and a training data set, and wherein said data set is the training data set.

18 . The one or more hardware storage devices of claim 14 , wherein a stability evaluation is performed prior to compressing the data set, and wherein the stability evaluation involves identification of an instability pattern of a data compressor, which performs said compressing, and the ML classifier.

19 . The one or more hardware storage devices of claim 14 , wherein determining the classification accuracy of the classification operation is based on a relation between (i) a first accuracy obtained when the ML classifier operates on the data that is identified as belonging to the selected class of data and (ii) a second accuracy obtained when the ML classifier operates on uncompressed data obtained from the data set, where the uncompressed data is also data that is identified as belonging to the selected class of data.

20 . The one or more hardware storage devices of claim 14 , wherein compressing the data set and subsequently decompressing the compressed data is performed to mimic a data transmission between an edge device and a cloud node.