IP Library Granted Patent US 11,789,914
Granted Patent B2
US 11,789,914 · App. 16/878,713 · Granted Oct 17, 2023

Data correctness optimization

Inventors: Swapnasarit Sahu (Berlin, DE); Ernest Kirubakaran Selvaraj (Berlin, DE); Tushar Agarwal (Berlin, DE); Projjol Banerjea (Berlin, DE); Daniel Heer (Berlin, DE); Sathish Kumar K S (Berlin, DE)
Assignee: zeotap GmbH
G06F16/215G06F17/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,789,914
App. No.
16/878,713
Filed
May 20, 2020
Granted
Oct 17, 2023
Kind
B2
Examiner
VU, BAI DUC
Art Unit
2162
USPC
707/687
Abstract

Method and system for providing ground truth dataset and use thereof for improving data correctness. Datasets comprising data elements are received from different sources, each data element including an identifier and at least one attribute value associated therewith. Data correctness values are determined for the attribute values, each associated with a probability that an attribute value is correct. Data element with single data correctness value is added to the ground truth dataset for each attribute value for each identifier with which a respective attribute value is associated based on the determined data correctness values for the attribute values, whereby the data correctness values in the ground truth dataset define probability distributions of data correctness for the attribute values. Data correctness values for attribute values of data elements of a new dataset can be determined based on overlapping data compared in the ground truth dataset and the new dataset.

Claims (80)

1. A computer implemented method for improving data correctness of measured values using a ground truth dataset, comprising:

using a transceiver and at least one processor coupled to the transceiver for:

receiving a plurality of datasets from different data sources, the datasets comprising a plurality of data elements, wherein each data element includes an identifier and at least one measured value associated with the identifier;

determining data correctness values for the measured values using at least one of a panel and calibration measurements for checking correctness of the measured values, wherein a data correctness value is associated with a probability that a measured value is correct;

adding a data element including a single data correctness value for each measured value for each identifier with which a respective measured value is associated to the ground truth dataset, wherein the single data correctness value being based on the determined data correctness values for the measured values, whereby the data correctness values in the ground truth dataset define probability distributions of data correctness for the measured values; and

outputting the ground truth dataset, wherein the ground truth dataset is used to determine data correctness of measured values included in at least one new dataset by

receiving the new dataset from a data source, the new dataset comprising a plurality of data elements, wherein each data element of the new dataset includes an identifier and at least one measured value associated with the identifier; and

using respective data correctness values in the ground truth dataset for the measured values of respective data elements of the ground truth dataset determined as being with known identifiers, in determining data correctness values for measured values of data elements of the new dataset, wherein a known identifier is an identifier which is included in the ground truth dataset and in the new dataset, wherein determining data elements of the ground truth dataset with known identifiers comprising comparing identifiers included in the new dataset with identifiers included in the ground truth dataset, wherein data correctness values in the ground truth dataset are assigned for measured values of data elements of the new dataset based on overlapping of the measured values of data elements of the new dataset with measured values of data elements of the ground truth dataset with known identifiers.

2. The computer implemented method of claim 1 , wherein the at least one processor is further configured to determine a data correctness value for a measured value received from a respective data source for a subset of data elements received from the respective data source and to assign the data correctness value for the measured value determined for the subset of the data elements to the measured values of a same sensor information of the other data elements received from the respective data source.

3. The computer implemented method of claim 1 , wherein the at least one processor is used to determine the single data correctness value for the respective measured value for a respective identifier based on data correctness values for the measured value received from different data sources, when two or more data elements with the respective identifier and an identical measured value are included in the plurality of datasets received from the different data sources.

4. The computer implemented method of claim 3 , wherein the at least one processor is used to determine the single data correctness value for the respective measured value for the respective identifier by calculating

p

=

i

n

p

i

i

n

p

i

+

i

n

(

1

-

p

i

)

wherein p is a single data correctness value, n is a number of the different data sources, and p i is a probability for the measured value received from data source i to be correct.

5. The computer implemented method of claim 1 , wherein the at least one processor is used to determine the single data correctness value for the respective measured value for a respective identifier based on calculating

p

i

=

j

k

p

ij

k

wherein p i is a mean probability for the measured value received from the data source i to be correct, k is a number of probabilities for the measured value of the data element received from the data source i to be correct, and p ij is a j-th probability for the measured value of the data element of the data source i to be correct, when two or more data elements with the respective identifier and an identical measured value are included in one or more datasets received from the same data source.

6. The computer implemented method of claim 1 , wherein the at least one processor is further used to determine a data correctness value for at least one measured value of data elements of the new dataset with unknown identifier based on a probability distribution of data correctness for the respective measured value, when the new dataset comprises a threshold level of data elements including a known identifier, wherein an unknown identifier is an identifier which is not included in the ground truth dataset.

7. The computer implemented method of claim 6 , wherein the at least one processor is used for:

obtaining the probability distribution of data correctness for the respective measured value based on the ground truth dataset,

binning of data correctness values of the probability distribution of data correctness for the respective measured value, wherein each bin includes data elements of the ground truth dataset with data correctness values for the respective measured value within a range of data correctness values for the respective measured value,

extracting, from each bin, a number of samples each including a number of data elements,

determining, for each sample, a number of data elements with identifiers in a respective sample which are identical to identifiers in the new dataset, wherein the respective sample is discarded if the number of data elements with identifiers in the respective sample which are identical to identifiers in the new dataset is below a threshold identifier number and wherein the respective sample is further processed if the number of data elements with identifiers in the respective sample which are identical to identifiers in the new dataset is equal to or above the threshold identifier number,

determining for each sample a sample probability score based on the data correctness values for the respective measured value of the data elements with identical identifiers in the ground truth dataset and the new dataset,

determining for each bin a bin probability score of a respective bin based on the sample probability scores of the respective bin,

determining for each bin a weighted bin score by multiplying a respective bin probability score with a respective number of data elements with identifiers in the respective bin which are identical to identifiers in the new dataset,

determining for each bin a confidence score based on a respective weighted bin score and the respective number of data elements with identifiers in the respective bin which are identical to identifiers in the new dataset,

determining a final score by dividing a sum of the confidence scores by a sum of the number of data elements with identifiers in the bins which are identical to identifiers in the new dataset, and

assigning the final score as data correctness value to the measured value of the data elements of the new dataset with unknown identifier.

8. The computer implemented method of claim 7 , wherein the at least one processor is used to determine the confidence score for each of the bins based on Wilson confidence interval.

9. The computer implemented method of claim 1 ; wherein the at least one processor is used to add, to the ground truth dataset, the data elements of the new dataset with unknown identifier for which the data correctness values for the measured values were determined.

10. The computer implemented method of claim 1 , wherein the at least one processor is used to compare the data correctness values for a respective measured value of respective data elements of the ground truth dataset to data correctness values for the respective measured value of the respective data elements determined using at least one of a panel and calibration measurements for checking correctness of the measured values in dependence of a triggering event.

11. The computer implemented method of claim 10 , wherein the at least one processor is used to adjust the data correctness values for the respective measured value of the respective data elements of the ground truth dataset based on a difference between the data correctness values for the respective measured value of the respective data elements of the ground truth dataset and the data correctness values for the respective measured value of the respective data elements determined using at least one of a panel and calibration measurements for checking correctness of the measured values.

12. The computer implemented method of claim 10 , wherein the at least one processor is used for:

obtaining a probability distribution of data correctness for the respective measured value based on the ground truth dataset,

binning of data correctness values of the probability distribution of data correctness for the respective measured value, wherein each bin includes data elements of the ground truth dataset with data correctness values for the respective measured value within a range of data correctness values for the respective measured value,

determining, for each bin, a bin data correctness value for the respective measured value using at least one of a panel and calibration measurements for checking correctness of the measured values,

learning an adjustment weight for the data correctness values for the respective measured value for the data elements for each bin based on minimizing a difference between a median data correctness value for the respective bin determined based on the data correctness values for the respective measured value in a respective bin weighted based on a current adjustment weight and the bin data correctness value for the respective bin, and

adjusting, for each data element, the data correctness value for the respective measured value in the ground truth dataset based on the respective adjustment weight.

13. The computer implemented method of claim 12 , wherein the at least one processor is used to perform determining, for each bin, a bin data correctness value for the respective measured value using at least one of a panel and calibration measurements for checking correctness of the measured values based on a subset of data elements selected from the respective bin.

14. The computer implemented method of claim 1 , wherein the measured values are received from one or more sensor devices.

15. A data correctness management system for improving data correctness using a ground truth dataset, comprising:

a transceiver; and

at least one processor coupled to the transceiver, the at least one processor executing a code comprising:

program instructions to receive, via the transceiver, a plurality of datasets from different data sources, the datasets comprising a plurality of data elements, wherein each data element includes an identifier and at least one measured value associated with the identifier;

program instructions to determine data correctness values for the measured values using at least one of a panel and calibration measurements for checking correctness of the measured values, wherein a data correctness value is associated with a probability that a measured value is correct; and

program instructions to add a data element including a single data correctness value for each measured value for each identifier with which a respective measured value is associated to the ground truth dataset, wherein the single data correctness value being based on the determined data correctness values for the measured values, whereby the data correctness values in the ground truth dataset define probability distributions of data correctness for the measured values, and

program instructions to receive a new dataset from a data source, the new dataset comprising a plurality of data elements, wherein each data element of the new dataset includes an identifier and at least one measured value associated with the identifier; and to

use respective data correctness values in the ground truth dataset for the measured values of respective data elements of the ground truth dataset determined as being with known identifiers, in determining data correctness values for measured values of data elements of the new dataset, wherein a known identifier is an identifier which is included in the ground truth dataset and in the new dataset, wherein determining data elements of the ground truth dataset with known identifiers comprising comparing identifiers included in the new dataset with identifiers included in the ground truth dataset, wherein data correctness values in the ground truth dataset are assigned for measured values of data elements of the new dataset based on overlapping of the measured values of data elements of the new dataset with measured values of data elements of the ground truth dataset with known identifiers.

16. The data correctness management system of claim 15 , the code comprising program instructions to determine a data correctness value for measured values of data elements of the new dataset with unknown identifier based on a probability distribution of data correctness for a respective measured value, when the new dataset comprises a threshold level of data elements including a known identifier, wherein an unknown identifier is an identifier which is not included in the ground truth dataset.

17. The data correctness management system of claim 15 , the code further comprising program instructions to compare the data correctness values for a respective measured value of respective data elements of the ground truth dataset to data correctness values for the respective measured value of the respective data elements determined using at least one of a panel and calibration measurements for checking correctness of the measured values and to adjust the data correctness values for the respective measured value of the respective data elements of the ground truth dataset based on a difference between the data correctness values for the respective measured value of the respective data elements of the ground truth dataset and the data correctness values for the respective measured value of the respective data elements determined using at least one of a panel and calibration measurements for checking correctness of the measured values.

18. The data correctness management system of claim 15 , wherein the measured values are received from one or more sensor devices.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2026
From: ZEOTAP GMBH
To: ZEOTAP DATA GMBH
Reel/Frame 075369/0283 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2023
From: K S, SATHISH KUMAR
To: ZEOTAP GMBH
Reel/Frame 063040/0184 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2020
From: ZEOTAP INDIA PVT. LTD.
To: ZEOTAP GMBH
Reel/Frame 053507/0339 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2020
From: BANERJEA, PROJJOL; HEER, DANIEL
To: ZEOTAP GMBH
Reel/Frame 053492/0992 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 7, 2020
From: SAHU, SWAPNASARIT; SELVARAJ, ERNEST KIRUBAKARAN; AGARWAL, TUSHAR
To: ZEOTAP INDIA PVT. LTD.
Reel/Frame 053427/0022 →
Continuity (1)
Related Publication 20210365420A1 · Nov 25, 2021