IP Library › Granted Patent US 12,141,319
Granted Patent B2
US 12,141,319 · App. 18/169,122 · Granted Nov 12, 2024

Systems and methods for dataset quality quantification in a zero-trust computing environment

Inventors: Mary Elizabeth Chalk (Austin, TX); Robert Derward Rogers (Oakland, CA)
Assignee: BeeKeeperAI, Inc.
G06F21/6245G06F16/2237G06F16/2458G06F16/2462G06F21/602G16H50/70G06F2221/2115
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,141,319
App. No.
18/169,122
Granted
Nov 12, 2024
Kind
B2
Abstract

Systems and methods for the quantification of sample set quality is provided. In some embodiments, a sample dataset and a sample vector set are received. A rule-based screening of the sample dataset is applied to generate a heuristic quality score. Additionally, a sample vector set is generated from the sample dataset. The difference between the sample vector set and the example vector set is calculated to generate a degree of difference quality score. The heuristic quality score and the degree of difference quality score are normalized and then combined into a quality metric. Calculating the difference between the sample vector set and the example vector set is by framing the distance as a p-value in a hypothesis test, compared against a threshold.

Claims (28)

1. A computerized method for dataset quality quantification in a zero-trust environment, the method comprising:

receiving a sample dataset from a data steward;

applying a rule based screening of the sample dataset to generate a heuristic quality score, wherein the heuristic quality score quantifies erroneous data within the sample dataset;

generating a sample vector set from the sample dataset;

receiving an example vector set generated from at least one of a synthetic dataset and an amalgamation of different vector sets from a plurality of different data stewards;

calculating a difference between the sample vector set and the example vector set to generate a degree of difference quality score; and

combining the heuristic quality score and the degree of difference quality score into a quality metric;

determining that the quality metric is above a threshold, and

when the quality metric is above the threshold then recommending the sample dataset to a data consumer.

2. The method of claim 1 , wherein the heuristic quality score and the degree of difference quality score are normalized before combining.

3. The method of claim 1 , wherein the combining is by at least one of adding, averaging, weighted averaging, and multiplying.

4. The method of claim 1 , wherein the heuristic quality score and the degree of difference quality score are bucketized by thresholds before combining.

5. The method of claim 1 , wherein the calculating the difference is by framing the distance as a p-value in a hypothesis test, compared against a threshold.

6. The method of claim 1 , wherein the generating the sample vector set includes:

encoding the dataset according to a set of classes;

generating a matrix of the encoded dataset, wherein each row of the matrix is a patient, and each column is a class or subset of classes in the set of classes; and

converting the generated matrix into a series of vector spaces.

7. A zero-trust computing system for dataset quality quantification, the system comprising:

a data store hosted within a secure computing node for receiving a sample dataset from a data steward, and receiving an example vector set generated from at least one of a synthetic dataset and an amalgamation of different vector sets from a plurality of different data stewards; and

a processor and memory implementing a runtime server within the secure computing node for applying a rule based screening of the sample dataset to generate a heuristic quality score, wherein the heuristic quality score quantifies erroneous data within the sample dataset, generating a sample vector set from the sample dataset, calculating a difference between the sample vector set and the example vector set to generate a degree of difference quality score, and combining the heuristic quality score and the degree of difference quality score into a quality metric, and when the quality metric is above a threshold then recommending the sample dataset to a data consumer.

8. The system of claim 7 , wherein the heuristic quality score and the degree of difference quality score are normalized before combining.

9. The system of claim 7 , wherein the combining is by at least one of adding, averaging, weighted averaging, and multiplying.

10. The system of claim 7 , wherein the heuristic quality score and the degree of difference quality score are bucketized by thresholds before combining.

11. The system of claim 7 , wherein the calculating the difference is by framing the distance as a p-value in a hypothesis test, compared against a threshold.

12. The system of claim 7 , wherein the generating the sample vector set includes:

encoding the dataset according to a set of classes;

generating a matrix of the encoded dataset, wherein each row of the matrix is a patient, and each column is a class or subset of classes in the set of classes; and

converting the generated matrix into a series of vector spaces.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2024
From: CHALK, MARY ELIZABETH; ROGERS, ROBERT DERWARD
To: BEEKEEPERAI, INC.
Reel/Frame 068547/0826 →
Continuity (3)
Continuation 18168560 · Feb 13, 2023
Provisional Application 63313774 · Feb 25, 2022
Related Publication 20230274025A1 · Aug 31, 2023