IP Library › Granted Patent US 12,339,993
Granted Patent B2
US 12,339,993 · App. 18/171,301 · Granted Jun 24, 2025

Synthetic and traditional data stewards for selecting, optimizing, verifying and recommending one or more datasets

Inventors: Mary Elizabeth Chalk (Austin, TX); Robert Derward Rogers (Oakland, CA); Alan Donald Czeszynski (Pleasanton, CA)
Assignee: BeeKeeperAI, Inc.
G06F21/6245G06F16/2458G06F21/602G16H50/70G06F2221/2115
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,339,993
App. No.
18/171,301
Granted
Jun 24, 2025
Kind
B2
Abstract

Confirmation of data selection in a zero-trust environment is provided. In some embodiments, a synthetic data steward and/or a traditional data steward can receive the dataset(s). Additionally, a script is received from the algorithm developer. The dataset(s) and script(s) reside within a secure computing node and are therefore inaccessible by any party. The script(s) are executed, resulting in at least one confirmation about the data within the dataset(s). The script(s) complete any of confirming a format for data in the at least one dataset, the expected class values for data within the at least one dataset, an overall characterization and completeness of the at least one dataset, and/or an expected class membership for different data attributes within the at least one dataset.

Claims (28)

1. A computerized method for dataset selection confirmation within a zero-trust computing system, the method comprising:

receiving at a sequestered computing node at least one dataset;

receiving at the sequestered computing node a set of requirements from an algorithm developer;

applying the set of requirements to the at least one dataset to generate a tabular format for each at least one dataset;

converting each of the tabular formatted at least one dataset into a vector set;

receive a target vector;

generating various combinations of the at least one dataset to generate combination vector sets;

applying a cost function to minimize a difference between the target vector and one of the combination vector sets; and

selecting the datasets forming the minimized combination vector set for further processing.

2. The method of claim 1 , wherein the set of requirements confirms a format for data in the at least one dataset.

3. The method of claim 1 , wherein the set of requirements confirms expected class values for data within the at least one dataset.

4. The method of claim 1 , wherein the set of requirements confirms an overall characterization and completeness of the at least one dataset.

5. The method of claim 1 , wherein the set of requirements confirms an expected class membership for different data attributes within the at least one dataset.

6. The method of claim 5 , wherein the expected class membership is the target vector.

7. The method of claim 1 , wherein the sequestered computing node is within one of a data steward or a synthetic data steward.

8. The method of claim 7 , wherein the synthetic data steward aggregated the at least one dataset from a plurality of data stewards.

9. The method of claim 1 , wherein the cost function applies a penalty to conditions responsive to the algorithm developer.

10. A zero-trust computing system for dataset selection confirmation comprising:

a data store for receiving at a sequestered computing node at least one dataset, and receiving at the sequestered computing node a set of requirements from an algorithm developer; and

a processor configured to apply the set of requirements to the at least one dataset to generate a tabular format for each at least one dataset, convert each of the tabular formatted at least one dataset into a vector set, receive a target vector, generate various combinations of the at least one dataset to generate combination vector sets, apply a cost function to minimize a difference between the target vector and one of the combination vector sets, and selecting the datasets forming the minimized combination vector set for further processing.

11. The zero-trust computing system of claim 10 , wherein the set of requirements confirms a format for data in the at least one dataset.

12. The zero-trust computing system of claim 10 , wherein the set of requirements confirms expected class values for data within the at least one dataset.

13. The zero-trust computing system of claim 10 , wherein the set of requirements confirms an overall characterization and completeness of the at least one dataset.

14. The zero-trust computing system of claim 10 , wherein the set of requirements confirms an expected class membership for different data attributes within the at least one dataset.

15. The zero-trust computing system of claim 14 , wherein the expected class membership is the target vector.

16. The zero-trust computing system of claim 10 , wherein the sequestered computing node is within one of a data steward or a synthetic data steward.

17. The zero-trust computing system of claim 16 , wherein the synthetic data steward aggregated the at least one dataset from a plurality of data stewards.

18. The zero-trust computing system of claim 10 , wherein the cost function applies a penalty to conditions responsive to the algorithm developer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2025
From: CHALK, MARY ELIZABETH; ROGERS, ROBERT DERWARD; CZESZYNSKI, ALAN DONALD
To: BEEKEEPERAI, INC.
Reel/Frame 070472/0972 →
Continuity (7)
Continuation In Part 18169867 · Feb 15, 2023
Continuation 18169111 · Feb 14, 2023
Continuation 18169122 · Feb 14, 2023
Continuation 18168560 · Feb 13, 2023
Continuation 18168560 · Feb 13, 2023
Provisional Application 63313774 · Feb 25, 2022
Related Publication 20230274026A1 · Aug 31, 2023
References Cited (17)
US 20040158581A1 · Kotlyar · 2004 [cited by examiner]
US 20170068477A1 · Yu · 2017 [cited by applicant]
US 20180018590A1 · Szeto · 2018 [cited by examiner]
US 20180114142A1 · Mueller · 2018 [cited by applicant]
US 20180165418A1 · Swartz · 2018 [cited by examiner]
US 20200111019A1 · Goodsitt et al. · 2020 [cited by applicant]
US 20200311300A1 · Callcut · 2020 [cited by examiner]
US 20200322820A1 · Carter et al. · 2020 [cited by applicant]
US 20210092160A1 · Crabtree et al. · 2021 [cited by applicant]
US 20210173854A1 · Wilshinsky · 2021 [cited by applicant]
US 20220021711A1 · Marsh et al. · 2022 [cited by applicant]
US 20220215243A1 · Narayanaswami · 2022 [cited by examiner]
US 20230325757A1 · Conway · 2023 [cited by examiner]
US 20230368070A1 · Bhargava · 2023 [cited by examiner]
Gonzalez-Abril et al. Statistical validation of synthetic data for lung cancer patients generated by using generative adversarial networks. 2022. Electronics, 11(20), 3277. doi:http://dx.doi.org/10.3390/electronics11203… [cited by examiner]
Zawafzki et al. “Synthetic Data Generation in Small Datasets to Improve Classification Performance for Chronic Heart Failure Prediction,” 2023 Computing in Cardiology (CinC), Atlanta, GA, USA, 2023, pp. 1-4, doi: 10.224… [cited by examiner]
ISA/US, “Notification of Transmittal of the ISR and the Written Opinion of the International Searching Authority, or the Declaration,” in PCT Application No. PCT/US2023/063083, Jul. 24, 2023, 16 pages. [cited by applicant]