IP Library › Granted Patent US 12,524,567
Granted Patent B2
US 12,524,567 · App. 18/168,560 · Granted Jan 13, 2026

Systems and methods for dataset selection optimization in a zero-trust computing environment

Inventors: Mary Elizabeth Chalk (Austin, TX); Robert Derward Rogers (Oakland, CA)
Assignee: BeeKeeperAI, Inc.
G06F21/6245G06F16/2237G06F16/2458G06F16/2462G06F21/602G16H50/70G06F2221/2115
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,567
App. No.
18/168,560
Granted
Jan 13, 2026
Kind
B2
Abstract

Systems and methods for the selection, verification and recommendation of cohort sample sets is provided. In some embodiments, a dataset selection optimization includes first receiving at data stewards classes of data required by the data consumer. The data stewards process their data (or a subset of their data) into a vector set within a sequestered computing node. These vector sets are transferred to a core management system for minimizing a difference between a target vector and any combination of the data stewards' vector sets. A cost function may also be applied to the vector sets during this optimization. Once the data steward(s) that best match the target vector are identified, they may be placed in contact with the data consumer for access of their information.

Claims (33)

1 . A computerized method for dataset selection optimization, the method comprising:

receiving at a plurality of data stewards a set of classes required by a data consumer, wherein each data steward has a dataset;

at each data steward, processing their respective dataset to generate a vector set responsive to the set of classes;

homomorphically encrypting a target vector set and the vector sets;

transmitting the encrypted vector set from each data steward to a central system; and

minimizing the difference between the target vector set and any given combination of the vector sets from the data stewards without decrypting the vector sets or the target vector set by subtracting the combination vector set from the target vector while applying a cost function to each vector set based upon a financial cost of the dataset.

2 . The method of claim 1 , wherein the processing occurs within a sequestered computing node, wherein the sequestered computing node preserves privacy of data assets and the set of required classes.

3 . The method of claim 1 , wherein the cost function is for one of number of data stewards, geography of the datasets, and data set quality.

4 . The method of claim 1 , wherein the minimizing is according to the equation of: Goal=minimize∥T{target}−T(Union({data steward}))∥.

5 . The method of claim 1 , further comprising selecting the datasets that minimize the difference.

6 . The method of claim 5 , further comprising facilitating contact between the data stewards associated with the selected datasets and the data consumer.

7 . The method of claim 6 , wherein the facilitating is acting as a broker.

8 . The method of claim 1 , wherein the data consumer is at least one of a clinical trial administrator, a researcher, a clinician, and a public health official.

9 . The method of claim 1 , wherein the generating the vector set includes:

encoding the dataset according to the set of classes;

generating a matrix of the encoded dataset, wherein each row of the matrix is a patient and each column is a class or subset of classes in the set of classes; and

converting the generated matrix into a series of vector spaces.

10 . The zero-trust computing system configured for data set selection optimization, the system comprising:

a plurality of data stewards configured to receive a set of classes required by a data consumer, wherein each data steward has a dataset, and wherein each data steward is configured to process its respective dataset a vector set responsive to the set of classes, and homomorphically encrypting the vector sets;

a first processor for homomorphically encrypting a target vector set; and

at least one central processing unit configured to receive the vector set from each data steward, and wherein the at least one central processing unit is further configured to minimize the difference between the target vector set and any given combination of the respective vector sets from the plurality of data stewards without decrypting the vector sets or the target vector set by subtracting the combination vector set from the target vector while applying a cost function to each vector set based upon a financial cost of the dataset.

11 . The system of claim 10 , wherein the processing at each data steward occurs within a sequestered computing node, wherein the sequestered computing node preserves privacy of data assets and the set of required classes.

12 . The system of claim 10 , wherein the at least one central processing unit is further configured to apply a cost function to the minimizing calculation.

13 . The system of claim 12 , wherein the cost function is for one of number of data stewards, geography of the datasets, financial cost of the datasets, and data set quality.

14 . The system of claim 10 , wherein the minimizing is according to the equation of: Goal=minimize∥T{target}−T(Union({data steward}))∥.

15 . The system of claim 10 , wherein the at least one central processing unit is configured to select the datasets that minimize the difference.

16 . The system of claim 15 , wherein the at least one central processing unit is further configured to facilitate contact between the data stewards associated with the selected datasets and the data consumer.

17 . The system of claim 16 , wherein the facilitating is acting as a broker.

18 . The system of claim 10 , wherein the data consumer is at least one of a clinical trial administrator, a researcher, a clinician, and a public health official.

19 . The system of claim 10 , wherein the generating the vector set includes:

encoding the dataset according to the set of classes;

generating a matrix of the encoded dataset, wherein each row of the matrix is a patient and each column is a class or subset of classes in the set of classes; and

converting the generated matrix into a series of vector spaces.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2025
From: CHALK, MARY ELIZABETH; ROGERS, ROBERT DERWARD
To: BEEKEEPERAI, INC.
Reel/Frame 070473/0622 →
Continuity (2)
Provisional Application 63313774 · Feb 25, 2022
Related Publication 20230274024A1 · Aug 31, 2023
References Cited (16)
US 20170068477A1 · Yu · 2017 [cited by applicant]
US 20180114142A1 · Mueller · 2018 [cited by applicant]
US 20200111019A1 · Goodsitt et al. · 2020 [cited by applicant]
US 20200322820A1 · Carter et al. · 2020 [cited by applicant]
US 20210092160A1 · Crabtree et al. · 2021 [cited by applicant]
US 20210173854A1 · Wilshinsky · 2021 [cited by applicant]
US 20210406346A1 · Shiue · 2021 [cited by examiner]
US 20220021711A1 · March et al. · 2022 [cited by applicant]
US 20220183571A1 · Johnson · 2022 [cited by examiner]
US 20230021563A1 · Narayanam · 2023 [cited by examiner]
Dhingra, Evaluation Metrics, pp. 1-10, May 21 (Year: 2021). [cited by examiner]
Luchenko, Zero Trust Technology Application For AI Medical Research, pp. 264-267, Nov. 2021. [cited by examiner]
Doyle, Accelerating healthcare AI innovation with Zero Trust technology, pp. 1-5, Oct. 26, 2021. [cited by examiner]
Roth, UCSF Joins Forces With Tech Companies to Eliminate Data-Sharing Risks, pp. 1-6, Jan. 12, 2021. [cited by examiner]
Kurtzman, UCSF, Fortanix, Intel, and Microsoft Azure Utilize Privacy-Preserving Analytics to Accelerate AI in Health Care, pp. 1-7, Oct. 8, 2020. [cited by examiner]
ISA/US, “Notification of Transmittal of the ISR and the Written Opinion of the International Searching Authority, or the Declaration,” in PCT Application No. PCT/US2023/063083, Jul. 24, 2023, 16 pages. [cited by applicant]
Cited By (1)
US 12,737,495