IP Library › Granted Patent US 12,462,165
Granted Patent B2
US 12,462,165 · App. 18/665,954 · Granted Nov 4, 2025

Distributed privacy-preserving computing on protected data

Inventors: Rachael A. Callcut (San Francisco, CA); Michael Blum (San Francisco, CA); Joseph H. Hesse (San Francisco, CA); Robert D. Rogers (Pleasanton, CA); Scott Hammond (Mill Valley, CA); Mary Elizabeth Chalk (Austin, TX)
Assignee: The Regents of the University of California
G06N5/02G06F16/256G06F21/53G06F21/602G06F30/20G06N20/00G06F21/6245
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,165
App. No.
18/665,954
Granted
Nov 4, 2025
Kind
B2
Abstract

The present disclosure relates to techniques for developing artificial intelligence algorithms by distributing analytics to multiple sources of privacy protected, harmonized data. Particularly, aspects are directed to a computer implemented method that includes receiving an algorithm and input data requirements associated with the algorithm, identifying data assets as being available from a data host based on the input data requirements, curating the data assets within a data storage structure that is within infrastructure of the data host, and integrating the algorithm into a secure capsule computing framework. The secure capsule computing framework serves the algorithm to the data assets within the data storage structure in a secure manner that preserves privacy of the data assets and the algorithm. The computer implemented method further includes running the data assets through the algorithm to obtain an inference.

Claims (38)

1 . A system comprising:

one or more data processors; and

a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform actions including:

receiving an algorithm and input data requirements associated with the algorithm, wherein the input data requirements include optimization and/or validation selection criteria for data assets to be run on the algorithm;

obtaining the data assets based on the optimization and/or validation selection criteria for the data assets;

preparing the data assets for processing by the algorithm, wherein the preparing comprises applying harmonizing transformers to the data assets to generate harmonized data assets; and

running the harmonized data assets through the algorithm, wherein the algorithm is in a secure capsule computing framework that serves the algorithm to the harmonized data assets in a secure manner that preserves privacy of the data assets and the algorithm.

2 . The system of claim 1 , wherein the preparing further comprises:

preparing a transformer prototype set of data to use as a guide for developing algorithms for data transformation, wherein the transformer prototype set of data captures key attributes of a harmonization process;

creating a first set of harmonizing transformers for transformation of the data assets based on a present format of data in the transformer prototype set of data;

applying the first set of harmonizing transformers to the data assets to generate transformed data assets;

preparing a harmonization prototype set of data to use as a guide for developing algorithms for data transformation, wherein the harmonization prototype set of data captures key attributes of the harmonization process;

creating a second set of harmonizing transformers for transformation of the transformed data assets based on a present format of data in the harmonization prototype set of data; and

applying the second set of harmonizing transformers to the transformed data assets to generate the harmonized data assets.

3 . The system of claim 2 , wherein the preparing further comprises de-identifying the transformer prototype set of data and making the de-identified transformer prototype set of data available to an algorithm developer for the purpose of creating the first set of harmonizing transformers for transformation of the data assets.

4 . The system of claim 2 , wherein the preparing further comprises annotating the transformed data assets according to a predefined annotation protocol to generate annotated data sets, wherein the second set of harmonizing transformers is applied to the annotated data sets to generate harmonized data assets.

5 . The system of claim 2 , wherein the applying the first set of harmonizing transformers to the data assets and the applying the second set of harmonizing transformers to the annotated data assets are performed within a data storage structure separate from the secure capsule computing framework.

6 . The system of claim 1 , wherein the harmonized data assets through the algorithm comprises executing a training workflow that includes: creating multiple instances of the algorithm, splitting the harmonized data assets into sets of training data and one or more sets of testing data, training the multiple instances of the algorithm on the sets of training data, integrating results from the training each of the multiple instances of the algorithm into a fully federated model, running the one or more sets of testing data through the fully federated model, and computing performance of the fully federated model based on the running of the one or more sets of testing data.

7 . The system of claim 1 , wherein the running the harmonized data assets through the algorithm comprises executing a validation workflow that includes: splitting the data assets in one or more sets of validation data, running the one or more sets of validation data through the algorithm, and computing performance of the algorithm based on the running of the one or more sets of validation data.

8 . The system of claim 1 , wherein the secure capsule computing framework is provisioned within a computing infrastructure configured to accept encrypted code required to run the algorithm.

9 . The system of claim 1 , wherein the actions further comprise reporting aggregate results of the training workflow and performance metrics for the fully federated model.

10 . A system comprising:

one or more data processors; and

a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform actions including:

receiving an algorithm and input data requirements associated with the algorithm, wherein the input data requirements include optimization and/or validation selection criteria for data assets to be run on the algorithm;

obtaining the data assets based on the optimization and/or validation selection criteria for the data assets;

preparing the data assets for processing by the algorithm; and

running the data assets through the algorithm, wherein the running comprises: passing the data assets from a data storage structure to the algorithm in a secure capsule computing framework, and optimizing, validating, or computing inference with the algorithm using the data assets, wherein the algorithm is in a secure capsule computing framework that serves the algorithm to the data assets in accordance with encrypted code stored inside the secure capsule computing framework.

11 . The system of claim 10 , wherein the algorithm is optimized using the data assets, and wherein the running the data assets through the algorithm comprises executing a training workflow that includes: creating multiple instances of the algorithm, splitting the data assets into sets of training data and one or more sets of testing data, training the multiple instances of the algorithm on the sets of training data, integrating results from the training each of the multiple instances of the algorithm into a fully federated model, running the one or more sets of testing data through the fully federated model, and computing performance of the fully federated model based on the running of the one or more sets of testing data.

12 . The system of claim 10 , wherein the running the data assets through the algorithm comprises executing a validation workflow that includes: splitting the data assets in one or more sets of validation data, running the one or more sets of validation data through the algorithm, and computing performance of the algorithm based on the running of the one or more sets of validation data.

13 . The system of claim 12 , wherein the actions further comprise reporting optimized hyperparameters and performance metrics of the algorithm.

14 . The system of claim 10 , wherein the optimization and/or validation selection criteria define characteristics, formats and requirements for the data assets to be run on the algorithm.

15 . The system of claim 10 , wherein the preparing the data assets comprises applying one or more transforms to the data assets, annotating the data assets, harmonizing the data assets, or a combination thereof.

16 . The system of claim 10 , wherein the obtaining comprises:

identifying the data assets as being available based on the optimization and/or validation selection criteria for the data assets;

selecting a data storage structure from multiple data storage structures based on a type of the algorithm, a type of data within the data assets, system requirements of a data processing system, or a combination thereof;

provisioning the data storage structure within an infrastructure; and

curating the data assets within the data storage structure.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2024
From: CALLCUT, RACHAEL A.; BLUM, MICHAEL; HESSE, JOSEPH H.; ROGERS, ROBERT D.; HAMMOND, SCOTT; CHALK, MARY ELIZABETH
To: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
Reel/Frame 068185/0028 →
Continuity (6)
Continuation 18335053 · Jun 14, 2023
Continuation 17988664 · Nov 16, 2022
Continuation 16831763 · Mar 26, 2020
Provisional Application 62948556 · Dec 16, 2019
Provisional Application 62824183 · Mar 26, 2019
Related Publication 20240386290A1 · Nov 21, 2024
References Cited (41)
US 8346534B2 · Csomai et al. · 2013 [cited by applicant]
US 9832226B2 · Epstein · 2017 [cited by applicant]
US 10133878B2 · Horvitz et al. · 2018 [cited by applicant]
US 10198399B1 · Fritchman · 2019 [cited by examiner]
US 11409993B2 · Ghosh et al. · 2022 [cited by applicant]
US 11531904B2 · Callcut et al. · 2022 [cited by applicant]
US 11748633B2 · Callcut et al. · 2023 [cited by applicant]
US 20090254971A1 · Herz et al. · 2009 [cited by applicant]
US 20170258390A1 · Howard · 2017 [cited by applicant]
US 20180150609A1 · Kim et al. · 2018 [cited by applicant]
US 20180182037A1 · Lange et al. · 2018 [cited by applicant]
US 20180294047A1 · Hosseini et al. · 2018 [cited by applicant]
US 20190377897A1 · Griffin · 2019 [cited by examiner]
US 20190392305A1 · Gu · 2019 [cited by examiner]
US 20200082270A1 · Gu et al. · 2020 [cited by applicant]
US 20200104705A1 · Bhowmick et al. · 2020 [cited by applicant]
US 20200125739A1 · Verma · 2020 [cited by examiner]
US 20200210867A1 · Banis et al. · 2020 [cited by applicant]
US 20200311617A1 · Swan · 2020 [cited by examiner]
US 20220092216A1 · Mohassel et al. · 2022 [cited by applicant]
EP 3449414A1 · 2019 [cited by applicant]
JP 6329333B1 · 2018 [cited by applicant]
WO 2017187207A1 · 2017 [cited by applicant]
“Predictive Model Markup Language”, Wikipedia, XP055978337, Available Online at: https://en.Wikipedia.org/w/index,php?title=Predictive_Model_Markup_Language&oldid=840563195, May 10, 2018, pp. 1-6. [cited by applicant]
U.S. Appl. No. 16/831,763, Non-Final Office Action, Mailed May 11, 2022, 10 pages. [cited by applicant]
U.S. Appl. No. 16/831,763, Notice of Allowance, Mailed Aug. 17, 2022, 12 pages. [cited by applicant]
Al-Rubaie et al., “Privacy Preserving Machine Learning: Threats and Solutions,” IEEE Security & Privacy, vol. 17, No. 2, Mar.-Apr. 2019, pp. 1-18. [cited by applicant]
Application No. CA3,133,466, Office Action, Mailed Dec. 6, 2022, 3 pages. [cited by applicant]
Application No. EP20778145.1, Extended European Search Report, Mailed Nov. 17, 2022, 10 pages. [cited by applicant]
Hesamifard et al., “Privacy-Preserving Machine Learning in Cloud,” In Proceedings of the 2017 on Cloud Computing Security Workshop (CCSW'17), Nov. 3, 2017, pp. 39-43. [cited by applicant]
Hunt et al., “Chiron: Privacy-Preserving Machine Learning as a Service,” Cornell University Library, Mar. 15, 2018, pp. 1-15. [cited by applicant]
Hynes, “Efficient Privacy-Preserving ML Using TVM”, XP055978065, Available Online at: https://tvm.apache.org/2018/10/09/ml-in-tees, Oct. 9, 2018, pp. 1-3. [cited by applicant]
Mohassel et al., “SecureML: A System for Scalable Privacy-Preserving Machine Learning,” IEEE Symposium on Security and Privacy, Jun. 26, 2017, pp. 1-38. [cited by applicant]
Application No. PCT/US2020/025083, International Preliminary Report on Patentability, Mailed Oct. 7, 2021, 13 pages. [cited by applicant]
Application No. PCT/US2020/025083, International Search Report and Written Opinion, Mailed Aug. 14, 2020, 17 pages. [cited by applicant]
PCT/US2020/025083, Invitation to Pay Additional Fees and, Where Applicable, Protest Fee, May 22, 2020, 3 pages. [cited by applicant]
U.S. Appl. No. 17/988,664, Notice of Allowance, Mailed Mar. 16, 2023, 10 pages. [cited by applicant]
Japanese Application No. JP2021-557379, Office Action, Mailed May 22, 2023, 3 pages. [cited by applicant]
Korean Application No. KR10-2021-7034758, Office Action, Mailed Jun. 26, 2023, 4 pages. [cited by applicant]
Canadian Application No. CA3133466, Office Action, Mailed Sep. 6, 2023, 4 pages. [cited by applicant]
“Office Action,” mailed Apr. 2, 2024, in Brazilian Patent Application No. BR112021018241-1. 13 pages (includes English translation). [cited by applicant]