IP Library › Granted Patent US 12,518,214
Granted Patent B2
US 12,518,214 · App. 18/137,812 · Granted Jan 6, 2026

Distributed machine learning systems including generation of synthetic data

Inventors: Christopher W. Szeto (Scotts Valley, CA); Stephen Charles Benz (Santa Cruz, CA); Nicholas J. Witchey (Laguna Hills, CA)
Assignees: NantOmics, LLC; Nant Holdings IP, LLC
G06N20/00G06F21/6254G06N20/10G16H10/60G16H40/20G16H50/20G16H50/50G06F21/6245
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,214
App. No.
18/137,812
Granted
Jan 6, 2026
Kind
B2
Abstract

A distributed, online machine learning system is presented. Contemplated systems include many private data servers, each having local private data. Researchers can request that relevant private data servers train implementations of machine learning algorithms on their local private data without requiring de-identification of the private data or without exposing the private data to unauthorized computing systems. The private data servers also generate synthetic or proxy data according to the data distributions of the actual data. The servers then use the proxy data to train proxy models. When the proxy models are sufficiently similar to the trained actual models, the proxy data, proxy model parameters, or other learned knowledge can be transmitted to one or more non-private computing devices. The learned knowledge from many private data servers can then be aggregated into one or more trained global models without exposing private data.

Claims (48)

1 . A computer-based distributed machine learning system comprising:

at least one private data server storing local private data and having a local modeling agent; and

a non-private data server coupled with the at least one private data server over a network, the non-private data server lacking authorized access to the local private data, the non-private data server comprising at least one processor that, upon execution of software instructions stored in a computer readable memory, performs operations of:

transmitting a definition of a machine learning task to the modeling agent of the at least one private data server, the definition of the machine learning task including machine learning model instructions and local private data features;

enabling the modeling agent of the at least one private data server to generate synthetic data capable of reproducing knowledge gained from execution of the machine learning model instructions on at least some of the local private data having the local private data features, wherein the synthetic data is generated based on at least one data distribution of the local private data which includes multiple patient sample points for different patients, the synthetic data includes at least one synthetic data distribution of synthetic patient data which is different than the at least one data distribution of the local private data, and each of the multiple patient sample points includes at least one a symptom, a test result, a provider name, a diagnosis, a current procedural terminology (CPT) code, an international classification of diseases (ICD) code, or a diagnostic and statistical manual of mental disorders (DSM) code;

receiving from the modeling agent of the at least one private data server, first proxy model data representative of the knowledge gained via the synthetic data;

comparing the first proxy model data with second proxy model data received from a second private data server;

in response to the first proxy model data and the second proxy model data having at least one of a same shape and a same overall property, aggregating a combination of the second proxy model data and the first proxy model data representative of knowledge gained via the synthetic data into a global model corresponding to the machine learning task; and

flagging the second proxy model data in response to the second proxy model data having a different shape than the first proxy model data, to determine whether an underlying private data distribution set of the second proxy model data is corrupted, has missing data, or contains multiple outliers.

2 . The system of claim 1 , wherein the first proxy model data comprises compressed learned data.

3 . The system of claim 1 , wherein the first proxy model data comprises lossy compressed learned data.

4 . The system of claim 1 , wherein the first proxy model data comprises the synthetic data.

5 . The system of claim 1 , wherein the first proxy model data comprises proxy model parameters derived from the synthetic data.

6 . The system of claim 5 , wherein the operations further include duplicating the synthetic data according to the proxy model parameters.

7 . The system of claim 6 , wherein the proxy model parameters include a seed for a deterministic function able to generate the synthetic data.

8 . The system of claim 1 , wherein the first proxy model data comprises a trained proxy model trained on the synthetic data.

9 . The system of claim 1 , wherein the synthetic data comprises Monte Carlo data.

10 . The system of claim 1 , wherein the local private data comprises private patient data.

11 . The system of claim 10 , wherein the private patient data comprises at least one of healthcare data or genomic data.

12 . The system of claim 1 , wherein the local private data includes at least one of insurance data, financial data, social media profile data, human capital data, proprietary experimental data, gaming or gambling data, military data, network traffic data, or shopping or marketing data.

13 . The system of claim 1 , wherein the operations further include paying a fee in exchange for accessing the modeling agent of the at least one private data server.

14 . The system of claim 1 , wherein the machine learning model instructions comprise at least one of supervised machine learning instructions, unsupervised machine learning instructions, or machine learning clustering instructions.

15 . The system of claim 1 , wherein the machine learning model instructions comprise at least one of machine learning regression instructions or machine learning classification instructions.

16 . The system of claim 1 , wherein the operations further include receiving private data metadata about the local private data from the local modeling agent.

17 . The system of claim 16 , wherein the operations further include generating the definition of the machine learning task based on the private data metadata.

18 . The system of claim 17 , wherein the private data metadata comprises attribute space information relating to the local private data.

19 . The system of claim 1 , wherein synthetic data is generated such that similarity score between a trained proxy model trained on the synthetic data and a trained actual model trained on at least some of the local private data, is below a threshold.

20 . The system of claim 19 , wherein the first proxy model data is received upon satisfaction of transmission requirements defined based on the similarity score.

21 . A method of computer-based distributed machine learning, the method comprising:

transmitting, by at least one processor of a non-private data server coupled with at least one private data server over a network, a definition of a machine learning task to a modeling agent of at least one private data server,

the at least one private data server storing local private data and having a local modeling engine,

the non-private data server lacking authorized access to the local private data, and

the definition of the machine learning task including machine learning model instructions and local private data features;

enabling the modeling agent of the at least one private data server to generate synthetic data capable of reproducing knowledge gained from execution of the machine learning model instructions on at least some of the local private data having the local private data features, wherein the synthetic data is generated based on at least one data distribution of the local private data which includes multiple patient sample points for different patients, the synthetic data includes at least one synthetic data distribution of synthetic patient data which is different than the at least one data distribution of the local private data, and each of the multiple patient sample points includes at least one a symptom, a test result, a provider name, a diagnosis, a current procedural terminology (CPT) code, an international classification of diseases (ICD) code, or a diagnostic and statistical manual of mental disorders (DSM) code;

receiving from the modeling agent of the at least one private data server, first proxy model data representative of the knowledge gained via the synthetic data;

comparing the first proxy model data with second proxy model data received from a second private data server;

in response to the first proxy model data and the second proxy model data having at least one of a same shape and a same overall property, aggregating a combination of the second proxy model data and the first proxy model data representative of knowledge gained via the synthetic data into a global model corresponding to the machine learning task; and

flagging the second proxy model data in response to the second proxy model data having a different shape than the first proxy model data, to determine whether an underlying private data distribution set of the second proxy model data is corrupted, has missing data, or contains multiple outliers.

22 . A non-transitory computer-readable medium comprising computer-executable instructions configured to, when executed by at least one processor, cause the processor to perform operations including:

transmitting, by at least one processor of a non-private data server coupled with at least one private data server over a network, a definition of a machine learning task to a modeling agent of at least one private data server,

the at least one private data server storing local private data and having a local modeling engine,

the non-private data server lacking authorized access to the local private data, and

the definition of the machine learning task including machine learning model instructions and local private data features;

enabling the modeling agent of the at least one private data server to generate synthetic data capable of reproducing knowledge gained from execution of the machine learning model instructions on at least some of the local private data having the local private data features, wherein the synthetic data is generated based on at least one data distribution of the local private data which includes multiple patient sample points for different patients, the synthetic data includes at least one synthetic data distribution of synthetic patient data which is different than the at least one data distribution of the local private data, and each of the multiple patient sample points includes at least one a symptom, a test result, a provider name, a diagnosis, a current procedural terminology (CPT) code, an international classification of diseases (ICD) code, or a diagnostic and statistical manual of mental disorders (DSM) code;

receiving from the modeling agent of the at least one private data server, first proxy model data representative of the knowledge gained via the synthetic data;

comparing the first proxy model data with second proxy model data received from a second private data server;

in response to the first proxy model data and the second proxy model data having at least one of a same shape and a same overall property, aggregating a combination of the second proxy model data and the first proxy model data representative of knowledge gained via the synthetic data into a global model corresponding to the machine learning task; and

flagging the second proxy model data in response to the second proxy model data having a different shape than the first proxy model data, to determine whether an underlying private data distribution set of the second proxy model data is corrupted, has missing data, or contains multiple outliers.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2023
From: WITCHEY, NICHOLAS J.
To: NANTWORKS, LLC
Reel/Frame 063646/0452 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2023
From: SZETO, CHRISTOPHER; BENZ, STEPHEN CHARLES
To: NANTOMICS, LLC
Reel/Frame 063646/0524 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2023
From: NANTWORKS, LLC
To: NANT HOLDINGS IP, LLC
Reel/Frame 063646/0561 →
Continuity (4)
Continuation 17890953 · Aug 18, 2022
Continuation 15651345 · Jul 17, 2017
Provisional Application 62363697 · Jul 18, 2016
Related Publication 20230267375A1 · Aug 24, 2023
References Cited (87)
US 7899225B2 · Collins et al. · 2011 [cited by applicant]
US 7974714B2 · Hoffberg · 2011 [cited by applicant]
US 8429103B1 · Aradhye · 2013 [cited by examiner]
US 8755837B2 · Rhoads et al. · 2014 [cited by applicant]
US 8873813B2 · Tadayon et al. · 2014 [cited by applicant]
US 8954365B2 · Criminisi et al. · 2015 [cited by applicant]
US 9202253B2 · Macoviak et al. · 2015 [cited by applicant]
US 9224180B2 · Macoviak et al. · 2015 [cited by applicant]
US 9646134B2 · Sanborn et al. · 2017 [cited by applicant]
US 11461690B2 · Szeto et al. · 2022 [cited by applicant]
US 20040267773A1 · Levine · 2004 [cited by examiner]
US 20050197783A1 · Kuchinsky et al. · 2005 [cited by applicant]
US 20060063156A1 · Willman et al. · 2006 [cited by applicant]
US 20090037351A1 · Kristal et al. · 2009 [cited by applicant]
US 20090203588A1 · Willman et al. · 2009 [cited by applicant]
US 20100168533A1 · Johnsen et al. · 2010 [cited by applicant]
US 20110228976A1 · Fitzgibbon et al. · 2011 [cited by applicant]
US 20120303558A1 · Jaiswal · 2012 [cited by applicant]
US 20130085773A1 · Yao et al. · 2013 [cited by applicant]
US 20130090254A1 · Lamb et al. · 2013 [cited by applicant]
US 20130291118A1 · Li · 2013 [cited by examiner]
US 20140012843A1 · Soon-Shiong · 2014 [cited by examiner]
US 20140038836A1 · Higgins et al. · 2014 [cited by applicant]
US 20140046696A1 · Higgins et al. · 2014 [cited by applicant]
US 20140074760A1 · Boldyrev et al. · 2014 [cited by applicant]
US 20140088989A1 · Krishnapuram · 2014 [cited by examiner]
US 20140101090A1 · Gordon · 2014 [cited by examiner]
US 20140142970A1 · Baronov et al. · 2014 [cited by applicant]
US 20140201126A1 · Zadeh et al. · 2014 [cited by applicant]
US 20140222349A1 · Higgins et al. · 2014 [cited by applicant]
US 20140222737A1 · Chen et al. · 2014 [cited by applicant]
US 20140344343A1 · Zarkesh · 2014 [cited by examiner]
US 20150170055A1 · Beymer et al. · 2015 [cited by applicant]
US 20150199010A1 · Coleman et al. · 2015 [cited by applicant]
US 20160055307A1 · Macoviak et al. · 2016 [cited by applicant]
US 20160055410A1 · Spagnola · 2016 [cited by applicant]
US 20160055427A1 · Adjaoute · 2016 [cited by applicant]
US 20160071017A1 · Adjaoute · 2016 [cited by applicant]
US 20160078367A1 · Adjaoute · 2016 [cited by applicant]
US 20160149862A1 · Kilgallon · 2016 [cited by examiner]
US 20160210427A1 · Mynhier · 2016 [cited by examiner]
US 20160300252A1 · Frank · 2016 [cited by examiner]
US 20180018590A1 · Szeto et al. · 2018 [cited by applicant]
US 20220405644A1 · Szeto et al. · 2022 [cited by applicant]
CA 2795554A1 · 2011 [cited by applicant]
CN 102413872A · 2012 [cited by applicant]
DE 102005020618A1 · 2005 [cited by applicant]
JP 2013523154A · 2013 [cited by applicant]
WO WO2010126624A1 · 2010 [cited by applicant]
WO WO2010126625A1 · 2010 [cited by applicant]
WO WO2011127150A2 · 2011 [cited by applicant]
WO WO2015149035A1 · 2015 [cited by applicant]
Evan et al., “Create Once, Use Many Times: The Clever Use of Recordkeeping Metadata for Multiple Archival Purposes,” 15th Int'l Congress on Archives (2004) (Year: 2004). [cited by examiner]
Johnson et al., “On Compressing Encrypted Data,” IEEE (2004) (Year: 2004). [cited by examiner]
Berkvosky et al., “Enhancing Privacy and Preserving Accuracy of a Distributed Collaborative Filtering,” RecSys '07 (2007) (Year: 2007). [cited by examiner]
Poh et al., “Challenges in Designing an Online Healthcare Platform for Personalised Patient Analytics,” IEEE (2014) [hereinafter Poh] (Year: 2014). [cited by examiner]
Shokri et al., “Preserving Privacy in Collaborative Filtering through Distributed Aggregation of Offline Profiles,” ACM (2009) (Year: 2009). [cited by examiner]
Berkvosky et al., “Enhancing Privacy and Preserving Accuracy of a Distributed Collaborative Filtering,” RecSys (2007) (Year: 2007). [cited by examiner]
Bareinboim, Elias et al.; Casual Inference and the Data-Fusion Problem; CrossMark Colloquium Paper; Jun. 29, 2015; 8 pages. [cited by applicant]
Daemen, Anneleen et al.; Development of a Kernel Function for Clinical Data; 31st Annual International Conference of the IEEE EMBS; Sep. 2009; 5 pages. [cited by applicant]
Feldman et al.; Dimensionality Reduction of Massive Sparse Datasets Using Coresets; Mar. 5, 2015; 11 pages. [cited by applicant]
Hardesty; Big data technique shrinks data sets while preserving their fundamental mathematical relationships; Dec. 15, 2016; 5 pages. [cited by applicant]
He, B., et al., “CRFs based de-identification of medical records,” Journal of Biomedical Informatics, 58: S39-S40 (2015). [cited by applicant]
Huang et al.; Patient Clustering Improves Efficiency of Federated Machine Learning to Predict Mortality and Hospital Stay Time using Distributed Electronic Medical Records; Mar. 22, 2019; 21 pages. [cited by applicant]
Intemational Search Report and Written Opinion in counterpart International Application No. PCT/US17/42356, mailed Sep. 26, 2017 (9 pages). [cited by applicant]
International Preliminary Report on Patentability for PCT Application No. PCT/US17/42356; Sep. 26, 2017; 7 pages. [cited by applicant]
International Search Report and Written Opinion for PCT Application No. PCT/US17/42356; Sep. 26, 2017; 8 pages. [cited by applicant]
Knight; How AI could save lives without spilling medical secrets; May 14, 2019; 3 pages. [cited by applicant]
Koperniak; Artificial data give the same results as real data—without compromising privacy; Mar. 6, 2017; 3 pages. [cited by applicant]
McMahan et al.; Federated Learning: Collaborative Machine Learning without Centralized Training Data; Google AI Blog; Apr. 6, 2017. [cited by applicant]
Mitra, A., et al., “Eigen-Profiles of Spation-Temporal Fragments for Adaptive Region-Based Tracking,” ICASSP, pp. 1497-1500 (2012). [cited by applicant]
Office Action for Canadian Application No. 3,031,067; dated Jan. 6, 2020; 7 pages. [cited by applicant]
Office Action for Japanese Application No. 2019-502045; dates Jun. 17, 2020; 2 pages. [cited by applicant]
Office Action for Japanese Application No. 2019-502045; dates Mar. 10, 2020; 2 pages. [cited by applicant]
Office Action from corresponding U.S. Appl. No. 15/651,345 dated Jan. 13, 2020. [cited by applicant]
Office Action from corresponding U.S. Appl. No. 15/651,345 dated Mar. 16, 2021. [cited by applicant]
Office Action from corresponding U.S. Appl. No. 15/651,345 dated Nov. 16, 2020. [cited by applicant]
Office Action from corresponding U.S. Appl. No. 15/651,345 dated Nov. 17, 2021. [cited by applicant]
Office Action from corresponding U.S. Appl. No. 15/651,345 dated Aug. 22, 2019. [cited by applicant]
Office Action from corresponding U.S. Appl. No. 15/651,345 dated Jul. 28, 2021. [cited by applicant]
OneLook Online Thesaurus, https://onelook.com/thesaurus/?s=transmit, 2021. [cited by applicant]
Patki et al.; The Synthetic data vault; Oct. 17, 2016; 12 pages. [cited by applicant]
Poh, N., et al., “Challenges in Designing an Online Healthcare Platform for Personalised Patient Analytics,” IEEE, pp. 1-6 (2014). [cited by applicant]
Stojanovic, J., et al., “Modeling Healthcare Quality via Compact Representations of Electronic Health Records,” IEEE, p. 1-10 (2016). [cited by applicant]
Wiggers; Federated Learning Technique Predicts Hospital Stay and Patient Mortality; Mar. 25, 2019; 2 pages. [cited by applicant]
Office Action from corresponding U.S. Appl. No. 17/890,953 dated Dec. 27, 2022. [cited by applicant]
Nandapalan, N., et al., “High-Performance Pseudo-Random Number Generation on Graphics Processing Units,” arXiv (2011). [cited by applicant]