IP Library Granted Patent US 11,675,921
Granted Patent B2
US 11,675,921 · App. 16/838,163 · Granted Jun 13, 2023

Device and method for secure private data aggregation

Inventors: James Reid Desmond Arthur (Shrewsbury, GB); Luke Anthony William Robinson (Shrewsbury, GB); Harry Richard Keen (Shrewsbury, GB); Garry Hill (Sevenoaks, GB)
Assignee: Hazy Limited
G06F21/6218G06N5/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,675,921
App. No.
16/838,163
Granted
Jun 13, 2023
Kind
B2
Abstract

A computing system for enabling the analysis of multiple raw data sets whilst protecting the privacy of information within the raw data sets, the system comprising a plurality of synthetic data generators and a data hub. Each synthetic data generator is configured to: access a corresponding raw data set stored in a corresponding one of a plurality of raw data stores; produce, based on the corresponding raw data set, a synthetic data generator model configured to generate a synthetic data set representative of the corresponding raw data set; and push synthetic information including at least one of the corresponding synthetic data set and the synthetic data generator model to the data hub. The data hub is configured to store the synthetic information received from the synthetic data generators for access by one or more clients for analysis. The system is configured such that the data hub cannot directly access the raw data sets and such that the synthetic data information can only be pushed from the synthetic data generators to the data hub.

Claims (39)

1. A computing system for enabling the analysis of multiple raw data sets whilst protecting the privacy of information within the raw data sets, the computing system comprising a plurality of synthetic data generators and a data hub, wherein:

each synthetic data generator comprises one or more processors configured to:

access a corresponding raw data set stored in a corresponding one of a plurality of raw data stores;

produce, based on the corresponding raw data set, a synthetic data generator model configured to generate a synthetic data set representative of the corresponding raw data set; and

push synthetic information including at least one of the corresponding synthetic data set and the synthetic data generator model to the data hub;

the data hub comprises:

memory configured to store the synthetic information received from the synthetic data generators for access by one or more clients for analysis; and

one or more processors configured to determine a relative contribution provided by one or more of the synthetic data generator models towards an objective, wherein determining the relative contribution provided by the one or more of the synthetic data generator models towards the objective comprises determining a difference in performance between a first and second model, wherein:

the first model is trained to achieve the objective based on first training data including the one or more synthetic data generator models or synthetic data generated by the one or more synthetic data generator models; and

the second model is trained to achieve the objective based on second training data that does not include the one or more synthetic data generator models or synthetic data generated by the one or more synthetic data generator models; and

the system is configured such that the data hub cannot directly access the raw data sets and such that the synthetic data information can only be pushed from the synthetic data generators to the data hub,

wherein each synthetic generator model is configured to generate its corresponding synthetic data set to comply with a corresponding privacy level relative to its corresponding raw data set, and

wherein each synthetic generator model is configured to generate its corresponding synthetic data such that the corresponding synthetic data set is differentially private according to the corresponding privacy level.

2. The system of claim 1 wherein each synthetic data generator is configured to update one or more parameters of its corresponding synthetic data generator model based on its corresponding raw data set.

3. The system of claim 1 wherein each synthetic data generator is limited to read only privileges with respect to its corresponding raw data store.

4. The system of claim 1 wherein at least one of the synthetic data generators is configured to push its corresponding synthetic data generator model to the data hub and the data hub is configured to, for each synthetic data generator model received, generate a corresponding synthetic data set.

5. The system of claim 1 further comprising the one or more clients, wherein the one or more clients are configured to access, from the data hub, synthetic data information originating from at least two of the synthetic data generators and to aggregate the accessed synthetic data information to determine one or more attributes shared across the accessed synthetic data information.

6. The system of claim 5 wherein accessing the synthetic data information originating from at the least two of the synthetic data generators comprises one or more of:

pulling at least two synthetic data sets from the data hub; and

pulling at least two synthetic data generator models from the data hub and, for each synthetic data generator model that has been pulled from the data hub, generating a corresponding synthetic data set using the synthetic data model.

7. The system of claim 5 wherein aggregating the accessed synthetic data information comprises training a machine learning system based on the accessed synthetic data information to determine one or more attributes of the corresponding synthetic data sets.

8. The system of claim 1 wherein determining the relative contribution provided by the one or more of the synthetic data generator models towards the objective comprises:

training the first model based on the first training data;

evaluating the performance of the first model with respect to the objective;

training the second model based on the second training data; and

evaluating the performance of the second model with respect to the objective.

9. The system of claim 1 further configured to determine, for each of a plurality of the synthetic data generator models, a relative contribution provided the synthetic data generator model towards the objective.

10. A computer-implemented method for enabling the analysis of multiple raw data sets whilst protecting the privacy of information within the raw data sets, the method comprising:

for each of a plurality of synthetic data generators:

accessing a corresponding raw data set stored in a corresponding one of a plurality of raw data stores;

producing, based on the corresponding raw data set, a synthetic data generator model configured to generate a synthetic data set representative of the corresponding raw data set; and

pushing synthetic information including at least one of the corresponding synthetic data set and the synthetic data generator model to a data hub;

storing at the data hub the synthetic information received from the synthetic data generators for access by one or more clients for analysis; and

configuring a network comprising the synthetic data generators and the data hub such that the data hub is prevented from directly accessing the raw data sets and synthetic data information can only be pushed from the synthetic data generators to the data hub; and

determining a relative contribution provided by one or more of the synthetic data generator models towards an objective, wherein determining the relative contribution provided by the one or more of the synthetic data generator models towards the objective comprises determining a difference in performance between a first and second model, wherein:

the first model is trained to achieve the objective based on first training data including the one or more synthetic data generator models or synthetic data generated by the one or more synthetic data generator models; and

the second model is trained to achieve the objective based on second training data that does not include the one or more synthetic data generator models or synthetic data generated by the one or more synthetic data generator models,

wherein each synthetic generator model is configured to generate its corresponding synthetic data set to comply with a corresponding privacy level relative to its corresponding raw data set, and

wherein each synthetic generator model is configured to generate its corresponding synthetic data such that the corresponding synthetic data set is differentially private according to the corresponding privacy level.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 20, 2024
From: HAZY LIMITED
To: SAS INSTITUTE INC.
Reel/Frame 069340/0083 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2022
From: HILL, GARRY
To: HAZY LIMITED
Reel/Frame 061764/0357 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2020
From: ARTHUR, JAMES REID DESMOND; ROBINSON, LUKE ANTHONY WILLIAM; KEEN, HARRY RICHARD
To: HAZY LIMITED
Reel/Frame 053059/0530 →
Continuity (1)
Related Publication 20210312064A1 · Oct 7, 2021
Cited By (2)
US 12,608,394 US 12,711,155