IP Library Granted Patent US 11,574,075
Granted Patent B2
US 11,574,075 · App. 16/432,712 · Granted Feb 7, 2023

Distributed machine learning technique used for data analysis and data computation in distributed environment

Inventors: Raajen Patel (Houston, TX); Vincent Gagne (Bellaire, TX); Christian Pesantes (Houston, TX)
Assignee: Medical Informatics Corp.
G06F21/6245G06F9/4881G06F16/2471G06N20/00G16H10/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,574,075
App. No.
16/432,712
Granted
Feb 7, 2023
Kind
B2
Abstract

A system for distributed computing allows researchers performing analysis on data having legal or policy restrictions on movement of the data to perform the analysis on data at multiple sites without exporting the restricted data from each site. A primary system determines which sites contain the data to be executed by a compute job and sends the compute job to the associated site. Resulting calculated data can be exported to a primary collecting and recording system. The system allows the rapid development of analysis code for the compute jobs that can be executed on a local data set then passed to the primary system for distributed machine learning on the multiple sites without any changes to the code of the compute job.

Claims (35)

1. A system for hierarchical distributed machine learning, comprising:

a plurality of secondary distributed computing systems, each associated with an organization of a plurality of organizations, each holding personal health information data that cannot be exported from the associated organization;

a recording data collector system, outside of the plurality of organizations, programmed to aggregate job completion information received from the plurality of secondary distributed computing systems; and

a primary distributed computing system, outside of the plurality of organizations, comprising:

a primary compute job scheduler of a primary distributed computing system programmed to send compute jobs to each of the plurality of secondary distributed computing systems for execution on personal health information data held by the associated organization that cannot be exported from the associated organization,

wherein each compute job exports job completion information to the recording data collector system, and

wherein the job completion information comprises non-personal health information metadata calculated by the compute job.

2. The system for hierarchical distributed machine learning of claim 1 , wherein each of the plurality of secondary distributed computing systems comprises:

a plurality of informatics servers, each associated with a portion of the personal health information data held by the secondary distributed computing system;

a job scheduler, programmed to:

determine which of the plurality of informatics servers is associated with data needed by the compute job; and

schedule execution of the compute job on informatics servers determined to be associated with data needed by the compute job.

3. The system for hierarchical distributed machine learning of claim 1 , wherein the recording data collector system is the primary distributed computing system.

4. The system for hierarchical distributed machine learning of claim 1 , wherein a secondary distributed computing system of the plurality of secondary distributed computing systems is a primary distributed computing system for a second plurality of secondary computing systems.

5. A system for distributed machine learning, comprising:

a plurality of informatics servers, each associated with a corresponding database of personal health information data that is prohibited from being provided outside of an organization associated with that informatics server;

a compute job for performing machine learning on one or more desired databases of personal health information data held by the plurality of informatics servers;

a job scheduler, programmed to:

determine a subset of the plurality of informatics servers is associated with a database of the one or more desired databases of data;

send the compute job to the subset of the plurality of informatics servers;

instruct the subset of the plurality of informatics servers to execute the compute job; and

aggregate job completion information returned to the job scheduler from the compute job, wherein the job completion information is non-personal health information metadata calculated by the compute job.

6. The system of claim 5 , wherein the compute job executes in parallel on the subset of the plurality of informatics servers.

7. The system of claim 5 , wherein the compute job comprises code developed by a researcher on a researcher computer system that executes unchanged on the plurality of informatics servers.

8. The system of claim 5 , wherein the corresponding database associated with each of the plurality of informatics servers comprises a collection of files, and

wherein metadata about a meaning of each of the collection of files and a location of each of the collection of files is used by the job scheduler.

9. A method of performing distributed machine learning on restricted data, comprising:

receiving at a primary distributed computing system a compute job for performing machine learning from a researcher computing system;

determining which of a plurality of secondary distributed computing systems have data needed by the compute job, wherein each of the plurality of secondary distributed computing systems is associated with a different organization, wherein the primary distributed computing system is outside the organizations associated with the plurality of secondary distributed computing systems;

sending the compute job to secondary distributed computing systems of the plurality of secondary distributed computing platforms systems determined to have the data needed by the compute job, wherein the data needed by the compute job is personal health information data that is restricted from being exported outside the associated organization;

receiving by the primary distributed computing system job completion information from each of the secondary distributed computing platforms upon completion of the compute job, wherein the job completion information comprises non-personal health information metadata calculated by the compute job; and

aggregating the job completion information received from each of the secondary distributed computing platforms.

10. The method of claim 9 , further comprising:

creating the compute job on the researcher computing system.

11. The method of claim 9 , wherein receiving job completion information comprises receiving job completion information by a collecting and recording system.

Assignments (1)
SECURITY INTEREST Recorded Jun 13, 2025
From: MEDICAL INFORMATICS CORP.
To: WESTERN ALLIANCE BANK
Reel/Frame 071411/0019 →
Continuity (2)
Provisional Application 62680705 · Jun 5, 2018
Related Publication 20190370490A1 · Dec 5, 2019