IP Library Granted Patent US 12670445
Granted Patent B2
US 12670445 · App. 18/633,486 · Granted Jun 30, 2026

Removing biases within a distributed model

Inventors: Jason Crabtree (Vienna, VA); Andrew Sellers (Monument, CO)
Assignee: QOMPLX LLC
G06N20/00G06F16/215G06F16/951G06F18/214G06F18/24G06N5/022H04L67/10G06F18/29
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670445
App. No.
18/633,486
Granted
Jun 30, 2026
Kind
B2
Abstract

Improving a distributed model with distributed data receives an instance of the distributed model from the network-connected computing system, create a cleansed dataset from data stored in the memory with at least biases within the data stored in memory corrected, train the instance of the distributed model with the cleansed dataset, and generate an update report based at least in part by updates to the instance of the distributed model.

Claims (80)

1 . A system for improving a distributed model with distributed data, comprising one or more computers with executable instructions that, when executed, cause the system to:

receive an instance of a distributed model via a computer network;

locally process data on a computing device to maintain data privacy, wherein the local processing comprises:

locally identifying a plurality of data biases within the instance of the distributed model, wherein the plurality of data biases are identified by computing differences between data stored locally on the computing device and data of the distributed model;

creating a modified dataset based on the distributed model and one or more identified biases of the identified plurality of data biases, wherein creating the modified dataset comprises at least one of:

removing one or more of the identified biases from the data stored locally on the computing device; or

generating synthetic data that approximates characteristics of the data stored locally on the computing device while reducing one or more of the identified biases;

wherein one or more of the identified biases are automatically weighted and corrected by a local machine learning algorithm to generalize the modified dataset for use in the distributed model while adhering to local laws and data handling requirements;

locally training the instance of the distributed model using the modified dataset; and

generating an update report based upon which remote instances of the distributed model may be adjusted to use the modified dataset, wherein the update report is structured to enable integration with the distributed model while maintaining local data requirements and reducing network resource utilization compared to centralized model training approaches; and

transfer the update report to the distributed model to improve the distributed model via the computer network.

2 . The system of claim 1 , wherein at least a portion of the modified dataset is data that has had sensitive information removed.

3 . The system of claim 1 , wherein the system is further caused to improve the distributed model using at least a portion of the update report.

4 . The system of claim 1 , wherein at least a portion of the data processed locally on the computing device is medical-related data.

5 . The system of claim 1 , wherein at least a portion of the data processed locally on the computing device is crime-related data.

6 . The system of claim 1 , wherein at least a portion of the data processed locally on the computing device is banking-related data.

7 . The system of claim 3 , wherein improving the distributed model comprises classifying an incoming update report based at least in part by geographical origin of the incoming update report.

8 . The system of claim 7 , wherein the update report is used to improve a distributed model specific to the geographical origin.

9 . The system of claim 8 , wherein the distributed model specific to the geographical origin is updated using only update reports from the geographical origin.

10 . The system of claim 1 , wherein the modified dataset is created by removing biased data.

11 . A method for improving a distributed model with distributed data, comprising the steps of:

receiving an instance of a distributed model via a computer network;

locally processing data on a computing device to maintain data privacy, wherein the local processing comprises:

locally identifying a plurality of data biases within the instance of the distributed model, wherein the plurality of data biases are identified by computing differences between data stored locally on the computing device and data of the distributed model;

creating a modified dataset based on the distributed model and one or more identified biases of the identified plurality of data biases, wherein creating the modified dataset comprises at least one of:

removing one or more of the identified biases from the data stored locally on the computing device; or

generating synthetic data that approximates characteristics of the data stored locally on the computing device while reducing one or more of the identified biases;

wherein one or more of the identified biases are automatically weighted and corrected by a local machine learning algorithm to generalize the modified dataset for use in the distributed model while adhering to local laws and data handling requirements;

locally training the instance of the distributed model using the modified dataset; and

generating an update report based upon which remote instances of the distributed model may be adjusted to use the modified dataset, wherein the update report is structured to enable integration with the distributed model while maintaining local data requirements and reducing network resource utilization compared to centralized model training approaches; and

transferring the update report to the distributed model to improve the distributed model via the computer network.

12 . The method of claim 11 , wherein at least a portion of the modified dataset is data that has had sensitive information removed.

13 . The method of claim 11 , further comprising the step of using a distributed model source to improve the distributed model using at least a portion of the update report.

14 . The method of claim 11 , wherein at least a portion of the data processed locally on the computing device is medical-related data.

15 . The method of claim 11 , wherein at least a portion of the data processed locally on the computing device is crime-related data.

16 . The method of claim 11 , wherein at least a portion of the data processed locally on the computing device is banking-related data.

17 . The method of claim 13 , wherein the distributed model source categorizes an incoming update report based at least in part by geographical origin of the incoming update report.

18 . The method of claim 17 , wherein the update report is used to improve a distributed model specific to the geographical origin.

19 . The method of claim 18 , wherein the distributed model specific to the geographical origin is modified using only update reports from the geographical origins.

20 . The method of claim 11 , wherein the modified dataset is created by removing biased data.

21 . A computing system for improving a distributed model with distributed data, the computing system comprising:

one or more hardware processors configured for:

receiving an instance of a distributed model via a computer network;

locally processing data on a computing device to maintain data privacy, wherein the local processing comprises:

locally identifying a plurality of data biases within the instance of the distributed model, wherein the plurality of data biases are identified by computing differences between data stored locally on the computing device and data of the distributed model;

creating a modified dataset based on the distributed model and one or more identified biases of the identified plurality of biases, wherein creating the modified dataset comprises at least one of:

removing one or more of the identified biases from the data stored locally on the computing device; or

generating synthetic data that approximates characteristics of the data stored locally on the computing device while reducing one or more of the identified biases;

wherein one or more of the identified biases are automatically weighted and corrected by a local machine learning algorithm to generalize the modified dataset for use in the distributed model while adhering to local laws and data handling requirements;

locally training the instance of the distributed model using the modified dataset; and

generating an update report based upon which remote instances of the distributed model may be adjusted to use the modified dataset, wherein the update report is structured to enable integration with the distributed model while maintaining local data requirements and reducing network resource utilization compared to centralized model training approaches; and

transferring the update report to the distributed model to improve the distributed model via the computer network.

22 . The computing system of claim 21 , wherein at least a portion of the modified dataset is data that has had sensitive information removed.

23 . The computing system of claim 21 , wherein the one or more hardware processors are further configured for using a distributed model source to improve the distributed model using at least a portion of the update report.

24 . The computing system of claim 21 , wherein at least a portion of the data processed locally on the computing device is medical-related data.

25 . The computing system of claim 21 , wherein at least a portion of the data processed locally on the computing device is crime-related data.

26 . The computing system of claim 21 , wherein at least a portion of the data processed locally on the computing device is banking-related data.

27 . The computing system of claim 23 , wherein the distributed model source categorizes an incoming update report based at least in part by geographical origin of the incoming update report.

28 . The computing system of claim 27 , wherein the update report is used to improve a distributed model specific to the geographical origin.

29 . The computing system of claim 28 , wherein the distributed model specific to the geographical origin is modified using only update reports from the geographical origins.

30 . The computing system of claim 21 , wherein the modified dataset is created by removing biased data.

31 . Non-transitory, computer-readable storage media having computer-executable instructions embodied thereon that, when executed by one or more processors of a computing system for improving a distributed model with distributed data, cause the computing system to:

receive an instance of a distributed model via a computer network;

locally process data on a computing device to maintain data privacy, wherein the local processing comprises:

locally identifying a plurality of data biases within the instance of the distributed model, wherein the plurality of data biases are identified by computing differences between data stored locally on the computing device and data of the distributed model;

creating a modified dataset based on the distributed model and one or more identified biases of the identified plurality of data biases, wherein creating the modified dataset comprises at least one of:

removing one or more of the identified biases from the data stored locally on the computing device; or

generating synthetic data that approximates characteristics of the data stored locally on the computing device while reducing one or more of the identified biases;

wherein one or more of the identified biases are automatically weighted and corrected by a local machine learning algorithm to generalize the modified dataset for use in the distributed model while adhering to local laws and data handling requirements;

locally training the instance of the distributed model using the modified dataset; and

generating an update report based upon which remote instances of the distributed model may be adjusted to use the modified dataset, wherein the update report is structured to enable integration with the distributed model while maintaining local data requirements and reducing network resource utilization compared to centralized model training approaches; and

transfer the update report to the distributed model to improve the distributed model via the computer network.

32 . The non-transitory, computer-readable storage media of claim 31 , wherein at least a portion of the modified dataset is data that has had sensitive information removed.

33 . The non-transitory, computer-readable storage media of claim 31 , further wherein the system is further caused to improve the distributed model using at least a portion of the update report.

34 . The non-transitory, computer-readable storage media of claim 31 , wherein at least a portion of the data processed locally on the computing device is medical-related data.

35 . The non-transitory, computer-readable storage media of claim 31 , wherein at least a portion of the data processed locally on the computing device is crime-related data.

36 . The non-transitory, computer-readable storage media of claim 31 , wherein at least a portion of the data processed locally on the computing device is banking-related data.

37 . The non-transitory, computer-readable storage media of claim 36 , wherein the distributed model source classifies an incoming update report based at least in part by geographical origin of the incoming update report.

38 . The non-transitory, computer-readable storage media of claim 37 , wherein the update report is used to improve a distributed model specific to the geographical origin.

39 . The non-transitory, computer-readable storage media of claim 31 , wherein the modified dataset is created by removing biased data.