IP Library › Granted Patent US 11,449,805
Granted Patent B2
US 11,449,805 · App. 17/498,978 · Granted Sep 20, 2022

Target data party selection methods and systems for distributed model training

Inventors: Longfei Zheng (Hangzhou, CN); Chaochao Chen (Hangzhou, CN); Yinggui Wang (Hangzhou, CN); Li Wang (Hangzhou, CN); Jun Zhou (Hangzhou, CN)
Assignee: Alipay (Hangzhou) Information Technology Co., Ltd.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,449,805
App. No.
17/498,978
Granted
Sep 20, 2022
Kind
B2
Abstract

A computer-implemented method, medium, and system are disclosed. One example computer-implemented method performed by a server includes obtaining training task information from a task party. The training task information includes information about a to-be-pretrained model and information about a to-be-trained target model. A respective task acceptance indication from each of at least one of a plurality of data parties is received to obtain a candidate data party set. The information about the to-be-pretrained model is sent to each data party in the candidate data party set. A respective pre-trained model of each data party is received. A respective performance parameter of the respective pre-trained model of each data party is obtained. One or more target data parties from the candidate data party set is determined. The information about the to-be-trained target model is sent to the one or more target data parties to obtain a target model.

Claims (66)

1. A computer-implemented method, comprising:

obtaining, by a server, training task information from a task party, wherein the training task information comprises information about a to-be-pretrained model and information about a to-be-trained target model;

receiving, by the server, a respective task acceptance indication from each of at least one of a plurality of data parties, to obtain a candidate data party set;

sending, by the server, the information about the to-be-pretrained model to each data party in the candidate data party set;

receiving, by the server, a respective pre-trained model of each data party, wherein the respective pre-trained model of each data party is obtained by each data party through model training based on respective training samples of each data party and the information about the to-be-pretrained model;

obtaining, for each data party, a respective first performance parameter of the respective pre-trained model of a respective data party;

testing, for each data party, the respective pre-trained model based on a test set, and obtaining a respective second performance parameter of the respective pre-trained model, wherein the test set comprises a plurality of test samples;

obtaining, for each data party, a respective overfitting parameter of the respective pre-trained model based on the respective first performance parameter and the respective second performance parameter of the respective pre-trained model;

determining one or more target data parties from the candidate data party set based on at least the respective overfitting parameter of the respective pre-trained model, wherein each target data party of the one or more target data parties participates in distributed model training to obtain a target model; and

sending the information about the to-be-trained target model to each target data party of the one or more target data parties, to obtain the target model through cooperative training among the one or more target data parties.

2. The computer-implemented method of claim 1 , wherein the plurality of test samples in the test set are from one or more data parties, or the test set is from the task party.

3. The computer-implemented method of claim 1 , wherein the training task information further comprises a performance screening threshold, and wherein the determining one or more target data parties from the candidate data party set based on at least the respective overfitting parameter of the respective pre-trained model comprises:

determining the one or more target data parties from the candidate data party set based on the respective overfitting parameter of the respective pre-trained model and the performance screening threshold.

4. The computer-implemented method of claim 3 , wherein determining the one or more target data parties from the candidate data party set based on the respective overfitting parameter of the respective pre-trained model and the performance screening threshold comprises:

comparing the respective overfitting parameter of the respective pre-trained model with the performance screening threshold;

sorting, in descending order, overfitting parameters of pre-trained models whose comparison results satisfy a predetermined condition; and

determining, as the one or more target data parties, data parties corresponding to a first N pre-trained models associated with a first N sorted overfitting parameters, wherein N is an integer greater than 0.

5. The computer-implemented method of claim 1 , wherein the training task information further comprises a total task reward.

6. The computer-implemented method of claim 5 , further comprising:

obtaining, from each target data party, a respective quantity of training samples used for model training; and

determining a respective task reward for each target data party based on the respective quantity of training samples of each target data party and the total task reward.

7. The computer-implemented method of claim 1 , wherein the training task information further comprises description information of the respective training samples of each data party in the candidate data party set, wherein the method further comprises receiving data description information of the plurality of data parties from the plurality of data parties, and

wherein receiving, by the server, a respective task acceptance indication from each of at least one of a plurality of data parties, to obtain a candidate data party set further comprises:

determining, based on the description information of the respective training samples in the training task information and data description information of each data party that sends the respective task acceptance indication, whether each data party that sends the respective task acceptance indication is a data party in the candidate data party set.

8. A non-transitory, computer-readable medium storing one or more instructions executable by a computer system to perform operations comprising:

obtaining, by a server, training task information from a task party, wherein the training task information comprises information about a to-be-pretrained model and information about a to-be-trained target model;

receiving, by the server, a respective task acceptance indication from each of at least one of a plurality of data parties, to obtain a candidate data party set;

sending, by the server, the information about the to-be-pretrained model to each data party in the candidate data party set;

receiving, by the server, a respective pre-trained model of each data party, wherein the respective pre-trained model of each data party is obtained by each data party through model training based on respective training samples of each data party and the information about the to-be-pretrained model;

obtaining, for each data party, a respective first performance parameter of the respective pre-trained model of a respective data party;

testing, for each data party, the respective pre-trained model based on a test set, and obtaining a respective second performance parameter of the respective pre-trained model, wherein the test set comprises a plurality of test samples;

obtaining, for each data party, a respective overfitting parameter of the respective pre-trained model based on the respective first performance parameter and the respective second performance parameter of the respective pre-trained model;

determining one or more target data parties from the candidate data party set based on at least the respective overfitting parameter of the respective pre-trained model, wherein each target data party of the one or more target data parties participates in distributed model training to obtain a target model; and

sending the information about the to-be-trained target model to each target data party of the one or more target data parties, to obtain the target model through cooperative training among the one or more target data parties.

9. The non-transitory, computer-readable medium of claim 8 , wherein the plurality of test samples in the test set are from one or more data parties, or the test set is from the task party.

10. The non-transitory, computer-readable medium of claim 8 , wherein the training task information further comprises a performance screening threshold, and wherein the determining one or more target data parties from the candidate data party set based on at least the respective overfitting parameter of the respective pre-trained model comprises:

determining the one or more target data parties from the candidate data party set based on the respective overfitting parameter of the respective pre-trained model and the performance screening threshold.

11. The non-transitory, computer-readable medium of claim 10 , wherein determining the one or more target data parties from the candidate data party set based on the respective overfitting parameter of the respective pre-trained model and the performance screening threshold comprises:

comparing the respective overfitting parameter of the respective pre-trained model with the performance screening threshold;

sorting, in descending order, overfitting parameters of pre-trained models whose comparison results satisfy a predetermined condition; and

determining, as the one or more target data parties, data parties corresponding to a first N pre-trained models associated with a first N sorted overfitting parameters, wherein N is an integer greater than 0.

12. The non-transitory, computer-readable medium of claim 8 , wherein the training task information further comprises description information of the respective training samples of each data party in the candidate data party set, wherein the operations further comprise receiving data description information of the plurality of data parties from the plurality of data parties, and

wherein receiving, by the server, a respective task acceptance indication from each of at least one of a plurality of data parties, to obtain a candidate data party set further comprises:

determining, based on the description information of the respective training samples in the training task information and data description information of each data party that sends the respective task acceptance indication, whether each data party that sends the respective task acceptance indication is a data party in the candidate data party set.

13. A computer-implemented system, comprising:

one or more computers; and

one or more computer memory devices interoperably coupled with the one or more computers and having tangible, non-transitory, machine-readable media storing one or more instructions that, when executed by the one or more computers, perform one or more operations comprising:

obtaining, by a server, training task information from a task party, wherein the training task information comprises information about a to-be-pretrained model and information about a to-be-trained target model;

receiving, by the server, a respective task acceptance indication from each of at least one of a plurality of data parties, to obtain a candidate data party set;

sending, by the server, the information about the to-be-pretrained model to each data party in the candidate data party set;

receiving, by the server, a respective pre-trained model of each data party, wherein the respective pre-trained model of each data party is obtained by each data party through model training based on respective training samples of each data party and the information about the to-be-pretrained model;

obtaining, for each data party, a respective first performance parameter of the respective pre-trained model of a respective data party;

testing, for each data party, the respective pre-trained model based on a test set, and obtaining a respective second performance parameter of the respective pre-trained model, wherein the test set comprises a plurality of test s ample s;

obtaining, for each data party, a respective overfitting parameter of the respective pre-trained model based on the respective first performance parameter and the respective second performance parameter of the respective pre-trained model;

determining one or more target data parties from the candidate data party set based on at least the respective overfitting parameter of the respective pre-trained model, wherein each target data party of the one or more target data parties participates in distributed model training to obtain a target model; and

sending the information about the to-be-trained target model to each target data party of the one or more target data parties, to obtain the target model through cooperative training among the one or more target data parties.

14. The computer-implemented system of claim 13 , wherein the plurality of test samples in the test set are from one or more data parties, or the test set is from the task party.

15. The computer-implemented system of claim 13 , wherein the training task information further comprises a performance screening threshold, and wherein the determining one or more target data parties from the candidate data party set based on at least the respective overfitting parameter of the respective pre-trained model comprises:

determining the one or more target data parties from the candidate data party set based on the respective overfitting parameter of the respective pre-trained model and the performance screening threshold.

16. The computer-implemented system of claim 15 , wherein determining the one or more target data parties from the candidate data party set based on the respective overfitting parameter of the respective pre-trained model and the performance screening threshold comprises:

comparing the respective overfitting parameter of the respective pre-trained model with the performance screening threshold;

sorting, in descending order, overfitting parameters of pre-trained models whose comparison results satisfy a predetermined condition; and

determining, as the one or more target data parties, data parties corresponding to a first N pre-trained models associated with a first N sorted overfitting parameters, wherein N is an integer greater than 0.

17. The computer-implemented system of claim 13 , wherein the training task information further comprises description information of the respective training samples of each data party in the candidate data party set, wherein the one or more operations further comprise receiving data description information of the plurality of data parties from the plurality of data parties, and

wherein receiving, by the server, a respective task acceptance indication from each of at least one of a plurality of data parties, to obtain a candidate data party set further comprises:

determining, based on the description information of the respective training samples in the training task information and data description information of each data party that sends the respective task acceptance indication, whether each data party that sends the respective task acceptance indication is a data party in the candidate data party set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2021
From: ZHENG, LONGFEI; CHEN, CHAOCHAO; WANG, YINGGUI; WANG, LI; ZHOU, JUN
To: ALIPAY (HANGZHOU) INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 058163/0900 →
Priority Claims (1)
CN 202011082434.2 · Oct 12, 2020 · national
Continuity (1)
Related Publication 20220114492A1 · Apr 14, 2022
Cited By (2)
US 12,229,280 US 12,682,273