IP Library › Granted Patent US 12,198,029
Granted Patent B2
US 12,198,029 · App. 17/210,216 · Granted Jan 14, 2025

Joint training method and apparatus for models, device and storage medium

Inventors: Chuanyuan Song (Beijing, CN); Zhi Feng (Beijing, CN); Liangliang Lyu (Beijing, CN)
Assignee: Beijing Baidu Netcom Science and Technology Co., Ltd.
G06N20/20H04L9/008
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,029
App. No.
17/210,216
Granted
Jan 14, 2025
Kind
B2
Abstract

The present disclosure provides a joint training method and apparatus for models, a device and a storage medium. The method may include: training a first-party model to be trained using a first sample quantity of first-party training samples to obtain first-party feature gradient information; acquiring second-party feature gradient information and second sample quantity information from a second party, where the second-party feature gradient information is obtained by training, by the second party, a second-party model to be trained using a second sample quantity of second-party training samples; and determining model joint gradient information according to the first-party feature gradient information, the second-party feature gradient information, first sample quantity information and the second sample quantity information, and updating the first-party model and the second-party model according to the model joint gradient information.

Claims (51)

1. A joint training method for models, comprising:

training a first-party model to be trained using a first sample quantity of first-party training samples to obtain first-party feature gradient information;

acquiring second-party feature gradient information and second sample quantity information from a second party, wherein the second-party feature gradient information is obtained by training, by the second party, a second-party model to be trained using a second sample quantity of second-party training samples;

determining a model joint gradient cipher text according to a first-party feature gradient cipher text, a second-party feature gradient cipher text, a first sample quantity cipher text and a second sample quantity cipher text; wherein the first-party feature gradient cipher text and the first sample quantity cipher text are obtained by respectively encrypting, by the first party, the first-party feature gradient plain text and the first sample quantity plain text using a first homomorphic encryption key acquired from the second party;

wherein the second-party feature gradient cipher text and the second sample quantity cipher text are obtained by respectively encrypting, by the second party, the second-party feature gradient plain text and the second sample quantity plain text using the first homomorphic encryption key, and wherein the first sample quantity is a sample quantity of the first-party training samples, and the second sample quantity is a sample quantity of the second-party training samples; and

updating the first-party model and the second-party model according to the model joint gradient information.

2. The method according to claim 1 , wherein the first-party training samples are randomly drawn by the first party from first-party sample data, and a quantity of samples randomly drawn from the first-party sample data is used as the first sample quantity; the second-party training samples are randomly drawn by the second party from second-party sample data, and a quantity of samples randomly drawn from the second-party sample data is used as the second sample quantity.

3. The method according to claim 1 , wherein the updating the first-party model and the second-party model according to the model joint gradient information comprises:

sending the model joint gradient cipher text to the second party for instructing the second party to perform operations comprising: decrypting the model joint gradient cipher text using a second homomorphic encryption key to obtain the model joint gradient plain text; and updating the second-party model according to the model joint gradient plain text; and

acquiring the model joint gradient plain text from the second party, and updating the first-party model according to the model joint gradient plain text.

4. The method according to claim 1 , wherein data structures of the first-party training samples and the second-party training samples are same.

5. The method according to claim 4 , further comprising:

determining common feature dimensions between the first-party sample data and the second-party sample data through a security intersection protocol; and

determining the first-party training samples and the second-party training samples based on the common feature dimensions.

6. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used for a computer to execute the method according to claim 1 .

7. A joint training method for models, comprising:

training a second-party model to be trained using a second sample quantity of second-party training samples to obtain second-party feature gradient information; and

sending second sample quantity information and the second-party feature gradient information to a first party for instructing the first party to perform operations comprising:

training a first-party model to be trained using a first sample quantity of first-party training samples to obtain first-party feature gradient information; and determining model joint gradient information according to the first-party feature gradient information, the second-party feature gradient information, first sample quantity information and the second sample quantity information, and updating the first-party model and the second-party model according to the model joint gradient information;

wherein the method further comprises: determining a first homomorphic encryption key and a second homomorphic encryption key, and sending the first homomorphic encryption key to the first party; wherein the sending second sample quantity information and the second-party feature gradient information to the first party comprises:

sending a second sample quantity cipher text and a second-party feature gradient cipher text to the first party for instructing the first party to perform operations comprising:

determining a model joint gradient cipher text according to a first-party feature gradient cipher text, the second-party feature gradient cipher text, a first sample quantity cipher text and the second sample quantity cipher text; wherein the first-party feature gradient cipher text and the first sample quantity cipher text are obtained by respectively encrypting, by the first party, the first-party feature gradient plain text and the first sample quantity plain text using the first homomorphic encryption key; wherein the second-party feature gradient cipher text and the second sample quantity cipher text are obtained by respectively encrypting, by the second party, the second-party feature gradient plain text and the second sample quantity plain text using the first homomorphic encryption key, and wherein the first sample quantity is a sample quantity of the first-party training samples, and the second sample quantity is a sample quantity of the second-party training samples.

8. The method according to claim 7 , wherein the first-party training samples are randomly drawn by the first party from first-party sample data, and a quantity of samples randomly drawn from the first-party sample data is used as the first sample quantity; the second-party training samples are randomly drawn by the second party from second-party sample data, and a quantity of samples randomly drawn from the second-party sample data is used as the second sample quantity.

9. The method according to claim 7 , wherein the updating the first-party model and the second-party model according to the model joint gradient information comprises:

acquiring the model joint gradient cipher text from the first party, decrypting the model joint gradient cipher text using the second homomorphic encryption key to obtain the model joint gradient plain text, and updating the second-party model according to the model joint gradient plain text; and

sending the model joint gradient plain text to the first party for instructing the first party to update the first-party model according to the model joint gradient plain text.

10. The method according to claim 7 , wherein data structures of the first-party training samples and the second-party training samples are same.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used for a computer to execute the method according to claim 7 .

12. An electronic device, comprising:

at least one processor; and

a memory communicatively connected with the at least one processor;

wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to perform operations, comprising:

training a first-party model to be trained using a first sample quantity of first-party training samples to obtain first-party feature gradient information;

acquiring second-party feature gradient information and second sample quantity information from a second party, wherein the second-party feature gradient information is obtained by training, by the second party, a second-party model to be trained using a second sample quantity of second-party training samples;

determining a model joint gradient cipher text according to a first-party feature gradient cipher text, a second-party feature gradient cipher text, a first sample quantity cipher text and a second sample quantity cipher text; wherein the first-party feature gradient cipher text and the first sample quantity cipher text are obtained by respectively encrypting, by the first party, the first-party feature gradient plain text and the first sample quantity plain text using a first homomorphic encryption key acquired from the second party;

wherein the second-party feature gradient cipher text and the second sample quantity cipher text are obtained by respectively encrypting, by the second party, the second-party feature gradient plain text and the second sample quantity plain text using the first homomorphic encryption key, and wherein the first sample quantity is a sample quantity of the first-party training samples, and the second sample quantity is a sample quantity of the second-party training samples; and

updating the first-party model and the second-party model according to the model joint gradient information.

13. The electronic device according to claim 12 , wherein the first-party training samples are randomly drawn by the first party from first-party sample data, and a quantity of samples randomly drawn from the first-party sample data is used as the first sample quantity; the second-party training samples are randomly drawn by the second party from second-party sample data, and a quantity of samples randomly drawn from the second-party sample data is used as the second sample quantity.

14. The electronic device according to claim 12 , wherein the updating the first-party model and the second-party model according to the model joint gradient information comprises:

sending the model joint gradient cipher text to the second party for instructing the second party to perform operations comprising: decrypting the model joint gradient cipher text using a second homomorphic encryption key to obtain the model joint gradient plain text; and updating the second-party model according to the model joint gradient plain text; and

acquiring the model joint gradient plain text from the second party, and updating the first-party model according to the model joint gradient plain text.

15. An electronic device, comprising:

at least one processor; and

a memory communicatively connected with the at least one processor;

wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to perform operations, comprising:

training a second-party model to be trained using a second sample quantity of second-party training samples to obtain second-party feature gradient information; and

sending second sample quantity information and the second-party feature gradient information to a first party for instructing the first party to perform operations comprising:

training a first-party model to be trained using a first sample quantity of first-party training samples to obtain first-party feature gradient information; and determining model joint gradient information according to the first-party feature gradient information, the second-party feature gradient information, first sample quantity information and the second sample quantity information, and updating the first-party model and the second-party model according to the model joint gradient information;

wherein the operations further comprise: determining a first homomorphic encryption key and a second homomorphic encryption key, and sending the first homomorphic encryption key to the first party; wherein the sending second sample quantity information and the second-party feature gradient information to the first party comprises:

sending a second sample quantity cipher text and a second-party feature gradient cipher text to the first party for instructing the first party to perform operations comprising: determining a model joint gradient cipher text according to a first-party feature gradient cipher text, the second-party feature gradient cipher text, a first sample quantity cipher text and the second sample quantity cipher text; wherein the first-party feature gradient cipher text and the first sample quantity cipher text are obtained by respectively encrypting, by the first party, the first-party feature gradient plain text and the first sample quantity plain text using the first homomorphic encryption key; wherein the second-party feature gradient cipher text and the second sample quantity cipher text are obtained by respectively encrypting, by the second party, the second-party feature gradient plain text and the second sample quantity plain text using the first homomorphic encryption key, and wherein the first sample quantity is a sample quantity of the first-party training samples, and the second sample quantity is a sample quantity of the second-party training samples.

16. The electronic device according to claim 15 , wherein the first-party training samples are randomly drawn by the first party from first-party sample data, and a quantity of samples randomly drawn from the first-party sample data is used as the first sample quantity; the second-party training samples are randomly drawn by the second party from second-party sample data, and a quantity of samples randomly drawn from the second-party sample data is used as the second sample quantity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2021
From: SONG, CHUANYUAN; FENG, ZHI; LYU, LIANGLIANG
To: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
Reel/Frame 056798/0545 →
Priority Claims (1)
CN 202011025747.4 · Sep 25, 2020 · national
Continuity (1)
Related Publication 20210209515A1 · Jul 8, 2021
References Cited (22)
US 20170310643A1 · Hardy et al. · 2017 [cited by applicant]
US 20210097430A1 · Das · 2021 [cited by examiner]
US 20210150269A1 · Choudhury · 2021 [cited by examiner]
US 20210312334A1 · Liu et al. · 2021 [cited by applicant]
US 20230325529A1 · Sav · 2023 [cited by examiner]
CN 109165725A · 2019 [cited by applicant]
CN 109325584A · 2019 [cited by applicant]
CN 109886417A · 2019 [cited by applicant]
CN 111177768A · 2020 [cited by applicant]
CN 111428887A · 2020 [cited by applicant]
CN 111461215A · 2020 [cited by applicant]
CN 111523686A · 2020 [cited by applicant]
Song, Tianshu, Yongxin Tong, and Shuyue Wei. “Profit allocation for federated learning.” 2019 IEEE International Conference on Big Data (Big Data). IEEE, 2019. (Year: 2019). [cited by examiner]
Song, Lei, et al. “Privacy-preserving unsupervised domain adaptation in federated setting.” IEEE Access 8 (2020): 143233-143240. (Year: 2020). [cited by examiner]
Roy, Abhijit Guha, et al. “Braintorrent: A peer-to-peer environment for decentralized federated learning.” arXiv preprint arXiv: 1905.06731 (2019). (Year: 2019). [cited by examiner]
European Patent Office; Extended European Search Report for Application No. 21164246.7, dated Jan. 18, 2022 (43 pages). [cited by applicant]
Shengwen Yang et al.; “Parallel Distributed Logistic Regression for Vertical Federated Learning Without Third-Party Coordinator”; Arxiv.org, Cornell University Library, 201 Olin Library Cornell University, Ithaca, New Y… [cited by applicant]
Song Lei et al.; “Privacy-Preserving Unsupervised Domain Adaptation in Federated Setting”; IEEE Access, IEEE, USA; vol. 8, pp. 143233-143240; Aug. 3, 2020; XP011804655; DOI: 10.1109/ACCESS.2020.3014264 (8 pages). [cited by applicant]
Yang Qiang Qyang@CSE UST HK et al.; “Federated Machine Learning”; ACM Transactions on Intelligent Systems and Technology; Association for Computing Machinery Corporation, 2 Penn Plaza, Suite 701, New York, New York USA,… [cited by applicant]
Chen, G., et al., “Implementation of communication fraud identification model based on federated learning”, Telecommun. Sci., 2020, 36(S1), 300-306, English abstract only. [cited by applicant]
Jiang, J., et al., “Decentralised federated learning with adaptive partial gradient aggregation”, CAAI Trans. Intell. Technol., 2020, 5(3):230-236. [cited by applicant]
Zhou, J., et al., “Survey on Security and Privacy-preserving in Federated Learning”, Journal of Xihua University (Natural Science Edition), Jul. 2020, 39(4):9-17, English abstract only. [cited by applicant]