IP Library › Granted Patent US 12,393,874
Granted Patent B2
US 12,393,874 · App. 18/370,524 · Granted Aug 19, 2025

Data processing system and method

Inventors: Changzheng Zhang (Shenzhen, CN); Xiaolong Bai (Hangzhou, CN); Dandan Tu (Shenzhen, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06N20/00G06F17/16G06N7/00G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,393,874
App. No.
18/370,524
Granted
Aug 19, 2025
Kind
B2
Abstract

Embodiments of the present invention disclose a data processing apparatus. The apparatus is configured to: after calculating a set of gradient information of each parameter by using a sample data subset, delete the sample data subset, read a next sample data subset, calculate another set of gradient information of each parameter by using the next sample data subset, and accumulate a plurality of sets of calculated gradient information of each parameter, to obtain an update gradient of each parameter.

Claims (51)

1. A data processing system, wherein the data processing system comprises at least one processor; and

a memory coupled to the at least one processor to store program instructions, which when executed by the at least one processor, cause the at least one processor to:

in a process of one iterative operation, sequentially read a plurality of sample data subsets from a sample data set, wherein each sample data subset comprises at least one piece of sample data;

enter each read sample data subset into a machine learning model; and

calculate gradient information of each of a plurality of parameters of the machine learning model, wherein after a set of gradient information of each parameter is calculated using one sample data subset, another set of gradient information of each parameter is calculated using a next sample data subset, and wherein the machine learning model has an initialized global parameter or was updated in a last iterative operation;

in the process of the one iterative operation, accumulate a plurality of sets of gradient information of each parameter to obtain an update gradient of each parameter; and

send the update gradient of each parameter in the process of the one iterative operation, wherein the update gradient of each parameter is used to update the machine learning model.

2. The data processing system according to claim 1 , wherein the program instructions further cause the at least one processor to:

participate in a plurality of iterative operations after the one iterative operation, until the machine learning model converges or a specified quantity of iterations are completed in calculation;

in each of the plurality of iterative operations after the one iterative operation, repeat actions in the process of the one iterative operation; and in the one iterative operation and the plurality of iterative operations after the one iterative operation, after the machine learning model is updated using an update gradient obtained in an iterative operation, a next iterative operation is performed.

3. The data processing system according to claim 1 , wherein the program instructions further cause the at least one processor to:

for the plurality of sets of gradient information of each parameter that are obtained based on the plurality of read sample data subsets, accumulate a plurality of sets of gradient information of a same parameter to obtain the update gradient of each parameter.

4. The data processing system according to claim 1 , wherein the program instructions further cause the at least one processor to:

for one set of gradient information of each parameter obtained based on each sample data subset, accumulate one set of gradient information of a same parameter to obtain an accumulation gradient of each parameter, so that a plurality of accumulation gradients of each parameter are obtained based on the plurality of read sample data subsets, and

accumulate the plurality of accumulation gradients of each parameter to obtain the update gradient of each parameter.

5. The data processing system according to claim 1 , wherein the program instructions further cause the at least one processor to:

for the plurality of sets of gradient information of each parameter that are obtained based on the plurality of read sample data subsets, collect a plurality of sets of gradient information of a same parameter together, wherein the plurality of sets of gradient information of each parameter that are collected together are used as the update gradient of each parameter.

6. The data processing system according to claim 1 , wherein the program instructions further cause the at least one processor to:

in the process of the one iterative operation, sequentially read the plurality of sample data subsets from the sample data set;

enter each read sample data subset into the machine learning model;

read and use an intermediate calculation result to calculate the gradient information of each of the plurality of parameters of the machine learning model, wherein the intermediate calculation result is used as input information to calculate the gradient information, and after a set of gradient information of each parameter is calculated using one sample data subset, the sample data subset is deleted before a next sample data subset is read, and another set of gradient information of each parameter is calculated using the next sample data subset; and

after the intermediate calculation result is used, delete the intermediate calculation result, wherein an operation of deleting the intermediate calculation result needs to be completed before the next sample data subset is read.

7. A data processing method, comprising one iterative operation, and the one iterative operation comprises:

sequentially reading a plurality of sample data subsets from a sample data set, wherein each sample data subset comprises at least one piece of sample data;

entering each read sample data subset into a machine learning model; and

calculating gradient information of each of a plurality of parameters of the machine learning model, wherein after a set of gradient information of each parameter is calculated using one sample data subset, another set of gradient information of each parameter is calculated using a next sample data subset; and wherein the machine learning model has an initialized global parameter or is updated in a last iterative operation; and

accumulating a plurality of sets of gradient information of each parameter to obtain an update gradient of each parameter, wherein the update gradient of each parameter is used to update the machine learning model.

8. The method according to claim 7 , wherein the one iterative operation is iteratively performed in a plurality of iterative operations until the machine learning model converges or a specified quantity of iterations are completed in calculation, wherein

in each of the plurality of iterative operations after the one iterative operation, after the machine learning model is updated using an update gradient obtained in an iterative operation, a next iterative operation is performed.

9. The method according to claim 7 , further comprising updating the machine learning model using the update gradient of each parameter during each iteration operation.

10. The method according to claim 8 , wherein accumulating, in each iterative operation, a plurality of sets of gradient information of each parameter to obtain an update gradient of each parameter comprises:

for the plurality of sets of gradient information of each parameter that are obtained based on the plurality of read sample data subsets, accumulating a plurality of sets of gradient information of a same parameter to obtain the update gradient of each parameter.

11. The method according to claim 8 , wherein accumulating, in each iterative operation, a plurality of sets of gradient information of each parameter to obtain an update gradient of each parameter comprises:

for one set of gradient information of each parameter obtained based on each sample data subset, accumulating one set of gradient information of a same parameter to obtain an accumulation gradient of each parameter, so that a plurality of accumulation gradients of each parameter are obtained based on the plurality of read sample data subsets, and

accumulating the plurality of accumulation gradients of each parameter to obtain the update gradient of each parameter.

12. The method according to claim 8 , wherein accumulating, in each iterative operation, a plurality of sets of gradient information of each parameter to obtain an update gradient of each parameter comprises:

for the plurality of sets of gradient information of each parameter that are obtained based on the plurality of read sample data subsets, collecting a plurality of sets of gradient information of a same parameter together, wherein the plurality of sets of gradient information of each parameter that are collected together are used as the update gradient of each parameter.

13. The method according to claim 8 , wherein in a process of entering each sample data subset into the machine learning model and calculating gradient information of each of the plurality of parameters of the machine learning model during each iterative operation, one piece of gradient information of each parameter is obtained correspondingly based on one piece of sample data in the sample data subset, wherein the sample data subset comprises at least one piece of sample data; and correspondingly, one set of gradient information of each parameter is obtained correspondingly based on one sample data subset, wherein the one set of gradient information comprises at least one piece of gradient information.

14. The method according to claim 8 , wherein updating the machine learning model using an update gradient of each parameter in each iterative operation comprises:

updating the machine learning model according to a model update formula of stochastic gradient descent and using the update gradient of each parameter.

15. The method according to claim 8 , wherein the plurality of iterative operations are executed on at least one computing node, and the at least one computing node comprises at least one processor and a memory configured for the at least one processor.

16. The method according to claim 8 , wherein in each iterative operation, the plurality of sample data subsets that are sequentially read from the sample data set are stored in a memory, and after gradient information of each parameter is calculated using one sample data subset, the sample data subset is deleted from the memory before a next sample data subset is read into the memory.

17. The method according to claim 16 , wherein a storage space occupied by one sample data subset is less than or equal to a storage space reserved for the sample data subset in the memory, and a storage space occupied by two sample data subsets is greater than the storage space reserved for the sample data subset in the memory.

18. The method according to claim 16 , further comprising:

in each iterative operation, in a process of entering each sample data subset into the machine learning model and calculating gradient information of each of the plurality of parameters of the machine learning model, reading and using an intermediate calculation result stored in the memory, wherein the intermediate calculation result is used as input information to calculate the gradient information; and

after the intermediate calculation result is used, deleting the intermediate calculation result from the memory, wherein an operation of deleting the intermediate calculation result needs to be completed before a next sample data subset is read into the memory.

19. A non-transitory computer-readable storage medium, storing one or more instructions that, when executed by at least one processor, cause the at least one processor to:

sequentially reading a plurality of sample data subsets from a sample data set, wherein each sample data subset comprises at least one piece of sample data;

entering each read sample data subset into a machine learning model;

calculating gradient information of each of a plurality of parameters of the machine learning model, wherein after a set of gradient information of each parameter is calculated using one sample data subset, another set of gradient information of each parameter is calculated using a next sample data subset; and wherein the machine learning model has an initialized global parameter or is updated in a last iterative operation; and

accumulating a plurality of sets of calculated gradient information of each parameter to obtain an update gradient of each parameter, wherein the update gradient of each parameter is used to update the machine learning model.

Priority Claims (1)
CN 201611110243.6 · Dec 6, 2016 · national
Continuity (3)
Continuation 16432617 · Jun 5, 2019
Continuation PCTCN2017113581 · Nov 29, 2017
Related Publication 20240013098A1 · Jan 11, 2024
References Cited (43)
US 9350690B2 · Meijer et al. · 2016 [cited by applicant]
US 9652722B1 · Narsky · 2017 [cited by examiner]
US 10152676B1 · Strom · 2018 [cited by applicant]
US 10282809B2 · Jin et al. · 2019 [cited by applicant]
US 11308418B2 · Schiemenz · 2022 [cited by examiner]
US 20110264609A1 · Liu et al. · 2011 [cited by applicant]
US 20130325401A1 · Bouchard · 2013 [cited by applicant]
US 20140188462A1 · Zadeh · 2014 [cited by applicant]
US 20140236871A1 · Fujimaki et al. · 2014 [cited by applicant]
US 20150127590A1 · Gay et al. · 2015 [cited by applicant]
US 20150234781A1 · Angerer · 2015 [cited by examiner]
US 20150324690A1 · Chilimbi et al. · 2015 [cited by applicant]
US 20150379428A1 · Dirac et al. · 2015 [cited by applicant]
US 20160078359A1 · Csurka · 2016 [cited by examiner]
US 20160078361A1 · Brueckner · 2016 [cited by examiner]
US 20160350649A1 · Zhang et al. · 2016 [cited by applicant]
US 20170308789A1 · Langford et al. · 2017 [cited by applicant]
US 20180329798A1 · Zhou · 2018 [cited by applicant]
US 20190087744A1 · Schiemenz · 2019 [cited by examiner]
US 20190279088A1 · Zhang · 2019 [cited by examiner]
CN 101826166A · 2010 [cited by applicant]
CN 102521656A · 2012 [cited by applicant]
CN 103336877A · 2013 [cited by applicant]
CN 104463324A · 2015 [cited by applicant]
CN 104598972A · 2015 [cited by applicant]
CN 105469142A · 2016 [cited by applicant]
CN 105574585A · 2016 [cited by applicant]
CN 105677353A · 2016 [cited by applicant]
CN 105683944A · 2016 [cited by applicant]
CN 106062786A · 2016 [cited by applicant]
CN 106156810A · 2016 [cited by applicant]
CN 107330516A · 2017 [cited by applicant]
JP 2012079080A · 2012 [cited by applicant]
WO 2017185411A1 · 2017 [cited by applicant]
Song et al. Stochastic gradient descent with differentially private updates, 978-1-4799-0248-4/13/$31.00 © 2013 IEEE, GlobalSIP 2013. [cited by examiner]
Huang Jiuling, Research and Application on Imbalanced Data Set Based on Ensemble Learning Classification, Harbin University of Technology, May 2015, 2 pages. [cited by applicant]
Minsoo Rhu et al, vDNN: Virtualized Deep Neural Networks for Scalable, Memory-Efficient Neural Network Design, 49th IEEE/ACM International Symposium on Microarchitecture (MICRO-49, arXiv:1602.08124v3 [cs.DC] Jul. 28, 20… [cited by applicant]
Chen Mingzhong, Analysis and Compare of BP Neural Network's Training Arithmetic, Science Mosaic, 2010, Issue 03, 4 pages. [cited by applicant]
Small spoon digging Tarzan, Five algorithms for training neural networks, CSDN, Oct. 24, 2016, 14 pages, https://blog.csdn.net/baidu_32134295/article/details/52909687. [cited by applicant]
Aliaga R J et al: “SoC-Based Implementation of the Backpropagation Algorithm for MLP”, Hybrid Intelligent Systems, 2008. HIS″08. Sep. 2008, pp. 744-749, XP031321915. [cited by applicant]
Data Mining: Concepts, Models, Methods, and Algorithms, Chapter 3 Data Reduction, Mehmed Kantardzic, 2003, Wiley-IEEE Press. [cited by applicant]
Xiaosheng Liu et al.,“Scalable Parallel EM Algorithms for Latent Dirichlet Allocation in Multi-Core Systems”,May 18-22, 2015,total:11pages. [cited by applicant]
James McCaffrey et al.,“Variation on Back Propagation:Mini Batch Neural Network Training”,Jun. 6, 2023,total:8pags. [cited by applicant]