IP Library Granted Patent US 10,402,235
Granted Patent B2
US 10,402,235 · App. 15/980,196 · Granted Sep 3, 2019

Fine-grain synchronization in data-parallel jobs for distributed machine learning

Inventors: Asim Kadav (Jersey City, NJ); Erik Kruus (East Hillsborough, NJ)
Assignee: NEC CORPORATION
G06F9/522G06N20/00H04L67/1095
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,402,235
App. No.
15/980,196
Granted
Sep 3, 2019
Kind
B2
Abstract

A computer-implemented method and computer processing system are provided. The method includes synchronizing, by a processor, respective ones of a plurality of data parallel workers with respect to an iterative distributed machine learning process. The synchronizing step includes individually continuing, by the respective ones of the plurality of data parallel workers, from a current iteration to a subsequent iteration of the iterative distributed machine learning process, responsive to a satisfaction of a predetermined condition thereby. The predetermined condition includes individually sending a per-receiver notification from each sending one of the plurality of data parallel workers to each receiving one of the plurality of data parallel workers, responsive to a sending of data there between. The predetermined condition further includes individually sending a per-receiver acknowledgement from the receiving one to the sending one, responsive to a consumption of the data thereby.

Claims (16)

1. A computer-implemented method, comprising:

synchronizing, by a processor, respective ones of a plurality of data parallel workers with respect to an iterative distributed machine learning process,

wherein said synchronizing step includes individually continuing, by the respective ones of the plurality of data parallel workers, from a current iteration to a subsequent iteration of the iterative distributive machine learning process, responsive to a satisfaction of a predetermined condition thereby, and

wherein the predetermined condition includes:

individually sending a per-receiver notification from each sending one of the plurality of data parallel workers to each receiving one of the plurality of data parallel workers, responsive to a sending of data there between; and

individually sending a per-receiver acknowledgement from the receiving one to the sending one, responsive to a consumption of model parameters of the data thereby;

wherein at least some of the respective ones of the plurality of data parallel workers continue to the subsequent iteration at different times; and

wherein the different times are based on respective times at which the predetermined condition is satisfied by the at least some of the respective ones of the plurality of data parallel workers.

2. A computer program product for data synchronization, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

synchronizing, by a processor, respective ones of a plurality of data parallel workers with respect to an iterative distributed machine learning process,

wherein said synchronizing step includes individually continuing, by the respective ones of the plurality of data parallel workers, from a current iteration to a subsequent iteration of the iterative distributed machine learning process, responsive to a satisfaction of a predetermined condition thereby, and

wherein the predetermined condition includes:

individually sending a per-receiver notification from each sending one of the plurality of data parallel workers to each receiving one of the plurality of data parallel workers, responsive to a sending of data there between; and

individually sending a per-receiver acknowledgement from the receiving one to the sending one, responsive to a consumption of model parameters of the data thereby;

wherein at least some of the respective ones of the plurality of data parallel workers continue to the subsequent iteration at different times; and

wherein the different times are based on respective times at which the predetermined condition is satisfied by the at least some of the respective ones of the plurality of data parallel workers.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2019
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 049750/0034 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2018
From: KADAV, ASIM; KRUUS, ERIK
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 045809/0736 →
Continuity (3)
Continuation In Part 15480874 · Apr 6, 2017
Provisional Application 62322849 · Apr 15, 2016
Related Publication 20180260256A1 · Sep 13, 2018