IP Library Granted Patent US 11,334,819
Granted Patent B2
US 11,334,819 · App. 17/005,592 · Granted May 17, 2022

Method and system for distributed machine learning

Inventors: Andrew Feng (Cupertino, CA); Erik Ordentlich (San Jose, CA); Lee Yang (Campbell, CA); Peter Cnudde (Los Altos, CA)
Assignee: Verizon Patent and Licensing Inc.
G06N20/00H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,334,819
App. No.
17/005,592
Granted
May 17, 2022
Kind
B2
Abstract

The present teaching relates to estimating one or more parameters on a system including a plurality of nodes. In one example, the system comprises: one or more learner nodes, each of which is configured for generating information related to a group of words for estimating the one or more parameters associated with a machine learning model; and a plurality of server nodes, each of which is configured for obtaining a plurality of sub-vectors each of which is a portion of a vector that represents a word in the group of words, updating the sub-vectors based at least partially on the information to generate a plurality of updated sub-vectors, and estimating at least one of the one or more parameters associated with the machine learning model based on the plurality of updated sub-vectors.

Claims (62)

1. A system including a plurality of nodes, each of which has at least one processor, storage, and a communication platform connected to a network for training a machine learning model, the system comprising:

a plurality of server nodes, each of which is communicatively coupled with a plurality of learner nodes; and

a driver node configured for

receiving a request for training a model,

generating a plurality of sub-vectors, each of which is a portion of a feature vector associated with a word in a group of words,

allocating different sub-vectors associated with the word to different server nodes,

partitioning training data into a plurality of training data subsets, each of which is allocated to a learner node,

receiving a notification indicating completion of training the model, and

transmitting the model in response to the request.

2. The system of claim 1 , wherein each of the plurality of learner nodes is configured for generating information related to a group of words, wherein the information is indicative of a center word and relationship thereof with other words in the group of words.

3. The system of claim 2 , wherein the information related to the group of words includes a first index for the center word, and a second set of one or more indices for one or more positive context words of the center word.

4. The system of claim 3 , wherein:

the information related to the group of words further includes a random seed; and

each of the plurality of server nodes is further configured for generating, based on the random seed and the first index, a third set of one or more indices for one or more negative words that are not semantically related to the center word.

5. The system of claim 1 , wherein the driver node is further configured for:

selecting, based on the request, the one or more learner nodes and the plurality of server nodes for training the model; and

initializing at least some feature vectors in the plurality of feature vectors with random numbers.

6. The system of claim 1 , wherein each of the plurality of server nodes is further configured for:

calculating a plurality of partial dot products each of which is a dot product of two sub-vectors; and

providing the plurality of partial dot products to the plurality of learner nodes.

7. The system of claim 1 , wherein:

the model is utilized to determine one or more advertisements based on a query; and

the one or more advertisements are to be provided in response to the query.

8. A method implemented on at least one computing device each of which has at least one processor, storage, and a communication platform connected to a network for training a machine learning model, the method comprising:

receiving a request for training a model;

generating a plurality of sub-vectors, each of which is a portion of a feature vector associated with a word in a group of words;

allocating different sub-vectors associated with the word to different server nodes of a plurality of server nodes;

partitioning training data into a plurality of training data subsets, each of which is allocated to a learner node of a plurality of learner nodes, wherein each server node is communicatively coupled with the plurality of learner nodes;

receiving a notification indicating completion of training the model; and

transmitting the model in response to the request.

9. The method of claim 8 , wherein each of the plurality of learner nodes is configured for generating information related to a group of words, wherein the information is indicative of a center word and relationship thereof with other words in the group of words.

10. The method of claim 9 , wherein the information related to the group of words includes a first index for the center word, and a second set of one or more indices for one or more positive context words of the center word.

11. The method of claim 10 , wherein:

the information related to the group of words further includes a random seed; and

each of the plurality of server nodes is further configured for generating, based on the random seed and the first index, a third set of one or more indices for one or more negative words that are not semantically related to the center word.

12. The method of claim 8 , further comprising:

selecting, based on the request, the one or more learner nodes and the plurality of server nodes for training the model; and

initializing at least some feature vectors in the plurality of feature vectors with random numbers.

13. The method of claim 8 , further comprising:

calculating a plurality of partial dot products each of which is a dot product of two sub-vectors; and

providing the plurality of partial dot products to the plurality of learner nodes.

14. The method of claim 8 , wherein:

the model is utilized to determine one or more advertisements based on a query; and

the one or more advertisements are to be provided in response to the query.

15. A machine-readable tangible and non-transitory medium having information for training a machine learning model, wherein the information, when read by the machine, causes the machine to perform the following:

receiving a request for training a model;

generating a plurality of sub-vectors, each of which is a portion of a feature vector associated with a word in a group of words;

allocating different sub-vectors associated with the word to different server nodes of a plurality of server nodes;

partitioning training data into a plurality of training data subsets, each of which is allocated to a learner node of a plurality of learner nodes, wherein each server node is communicatively coupled with the plurality of learner nodes;

receiving a notification indicating completion of training the model; and

transmitting the model in response to the request.

16. The medium of claim 15 , wherein each of the plurality of learner nodes is configured for generating information related to a group of words, wherein the information is indicative of a center word and relationship thereof with other words in the group of words.

17. The medium of claim 16 , wherein the information related to the group of words includes a first index for the center word, and a second set of one or more indices for one or more positive context words of the center word.

18. The medium of claim 17 , wherein:

the information related to the group of words further includes a random seed; and

each of the plurality of server nodes is further configured for generating, based on the random seed and the first index, a third set of one or more indices for one or more negative words that are not semantically related to the center word.

19. The medium of claim 15 , further comprising:

selecting, based on the request, the one or more learner nodes and the plurality of server nodes for training the model; and

initializing at least some feature vectors in the plurality of feature vectors with random numbers.

20. The medium of claim 15 , wherein:

the model is utilized to determine one or more advertisements based on a query; and

the one or more advertisements are to be provided in response to the query.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2021
From: VERIZON MEDIA INC.
To: VERIZON PATENT AND LICENSING INC.
Reel/Frame 057453/0431 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2020
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 054258/0635 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2020
From: FENG, ANDREW; ORDENTLICH, ERIK; YANG, LEE; CNUDDE, PETER
To: YAHOO! INC.
Reel/Frame 053626/0837 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2020
From: YAHOO HOLDINGS, INC.
To: OATH INC.
Reel/Frame 053636/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2020
From: YAHOO! INC.
To: YAHOO HOLDINGS, INC.
Reel/Frame 054157/0218 →
Continuity (2)
Continuation 15098415 · Apr 14, 2016
Related Publication 20210049507A1 · Feb 18, 2021
Cited By (2)
US 12,198,072 US 12,468,963