IP Library Granted Patent US 9,805,312
Granted Patent B1
US 9,805,312 · App. 14/105,262 · Granted Oct 31, 2017

Using an integerized representation for large-scale machine learning data

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,805,312
App. No.
14/105,262
Granted
Oct 31, 2017
Kind
B1
Abstract

Methods and systems for replacing feature values of features in training data with integer values selected based on a ranking of the feature values. The methods and systems are suitable for preprocessing large-scale machine learning training data.

Claims (40)

1. A method comprising:

receiving, by a data processing system, training data examples that contain a plurality of features for each of a plurality of templates, wherein each of the plurality of templates is a category of feature-types that includes multiple features, all of which are from the same category;

for each of the templates, preprocessing the features of the template in the training data examples, before processing the training data examples by a machine learning system, the preprocessing including, for each of the templates:

ranking the plurality of features of the template based on a score, a number of impressions, a number of occurrences, a type of machine learning system, and a type of training examples received by the data processing system;

determining a first set of features associated with the template, wherein each feature in the first set of features exceeds a threshold ranking criteria;

determining a second set of features associated with the template, wherein each feature in the second set of features falls below a minimum threshold ranking criteria; and

assigning an integer value to each feature of the first set of features in order based upon ranking; and

before the training data examples are processed by the machine learning system, for each of the templates and each of the features in the first set of features associated with the template, replacing the feature in the training data with the integer value assigned to the feature, and removing from the training data each of the features of the template in the second set of features associated with the template.

2. The method of claim 1 , wherein features in the received training data examples comprise text strings.

3. The method of claim 1 , wherein assigning the integer value further comprises assigning the lowest integer value to the highest ranked feature in the first set of features.

4. The method of claim 3 , wherein the lowest integer value is 1.

5. The method of claim 4 , wherein integer values assigned to features are non-unique between templates.

6. The method of claim 1 , wherein:

for each of the templates, the second set of features associated with the template is all the features other than the features in the first set of features associated with the template.

7. A data processing system comprising a computer configured to perform operations comprising:

receiving training data examples that contain a plurality of features for each of a plurality of templates, wherein each of the plurality of templates is a category of feature-types that includes multiple features, all of which are from the same category;

for each of the templates, preprocessing of the features of the template in the training data examples, before processing the training data examples by a machine learning system, the preprocessing including, for each of the templates:

ranking the plurality of features of the template based on a score, a number of impressions, a number of occurrences, a type of machine learning system, and a type of training examples received by the data processing system;

determining a first set of features associated with the template, wherein each feature in the first set of features exceeds a threshold ranking criteria;

determining a second set of features associated with the template, wherein each feature in the second set of features falls below a minimum threshold ranking criteria; and

assigning an integer value to each of the first set of features in order based upon ranking; and

before the training data examples are processed by the machine learning system, for each of the templates and each of the features in the first set of features associated with the template, replacing the feature in the training data with the integer value assigned to the feature, and removing from the training data each of the features of the template in the second set of features associated with the template.

8. The system of claim 7 , wherein features in the received training data comprise text strings.

9. The system of claim 7 , wherein assigning the integer value further comprises assigning the lowest integer value to the highest ranked feature in the first set of features.

10. The system of claim 9 , wherein the lowest integer value is 1.

11. The system of claim 10 , wherein integer values assigned to features are non-unique between templates.

12. The system of claim 7 , wherein:

for each of the templates, the second set of features associated with the template is all the features other than the features in the first set of features associated with the template.

13. The method of claim 1 , further comprising:

processing the training data examples by the machine learning system.

14. The system of claim 7 , wherein the operations further comprise:

processing the training data examples by the machine learning system.

15. One or more non-transitory computer-readable storage media encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

receiving training data examples that contain a plurality of features for each of a plurality of templates, wherein each of the plurality of templates is a category of feature-types that includes multiple features, all of which are from the same category;

for each of the templates, preprocessing the features of the template in the training data examples, before processing the training data examples by a machine learning system, the preprocessing including, for each of the templates:

ranking the plurality of features of the template based on a score, a number of impressions, a number of occurrences, a type of machine learning system, and a type of training examples received by the data processing system;

determining a first set of features associated with the template, wherein each feature in the first set of features exceeds a threshold ranking criteria;

determining a second set of features associated with the template, wherein each feature in the second set of features falls below a minimum threshold ranking criteria; and

assigning an integer value to each feature of the first set of features in order based upon ranking; and

before the training data examples are processed by the machine learning system, for each of the templates and each of the features in the first set of features associated with the template, replacing the feature in the training data with the integer value assigned to the feature, and removing from the training data each of the features of the template in the second set of features associated with the template.

Assignments (2)
CHANGE OF NAME Recorded Dec 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044695/0115 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2013
From: SHAKED, TAL; CHANDRA, TUSHAR DEEPAK; SINGER, YORAM; IE, TZE WAY EUGENE; REDSTONE, JOSHUA
To: GOOGLE INC.
Reel/Frame 031777/0153 →