IP Library Granted Patent US 9,483,740
Granted Patent B1
US 9,483,740 · App. 14/108,119 · Granted Nov 1, 2016

Automated data classification

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,483,740
App. No.
14/108,119
Granted
Nov 1, 2016
Kind
B1
Abstract

A system and method for data classification are presented. A plurality of training tokens are identified by at least one server communicatively coupled to a network. Each training token includes a token retrieved from a content source and a classification of the token. For each training token in the plurality of training tokens, a plurality of n-gram sequences are identified, a plurality of features for the plurality of n-gram sequences are generated, and first training data is generated using the token retrieved from the content source, the plurality of features, and the classification of the token. A first classifier is trained with the first training data, and the first classifier is stored into a storage system in communication with the at least one server.

Claims (57)

1. A method, comprising:

identifying, by at least one server communicatively coupled to a network, a plurality of training tokens, each training token including a token retrieved from a content source and a classification of the token;

for each training token in the plurality of training tokens:

identifying, by the at least one server, a plurality of n-gram sequences,

generating, by the at least one server, a plurality of features for the plurality of n-gram sequences, and

generating, by the at least one server, first training data using the token retrieved from the content source, the plurality of features, and the classification of the token;

training a first classifier with the first training data;

storing, by the at least one server, the first classifier into a storage system in communication with the at least one server;

for each training token in the plurality of training tokens:

identifying a plurality of related tokens in the content source,

for each of the related tokens in the content source:

identifying a second plurality of n-gram sequences, and

generating a second plurality of features using the second plurality of n-gram sequences and by executing the first classifier on the related token to generate a probable classification of the related token;

generating second training data using the second plurality of features;

training a second classifier with the second training data; and

storing, by the at least one server, the second classifier into the storage system in communication with the at least one server.

2. The method of claim 1 , wherein each training token includes an indication of a visual appearance of the token retrieved from the content source.

3. The method of claim 2 , wherein the indication of the visual appearance includes at least one of a font size, font style, color, and position.

4. The method of claim 3 , wherein the indication of the visual appearance includes an orientation of the token.

5. The method of claim 1 , wherein the plurality of related tokens share a visual attribute with at least one of the plurality of training tokens.

6. The method of claim 5 , where the visual attribute is a font style or a font size.

7. The method of claim 1 , wherein the content source is a web page.

8. The method of claim 1 , wherein the first classifier is trained using stochastic gradient boosting.

9. A method, comprising:

identifying, by at least one server communicatively coupled to a network, a training token including a token retrieved from a content source and a classification of the token;

generating, by the at least one server, features for the training token;

training, by the at least one server, a classifier using the token retrieved from the content source, the features for the training token, and the classification; and

storing, by the at least one server, the classifier into a storage system in communication with the at least one server;

identifying, by the at least one server, a related token;

identifying second features for the related token by executing the classifier on the related token to generate a probable classification of the related token;

training, by the at least one server, a second classifier using the related token and the second features; and

storing, by the at least one server, the second classifier into a storage system in communication with the at least one server.

10. The method of claim 9 , wherein the training token includes an indication of a visual appearance of the token retrieved from the content source.

11. The method of claim 10 , wherein the indication of the visual appearance includes at least one of a font size, font style, color, and position.

12. The method of claim 11 , wherein the indication of the visual appearance includes an orientation of the token.

13. The method of claim 9 , wherein the content source is a web page.

14. The method of claim 9 , wherein the classifier is trained using stochastic gradient boosting.

15. A system, comprising:

a server computer configured to communicate with a content source using a network, the server computer being configured to:

identify a plurality of training tokens, each training token including a token retrieved from the content source and a classification of the token;

for each training token in the plurality of training tokens:

identify a plurality of n-gram sequences,

generate a plurality of features for the plurality of n-gram sequences, and

generate first training data using the token retrieved from the content source, the plurality of features, and the classification of the token;

train a first classifier with the first training data;

store the first classifier into a storage system in communication with the server computer;

for each training token in the plurality of training tokens:

identify a plurality of related tokens in the content source,

for each of the related tokens in the content source:

identifying a second plurality of n-gram sequences, and

generating a second plurality of features using the second plurality of n-gram sequences and by executing the first classifier on the related token to generate a probable classification of the related token;

generate second training data using the second plurality of features;

train a second classifier with the second training data; and

store, by the server computer, the second classifier into the storage system in communication with the at least one server.

16. The system of claim 15 , wherein each training token includes an indication of a visual appearance of the token retrieved from the content source.

17. The system of claim 16 , wherein the indication of the visual appearance includes at least one of a font size, font style, color, and position.

18. The system of claim 17 , wherein the indication of the visual appearance includes an orientation of the token.

Assignments (3)
SECURITY AGREEMENT Recorded Feb 17, 2023
From: GO DADDY OPERATING COMPANY, LLC; GD FINANCE CO, LLC; GODADDY MEDIA TEMPLE INC.; GODADDY.COM, LLC; LANTIRN INCORPORATED; POYNT, LLC
To: ROYAL BANK OF CANADA
Reel/Frame 062782/0489 →
SECURITY INTEREST Recorded May 20, 2014
From: LOCU, INC., AS GRANTOR; GO DADDY OPERATING COMPANY, LLC. AS GRANTOR; MEDIA TEMPLE, INC., AS GRANTOR; OUTRIGHT INC., AS GRANTOR
To: BARCLAYS BANK PLC, AS COLLATERAL AGENT
Reel/Frame 032933/0221 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2013
From: ANSEL, JASON; MARCUS, ADAM; MIERLE, KEIR; OLSZEWSKI, MAREK
To: LOCU, INC.
Reel/Frame 031818/0736 →