IP Library Granted Patent US 10,339,423
Granted Patent B1
US 10,339,423 · App. 15/621,452 · Granted Jul 2, 2019

Systems and methods for generating training documents used by classification algorithms

Inventors: Jonathan J. Dinerstein (Draper, UT); Christian Larsen (Orem, UT); Daniel Hardman (American Fork, UT)
Assignee: Symantec Corporation
G06K9/6256G06F17/2705G06F17/2845
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,339,423
App. No.
15/621,452
Granted
Jul 2, 2019
Kind
B1
Abstract

The disclosed computer-implemented method for generating training documents used by classification algorithms may include (i) identifying a set of training documents used by a classification system to classify documents written in a first language, (ii) generating a list of tokens from within the training documents that indicate critical terms representative of classes defined by the classification system, (iii) translating the list of tokens from the first language to a second language, (iv) creating, based on the translated tokens, a set of simulated training documents that enables the classification system to classify documents written in the second language, and (v) classifying an additional document written in the second language based on the set of simulated training documents. Various other methods, systems, and computer-readable media are also disclosed.

Claims (46)

1. A computer-implemented method for generating training documents used by classification algorithms, at least a portion of the method being performed by a computing device comprising at least one processor, the method comprising:

identifying a set of training documents used by a classification system to classify documents written in a first language;

generating a list of tokens from within the training documents that indicate critical terms representative of classes defined by the classification system;

determining linguistic properties of at least one token by analyzing a context in which the token is used within the set of training documents;

translating, while retaining the linguistic properties of the token, the list of tokens from the first language to a second language;

creating, based on the translated tokens, a set of simulated training documents that enables the classification system to classify documents written in the second language; and

classifying an additional document written in the second language based on the set of simulated training documents.

2. The method of claim 1 , wherein generating the list of tokens comprises identifying phrases that are used with at least a certain frequency within the set of training documents.

3. The method of claim 1 , wherein generating the list of tokens comprises identifying non-linguistic tokens that indicate types of content within the set of training documents.

4. The method of claim 1 , wherein the linguistic properties of the token comprise at least one of:

a denotation of the token;

a connotation of the token; and

a syntactic structure of the token.

5. The method of claim 1 , wherein translating the list of tokens is performed using fewer computing resources than are required to translate the set of training documents.

6. The method of claim 1 , wherein creating the set of simulated training documents comprises structuring at least one simulated training document based on a structure of a training document within the set of training documents.

7. The method of claim 1 , wherein creating the set of simulated training documents comprises distributing a token, after the token has been translated, throughout a simulated training document based on a distribution of the token throughout a training document within the set of training documents.

8. The method of claim 1 , wherein creating the set of simulated training documents comprises:

identifying, within the set of training documents, noise signals comprising non-token terms that reduce a statistical significance of at least one token; and

replicating the noise signals within the set of simulated training documents.

9. The method of claim 1 , wherein creating the set of simulated training documents comprises placing non-token terms between the translated tokens within the simulated training documents such that the translated tokens are not directly adjacent to other translated tokens.

10. The method of claim 1 , wherein creating the set of simulated training documents comprises generating fake documents that contain text in the second language but do not contain content that is comprehensible by a speaker of the second language.

11. A system for generating training documents used by classification algorithms, the system comprising:

an identification module, stored in memory, that identifies a set of training documents used by a classification system to classify documents written in a first language;

a token module, stored in memory, that generates a list of tokens from within the training documents that indicate critical terms representative of classes defined by the classification system;

a translation module, stored in memory, that:

determines linguistic properties of at least one token by analyzing a context in which the token is used within the set of training documents; and

translates, while retaining the linguistic properties of the token, the list of tokens from the first language to a second language;

a simulation module, stored in memory, that creates, based on the translated tokens, a set of simulated training documents that enables the classification system to classify documents written in the second language;

a classification module, stored in memory, that classifies an additional document written in the second language based on the set of simulated training documents; and

at least one physical processor configured to execute the identification module, the token module, the translation module, the simulation module, and the classification module.

12. The system of claim 11 , wherein the token module generates the list of tokens by identifying phrases that are used with at least a certain frequency within the set of training documents.

13. The system of claim 11 , wherein the token module generates the list of tokens by identifying non-linguistic tokens that indicate types of content within the set of training documents.

14. The system of claim 11 , wherein the linguistic properties of the token comprise at least one of:

a denotation of the token;

a connotation of the token; and

a syntactic structure of the token.

15. The system of claim 11 , wherein the translation module translates the list of tokens using fewer computing resources than are required to translate the set of training documents.

16. The system of claim 11 , wherein the simulation module creates the set of simulated training documents by structuring at least one simulated training document based on a structure of a training document within the set of training documents.

17. The system of claim 11 , wherein the simulation module creates the set of simulated training documents by distributing a token, after the token has been translated, throughout a simulated training document based on a distribution of the token throughout a training document within the set of training documents.

18. A non-transitory computer-readable medium comprising one or more computer-executable instructions that, when executed by at least one processor of a computing device, cause the computing device to:

identify a set of training documents used by a classification system to classify documents written in a first language;

generate a list of tokens from within the training documents that indicate critical terms representative of classes defined by the classification system;

determine linguistic properties of at least one token by analyzing a context in which the token is used within the set of training documents;

translate, while retaining the linguistic properties of the token, the list of tokens from the first language to a second language;

create, based on the translated tokens, a set of simulated training documents that enables the classification system to classify documents written in the second language; and

classify an additional document written in the second language based on the set of simulated training documents.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2019
From: SYMANTEC CORPORATION
To: CA, INC.
Reel/Frame 051144/0918 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2017
From: DINERSTEIN, JONATHAN J.; LARSEN, CHRISTIAN; HARDMAN, DANIEL
To: SYMANTEC CORPORATION
Reel/Frame 042694/0596 →
Cited By (11)
US 12,288,039 US 12,417,506 US 12,423,525 US 12,462,114 US 12,468,694 US 12,505,093 US 12,608,416 US 12,614,042 US 12,632,445 US 12,675,732 US 12,681,997