IP Library Granted Patent US 8,868,469
Granted Patent B2
US 8,868,469 · App. 12/775,547 · Granted Oct 21, 2014

System and method for phrase identification

Inventors: Liqin Xu (Scarborough, CA); Hyun Chul Lee (Thornhill, CA)
Assignee: Rogers Communications Inc.
G06F17/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,868,469
App. No.
12/775,547
Granted
Oct 21, 2014
Kind
B2
Abstract

A phrase identification system and method are provided. The method comprises: identifying one or more phrase candidates in the electronic document; selecting one of the phrase candidates; numerically representing features of the selected phrase candidates to obtain a numeric feature representation associated with that phrase candidate; and inputting the numeric feature representation into a machine learning classifier, the machine learning classifier being configured to determine, based on each numeric feature representation, whether the phrase candidate associated with that numeric feature representation is a phrase.

Claims (50)

1. A method of identifying phrases in an electronic document comprising:

identifying one or more phrase candidates in the electronic document;

selecting one of the phrase candidates;

numerically representing features of the selected phrase candidates to obtain a numeric feature representation associated with that phrase candidate by:

identifying, from a predetermined dictionary map which maps words to unique numbers, the number associated with each word in the selected phrase candidate; and

including the identified number associated with each word in the numeric feature representation;

inputting the numeric feature representation into a machine learning classifier, the machine learning classifier being configured to determine, based on each numeric feature representation, that the phrase candidate associated with that numeric feature representation is a phrase, wherein identifying phrase candidates in the electronic document comprises identifying word level n-grams in the document; and

associating the electronic document with a label identifying the subject matter of the electronic document, wherein the label is based on the phrase.

2. The method of claim 1 , wherein identifying word level n-grams comprises identifying word level n-grams which are of a size that is greater than or equal to two words.

3. The method of claim 1 , wherein identifying word level n-grams comprises identifying word level n-grams which are of a size that is greater than or equal to a first predetermined threshold and less than or equal to a second predetermined threshold.

4. The method of claim 1 , wherein identifying phrase candidates in the electronic document further comprises applying a rule-based filter to the word level n-grams.

5. The method of claim 1 wherein numerically representing features of the selected phrase candidates comprises:

performing part of speech tagging on each word in the selected phrase candidate;

determining one or more numbers associated with the part of speech of the words in the selected phrase candidate; and

determining the numeric feature representation associated with the selected phrase candidate in accordance with the numbers associated with the part of speech of the words in the phrase candidate.

6. The method of claim 5 wherein numerically representing features of the phrase candidates further comprises:

performing part of speech tagging on one or more words adjacent to the selected phrase candidate;

determining one or more numbers associated with the part of speech of the words adjacent to the selected phrase candidate; and

determining the numeric feature representation associated with the phrase candidate in accordance with the numbers associated with the part of speech of the words adjacent to the selected phrase candidate.

7. The method of claim 1 further comprising, prior to identifying:

training the machine learning classifier with training data, the training data including one or more electronic training documents and one or more phrase labels which identify one or more phrases in the electronic training documents.

8. The method of claim 1 further comprising saving the identified phrases in a memory.

9. The method of claim 1 , further comprising:

categorizing the electronic document based on the associated label.

10. A phrase identification system for identifying phrases in an electronic document, comprising:

a memory;

one or more processors, configured to:

identify one or more phrase candidates in the electronic document;

select one of the phrase candidates;

numerically represent features of the selected phrase candidates to obtain a numeric feature representation associated with that phrase candidate by:

identifying, from a predetermined dictionary map which maps words to unique numbers, the number associated with each word in the selected phrase candidate; and

including the identified number associated with each word in the numeric feature representation;

input the numeric feature representation into a machine learning classifier, the machine learning classifier being configured to determine, based on each numeric feature representation, that the phrase candidate associated with that numeric feature representation is a phrase, wherein identifying phrase candidates in the electronic document comprises identifying word level n-grams in the document; and,

associate the electronic document with a label identifying the subject matter of the electronic document, wherein the label is based on the phrase.

11. The phrase identification system of claim 10 , wherein identifying word level n-grams comprises identifying word level n-grams which are of a size that is greater than or equal to two words.

12. The phrase identification system of claim 10 , wherein identifying word level n-grams comprises identifying word level n-grams which are of a size that is greater than or equal to a first predetermined threshold and less than or equal to a second predetermined threshold.

13. The phrase identification system of claim 10 , wherein identifying phrase candidates in the electronic document further comprises applying a rule-based filter to the word level n-grams.

14. The phrase identification system of claim 10 wherein numerically representing features of the selected phrase candidates comprises:

performing part of speech tagging on each word in the selected phrase candidate;

determine one or more numbers associated with the part of speech of the words in the selected phrase candidate; and

determine the numeric feature representation associated with the selected phrase candidate in accordance with the numbers associated with the part of speech of the words in the phrase candidate.

15. The phrase identification system of claim 14 wherein numerically representing features of the phrase candidates further comprises:

performing part of speech tagging on one or more words adjacent to the selected phrase candidate;

determining one or more numbers associated with the part of speech of the words adjacent to the selected phrase candidate; and

determining the numeric feature representation associated with the phrase candidate in accordance with the numbers associated with the part of speech of the words adjacent to the selected phrase candidate.

16. The phrase identification system of claim 10 wherein, prior to identifying, the one or more processors are further configured to:

train the machine learning classifier with training data, the training data including one or more electronic training documents and one or more phrase labels which identify one or more phrases in the electronic training documents.

17. The phrase identification system of claim 10 wherein the one or more processors are further configured to save the identified phrases in the memory.

18. The phrase identification system of claim 10 , wherein the one or more processors are further configured to:

categorize the electronic document based on the associated label.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2011
From: 2167959 ONTARIO INC.
To: ROGERS COMMUNICATIONS INC.
Reel/Frame 026885/0281 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2010
From: XU, LIQIN; LEE, HYUN CHUL
To: 2167959 ONTARIO INC.
Reel/Frame 024353/0957 →
Continuity (2)
Provisional Application 61251790 · Oct 15, 2009
Related Publication 20110093414A1 · Apr 21, 2011