IP Library Granted Patent US 10,083,167
Granted Patent B2
US 10,083,167 · App. 14/506,156 · Granted Sep 25, 2018

System and method for unsupervised text normalization using distributed representation of words

Inventor: Vivek Kumar Rangarajan Sridhar (Morristown, NJ)
Assignee: AT&T INTELLECTUAL PROPERTY I, L.P.
G06F17/273G06F17/289G06Q50/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,083,167
App. No.
14/506,156
Granted
Sep 25, 2018
Kind
B2
Abstract

A system, method and computer-readable storage devices for providing unsupervised normalization of noisy text using distributed representation of words. The system receives, from a social media forum, a word having a non-canonical spelling in a first language. The system determines a context of the word in the social media forum, identifies the word in a vector space model, and selects an “n-best” vector paths in the vector space model, where the n-best vector paths are neighbors to the vector space path based on the context and the non-canonical spelling. The system can then select, based on a similarity cost, a best path from the n-best vector paths and identify a word associated with the best path as the canonical version.

Claims (38)

1. A method comprising:

receiving, from a social media forum, a correctly-spelled word having a non-canonical spelling, the non-canonical spelling comprising a correct spelling of a variant of a canonical spelling of the correctly-spelled word;

forming a linear finite state machine based on the correctly-spelled word;

composing the linear finite state machine with a finite state transducer, wherein the finite state transducer comprises a vector space model trained from a corpus of noisy text, and wherein words within the finite state transducer are clustered based on context, to yield a modified finite state machine;

composing the modified finite state machine with a language model constructed from clean vocabulary sentences to yield a resulting finite state machine;

performing a best path function on the resulting finite state machine, wherein the best path function comprises:

selecting n-best vector paths in the vector space model which are neighbors to the non-canonical spelling; and

selecting, based on a similarity cost, a best path from the n-best vector paths; and nominating a proposed word associated with the best path as a canonical form.

2. The method of claim 1 , wherein the correctly-spelled word is classified in the finite state transducer based on a word context and the non-canonical spelling.

3. The method of claim 1 , wherein the correctly-spelled word comprises a compound word.

4. The method of claim 1 , wherein nominating of the correctly-spelled word is used in a translation.

5. The method of claim 1 , wherein the similarity cost is based on a type of the non-canonical spelling.

6. The method of claim 5 , wherein the type of the non-canonical spelling is an abbreviation.

7. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

receiving, from a social media forum, a correctly-spelled word having a non-canonical spelling, the non-canonical spelling comprising a correct spelling of a variant of a canonical spelling of the correctly-spelled word;

forming a linear finite state machine based on the correctly-spelled word;

composing the linear finite state machine with a finite state transducer, wherein the finite state transducer comprises a vector space model trained from a corpus of noisy text, and wherein words within the finite state transducer are clustered based on context, to yield a modified finite state machine;

composing the modified finite state machine with a language model constructed from clean vocabulary sentences to yield a resulting finite state machine;

performing a best path function on the resulting finite state machine, wherein the best path function comprises:

selecting n-best vector paths in the vector space model which are neighbors to the non-canonical spelling; and

selecting, based on a similarity cost, a best path from the n-best vector paths; and

nominating a proposed word associated with the best path as a canonical form.

8. The system of claim 7 , wherein the correctly-spelled word is classified in the finite state transducer based on a word context and the non-canonical spelling.

9. The system of claim 7 , wherein the correctly-spelled word comprises a compound word.

10. The system of claim 7 , nominating of the correctly-spelled word is used in a translation.

11. The system of claim 7 , wherein the similarity cost is based on a type of the non-canonical spelling.

12. The system of claim 11 , wherein the type of the non-canonical spelling is an abbreviation.

13. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

receiving, from a social media forum, a correctly-spelled word having a non-canonical spelling, the non-canonical spelling comprising a correct spelling of a variant of a canonical spelling of the correctly-spelled word;

forming a linear finite state machine based on the correctly-spelled word;

composing the linear finite state machine with a finite state transducer, wherein the finite state transducer comprises a vector space model trained from a corpus of noisy text, and wherein words within the finite state transducer are clustered based on context, to yield a modified finite state machine;

composing the modified finite state machine with a language model constructed from clean vocabulary sentences to yield a resulting finite state machine;

performing a best path function on the resulting finite state machine, wherein the best path function comprises:

selecting n-best vector paths in the vector space model which are neighbors to the non-canonical spelling; and

selecting, based on a similarity cost, a best path from the n-best vector paths; and nominating a proposed word associated with the best path as a canonical form.

14. The computer-readable storage device of claim 13 , wherein the correctly-spelled word is classified in the finite state transducer based on a word context and the non-canonical spelling.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2014
From: RANGARAJAN SRIDHAR, VIVEK KUMAR
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 033883/0984 →
Continuity (1)
Related Publication 20160098386A1 · Apr 7, 2016
Cited By (2)
US 1,101,770 US 1,102,444