IP Library Granted Patent US 7,689,408
Granted Patent B2
US 7,689,408 · App. 11/515,468 · Granted Mar 30, 2010

Identifying language of origin for words using estimates of normalized appearance frequency

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,689,408
App. No.
11/515,468
Granted
Mar 30, 2010
Kind
B2
Abstract

The language of origin of a word or named entity is predicted using estimates of frequency of occurrence of the word or named entity in different languages. In one embodiment, the normalized frequency of occurrence of the word or named entity in a variety of different languages is estimated and the values are used as features in a feature vector which is scored and used to identify language of origin.

Claims (33)

1. A method of identifying a language of origin of an input word, using a computer with a processor, comprising:

generating a wide area network query based on the input word to obtain, with the processor, search results, comprising web pages, in a plurality of different languages;

estimating, with the processor, a normalized frequency of occurrence of the input word in each of the different languages based on the search results;

identifying, with the processor, the language of origin of the input word based on the estimated frequencies of occurrence, and

outputting an indication of the language of origin;

wherein the search results comprise web pages and wherein estimating a normalized frequency of occurrence in a selected language comprises:

obtaining a count of a number of web pages in the selected language in the search results that contain the input word; and

estimating a total number of web pages in the selected language by generating a wide area network query based on one or more function words in the selected language to obtain function word search results, and estimating the total number of web pages based on the function word search result.

2. The method of claim 1 wherein identifying the language of origin comprises:

generating a feature vector indicative of the frequencies of occurrence of the input word, in the different languages; and

identifying the language of origin based on the feature vector.

3. The method of claim 1 wherein estimating a normalized frequency of occurrence in the selected language comprises:

estimating the normalized frequency of occurrence as a ratio of the count of the number of web pages that contain the input word to the total number of web pages.

4. The method of claim 1 wherein the function word search results comprise an indication of web pages containing the one or more function words and wherein estimating the total number of web pages based on the function word search result, comprises:

estimating the total number of web pages in the selected language as a number of web pages that contain one of the function words.

5. The method of claim 1 wherein generating a wide area network query based on the one or more function words comprises:

obtaining a predefined set of a plurality of function words; and

generating the query using the predefined set of function words.

6. The method of claim 1 and further comprising:

generating an indication of how likely the input word is to have a given language of origin based on a morphological structure of the input word.

7. The method of claim 2 wherein generating the feature vector comprises:

generating morphological structure features indicative of n-gram scores for the input word given each of the different languages; and

generating the feature vector to indicate the morphological structure features.

8. A system for identifying a language of origin of an input word, comprising:

a feature extraction system comprising a frequency of occurrence estimation system estimating a frequency of occurrence of the input word in each of a plurality of different languages;

a language identifier identifying the language of origin of the input word based on the frequency of occurrence estimated;

a search engine coupled to the feature extraction system, the feature extraction system generating a wide area network query based on one or more function words in a selected language to obtain function word search results and extracting features from the function word search results, the features being indicative of the frequency of occurrence of the input word;

the feature extraction system extracting normalized frequency of occurrence features based on a number of pages in the function word search results, in the selected language, that contain the one or more function words and an estimate of a total number of pages in the language; and

a computer processor, activated by the frequency of occurrence estimation system, to facilitate estimating the frequency of occurrence.

9. The system of claim 8 wherein the feature extraction component includes:

a morphological feature extractor configured to extract morphological features indicative of how likely the input word has a given language of origin based on a morphological structure of the input word.

10. The system of claim 9 wherein the morphological feature extractor is configured to extract n-gram scores for the input word in each language.

11. The system of claim 10 wherein the language identifier comprises a classifier.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034542/0001 →