IP Library Granted Patent US 7,191,115
Granted Patent B2
US 7,191,115 · App. 10/173,252 · Granted Mar 13, 2007

Statistical method and apparatus for learning translation relationships among words

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,191,115
App. No.
10/173,252
Granted
Mar 13, 2007
Kind
B2
Abstract

A parallel bilingual training corpus is parsed into its content words. Word association scores for each pair of content words consisting of a word of language L 1 that occurs in a sentence aligned in the bilingual corpus to a sentence of language L 2 in which the other word occurs. A pair of words is considered “linked” in a pair of aligned sentences if one of the words is the most highly associated, of all the words in its sentence, with the other word. The occurrence of compounds is hypothesized in the training data by identifying maximal, connected sets of linked words in each pair of aligned sentences in the processed and scored training data. Whenever one of these maximal, connected sets contains more than one word in either or both of the languages, the subset of the words in that language is hypothesized as a compound.

Claims (34)

1. A method of calculating translation relationships among words, comprising:

calculating word association scores for word pairs based on co-occurrences of words in each of a plurality of sets of aligned, bilingual units in a corpus;

identifying hypothesized compounds in the units based on the word association scores;

re-calculating the word association scores, given the hypothesized compounds; and

obtaining translation relationships based on the re-calculated word association scores, wherein obtaining translation relationships comprises:

repeating the step of re-calculating word association scores considering co-occurrences of pairs, including pairs of words, pairs of compounds, and compound/word pairs, in a pair of aligned units only if the pairs are uniquely most strongly associated with one another among all words in the pair of aligned units, to obtain ultimate word association scores; and

ranking pairs based on the ultimate word association scores.

2. The method of claim 1 wherein obtaining translation relationships further comprises:

selecting pairs as translations of one another if the corresponding ultimate word association scores are above a threshold level.

3. A method of calculating translation relationships among words, comprising:

calculating word association scores for word pairs based on co-occurrences of words in each of a plurality of sets of aligned, bilingual units in a corpus;

identifying hypothesized compounds in the units based on the word association scores;

re-calculating the word association scores, given the hypothesized compounds; and

obtaining translation relationships based on the re-calculated word association scores;

wherein calculating word association scores comprises:

calculating the word association scores based on a surface form of the words in each of the aligned, bilingual units.

4. A method of training a machine translation system, comprising:

obtaining a corpus of aligned, bilingual multi-word units;

calculating word association scores for word pairs in the corpus based on co-occurrence of words in the aligned units;

identifying hypothesized compounds based on an absence of one-to-one correspondence between words in the aligned units;

re-calculating the word association scores, given the hypothesized compounds;

training the machine translation system based on the word association scores and the hypothesized compounds;

repeating the step of re-calculating word association scores considering co-occurrences of pairs, including word pairs, compound pairs and word/compound pairs, in a pair of aligned units only if the pairs are uniquely most strongly associated with one another among all words in the pair of aligned units, to obtain ultimate word association scores; and

ranking pairs based on ultimate word association scores.

5. The method of claim 4 and further comprising:

selecting pairs as translations of one another if the corresponding ultimate word association scores are above a threshold level.

6. The method of claim 5 wherein training the machine translation system based on the word association scores and the hypothesized compounds, comprises:

generating transfer mappings mapping a unit in one of the languages to a unit in the other of the languages based on the selected translations.

7. A method of training a machine translation system, comprising:

obtaining a corpus of aligned, bilingual multi-word units;

calculating word association scores for word pairs in the corpus based on co-occurrence of words in the aligned units;

identifying hypothesized compounds based on an absence of one-to-one correspondence between words in the aligned units; and

training the machine translation system based on the word association scores and the hypothesized compounds;

wherein the words are surface forms of the words.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034541/0477 →