IP Library Granted Patent US 7,689,405
Granted Patent B2
US 7,689,405 · App. 10/150,532 · Granted Mar 30, 2010

Statistical method for building a translation memory

Assignee: Language Weaver, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,689,405
App. No.
10/150,532
Granted
Mar 30, 2010
Kind
B2
Abstract

A statistical translation memory (TMEM) may be generated by training a translation model with a naturally generated TMEM. A number of tuples may be extracted from each translation pair in the TMEM. The tuples may include a phrase in a source language and a corresponding phrase in a target language. The tuples may also include probability information relating to the phrases generated by the translation model.

Claims (27)

1. A method comprising:

training a statistical translation model on a plurality of translation pairs, each translation pair including a text segment in a source language and a corresponding text segment in a target language;

extracting a plurality of tuples from each of a plurality of said translation pairs, each tuple comprising a pair of phrases and a corresponding word level alignment between said pair of phrases, the pair of phrases including a phrase in the source language comprising at least two words and a corresponding phrase in the target language comprising at least two words, the corresponding word level alignment based on the trained statistical translation model; and

constructing automatically a statistical translation memory from said plurality of tuples, wherein the statistical translation memory encodes a plurality of said pair of phrases and the corresponding word level alignment between each said pair.

2. The method of claim 1 , wherein the phrase in the target language corresponds to the phrase in the source language occurring most frequently in the plurality of tuples.

3. The method of claim 1 , wherein the phrase in the target language corresponds to the phrase in the source language having an alignment of highest probability.

4. The method of claim 1 , further comprising:

judging a correctness of the plurality of tuples; and

selecting a tuple as a translation equivalent in response to said judgment.

5. An automatically constructed statistical translation memory comprising:

a plurality of tuples extracted from each of a plurality of translation pairs, each tuple comprising a pair of phrases and a corresponding word level alignment between said pair of phrases, the pair of phrases including a phrase in the source language comprising at least two words and a corresponding phrase in the target language comprising at least two words, the corresponding word level alignment based on a statistical translation model trained on the translation pairs.

6. The automatically constructed statistical translation memory of claim 5 , wherein the phrase in the target language corresponds to the phrase in the source language having an alignment of highest probability.

7. The automatically constructed statistical translation memory of claim 5 , wherein the phrase in the target language corresponds to the phrase in the source language occurring most frequently in the plurality of tuples.

8. Apparatus comprising:

a statistical translation model operative to train on a plurality of translation pairs, each translation pair including a text segment in a source language and a corresponding text segment in a target language;

an extraction module operative to extract a plurality of tuples from each of a plurality of said translation pairs, each tuple comprising a pair of phrases and a corresponding word level alignment between said pair of phrases, the pair of phrases including a phrase in the source language comprising at least two words and a corresponding phrase in the target language comprising at least two words, the corresponding word level alignment based of the training of the statistical translation model; and

an automatically constructed statistical translation memory based on said plurality of tuples, wherein the statistical translation memory encodes a plurality of said pair of phrases and the corresponding word level alignment between each said pair.

9. The apparatus of claim 8 , wherein the extraction module is operative to select a phrase occurring most frequently in the plurality of tuples.

10. The apparatus of claim 8 , wherein the extraction module is operative to select a phrase having a highest probability of being a correct translation of said phrase in the source language.

11. The apparatus of claim 8 , wherein the extraction module is operative to select a phrase having an alignment of highest probability with said phrase in the source language.

12. An article comprising a machine-readable medium including instructions operative to cause a machine to perform:

training a statistical translation model on a plurality of translation pairs, each translation pair including a text segment in a source language and a corresponding text segment in a target language;

extracting a plurality of tuples from each of a plurality of said translation pairs, each tuple comprising a pair of phrases and a corresponding word level alignment between said pair of phrases, the pair of phrases including a phrase in the source language comprising at least two words and a corresponding phrase in the target language comprising at least two words, the corresponding word level alignment based on the trained statistical translation model; and

constructing automatically a statistical translation memory from said plurality of tuples, wherein the statistical translation memory encodes a plurality of said pair of phrases and the corresponding word level alignment between each said pair.

13. The article of claim 12 , wherein the phrase in the target language corresponds to the phrase in the source language occurring most frequently in the plurality of tuples.

14. The article of claim 12 , wherein said constructing comprises selecting a phrase having a highest probability of being a correct translation of said phrase in the source language.

15. The article of claim 12 , wherein constructing comprises selecting a phrase having an alignment of highest probability with said phrase in the source language.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2002
From: MARCU, DANIEL
To: UNIVERSITY OF SOUTHERN CALIFORNIA
Reel/Frame 012998/0501 →
Continuity (2)
Provisional Application 6029185200 · May 17, 2001
Related Publication 20030009322A1 · Jan 9, 2003