IP Library Granted Patent US 9,213,694
Granted Patent B2
US 9,213,694 · App. 14/051,175 · Granted Dec 15, 2015

Efficient online domain adaptation

Inventors: Felix Hieber (Heidelberg, DE); Jonathan May (Culver City, CA)
Assignee: Language Weaver, Inc.
G06F17/2854G06F17/2827
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,213,694
App. No.
14/051,175
Granted
Dec 15, 2015
Kind
B2
Abstract

Systems and methods for efficient online domain adaptation are provided herein. Methods may include receiving a post-edited machine translated sentence pair, updating a machine translation model by adjusting translation weights for a translation memory and a language model while generating test machine translations of the machine translated sentence pair until one of the test machine translations approximately matches the post-edits for the machine translated sentence pair, and retranslating the remaining machine translation sentence pairs that have yet to be post-edited using the updated machine translation model.

Claims (48)

1. A method of immediately updating a machine translation system with post-edits during translation of a document, using a machine translation system that comprises a processor and a memory for storing logic that is executed by the processor to perform the method, comprising:

receiving a post-edited machine translated sentence pair, the post-edited machine translated sentence pair comprising a source sentence unit and a post-edited target sentence unit;

updating a machine translation model by:

performing an alignment of the post-edits of the machine translated sentence pair to generate phrases; and

adding the phrases to the machine translation model;

adapting a language model from the post-edited target sentence unit;

calculating translation statistics for the post-edits;

adjusting translation weights using the translation statistics while generating test machine translations of the machine translated sentence pair until one of the test machine translations approximately matches the post-edits for the machine translated sentence pair; and

retranslating remaining machine translation sentence pairs that have yet to be post-edited using the updated machine translation model and the adjusted translation weights.

2. The method according to claim 1 , further comprising:

receiving a document for translation from a source language into a target language; and

performing a machine translation of the document to generate a set of machine translated sentence pairs.

3. The method according to claim 1 , wherein the updating of the machine translation model further comprises tokenizing the post-edits of the machine translated sentence pair.

4. The method according to claim 3 , wherein the updating of the machine translation model further comprises updating a vocabulary of source sentence units with unknown source sentence units included in the post-edits of the machine translated sentence pair and updating a vocabulary of target sentence units with unknown target sentence units included in the post-edits of the machine translated sentence pair.

5. The method according to claim 4 , further comprising extracting fractional counts from the post-edits of the machine translated sentence pair and adding the extracted fractional counts to a fractional count table.

6. The method according to claim 5 , further comprising adjusting probability distributions for the source sentence unit of the post-edits of the machine translated sentence pair.

7. The method according to claim 1 , further comprising updating a phrase table with counts that define a number of occurrences of phrases in the phrase table; and reordering feature values for the phrases in the phrase table based upon the counts.

8. The method according to claim 1 , wherein the translation weights for the machine translation system are adjusted using discriminative ridge regression.

9. The method according to claim 1 , wherein the adapting of the language model includes executing an ngram-count of the post-edited machine translated sentence pair to update a count file that comprises counts for input sentence pairs; and recompiling the language model using a smoothing algorithm.

10. The method according to claim 1 , wherein the alignment includes both forward and reverse alignments of the post-edited machine translated sentence pair.

11. A machine translation system that immediately incorporates post-edits into a machine translation model during translation of a document, the machine translation system comprising:

a processor; and

a memory for storing logic that is executed by the processor to:

receive a post-edit of a target sentence unit of a machine translated sentence pair of a set of machine translated sentence pairs, wherein a machine translated sentence pair comprises a source sentence unit and the target sentence unit;

receive a post-edited machine translated sentence pair, the post-edited machine translated sentence pair comprising a post-edited source sentence unit and a post-edited target sentence unit;

update the machine translation model by:

performing an alignment of the post-edits of the machine translated sentence pair to generate phrases; and

adding the phrases to the machine translation model;

adapt a language model from the post-edited target sentence unit;

calculate translation statistics for the post-edits;

adjust translation weights using the translation statistics while generating test machine translations of the machine translated sentence pair until one of the test machine translations approximately matches the post-edits for the machine translated sentence pair; and

retranslate remaining machine translation sentence pairs that have yet to be post-edited using the updated machine translation model.

12. The machine translation system according to claim 11 , wherein the processor further executes the logic to:

receive the document for translation from a source language into a target language; and

perform a machine translation of the document to generate the set of machine translated sentence pairs.

13. The machine translation system according to claim 11 , wherein the machine translation system updates the machine translation model by further tokenizing the post-edits of the machine translated sentence pair.

14. The machine translation system according to claim 13 , wherein the machine translation system updates the machine translation model by updating a vocabulary of source sentence units with unknown source sentence units included in the post-edits of the machine translated sentence pair and updating a vocabulary of target sentence units with unknown target sentence units included in the post-edits of the machine translated sentence pair.

15. The machine translation system according to claim 14 , wherein the processor further executes the logic to extract fractional counts from the post-edits of the machine translated sentence pair and adding the extracted fractional counts to a fractional count table.

16. The machine translation system according to claim 15 , wherein the processor further executes the logic to adjust probability distributions for the source sentence unit of the post-edits of the machine translated sentence pair.

17. The machine translation system according to claim 11 , wherein the processor further executes the logic to update a phrase table with counts that define a number of occurrences of phrases in the phrase table; and reorder feature values for the phrases in the phrase table based upon the counts.

18. The machine translation system according to claim 11 , wherein the translation weights for the machine translation system are adjusted using discriminative ridge regression.

19. The machine translation system according to claim 11 , wherein the machine translation system is configured to adapt the language model by executing an ngram-count of the post-edited machine translated sentence pair to update a count file that comprises counts for input sentence pairs; and recompiling the language model using a smoothing algorithm.

20. The machine translation system according to claim 11 , wherein the machine translation system is configured to:

receive post-edits for a retranslated machine translation sentence pair;

re-updat the machine translation model;

calculate translation statistics for the post-edits of the retranslated machine translation sentence pair;

adjust translation weights using the translation statistics while generating test machine translations of the retranslated machine translated sentence pair until one of the test machine translations approximately matches the post-edits for the retranslated machine translation sentence pair; and

retranslate any remaining retranslated machine translation sentence pairs that have yet to be post-edited using the updated machine translation model and the translation weights.

Assignments (2)
MERGER Recorded Feb 16, 2016
From: LANGUAGE WEAVER, INC.
To: SDL INC.
Reel/Frame 037745/0391 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2013
From: HIEBER, FELIX; MAY, JONATHAN
To: LANGUAGE WEAVER, INC.
Reel/Frame 031531/0931 →
Continuity (1)
Related Publication 20150106076A1 · Apr 16, 2015