IP Library Granted Patent US 7,496,496
Granted Patent B2
US 7,496,496 · App. 11/725,435 · Granted Feb 24, 2009

System and method for machine learning a confidence metric for machine translation

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,496,496
App. No.
11/725,435
Granted
Feb 24, 2009
Kind
B2
Abstract

A machine translation system is trained to generate confidence scores indicative of a quality of a translation result. A source string is translated with a machine translator to generate a target string. Features indicative of translation operations performed are extracted from the machine translator. A trusted entity-assigned translation score is obtained and is indicative of a trusted entity-assigned translation quality of the translated string. A relationship between a subset of the extracted features and the trusted entity-assigned translation score is identified.

Claims (48)

1. A method of training a machine translation computing device to generate confidence scores indicative of a quality of a translation result, comprising:

translating a source string with a machine translation computing device to generate a target string;

extracting features from the machine translator, indicative of performance of translation steps in the machine translator;

obtaining a trusted entity-assigned translation score indicative of a trusted entity-assigned translation quality of the target string;

identifying a relationship between a subset of the extracted features and the trusted entity-assigned translation score;

parsing the source string into a source intermediate linguistic structure indicative of a meaning of the source string;

wherein translating includes translating the source intermediate linguistic structure to a target intermediate linguistic structure;

wherein translating the source intermediate linguistic structure comprises identifying mappings, in a mapping database, that map portions of the source intermediate linguistic structure to portions of the target intermediate linguistic structure; and

wherein extracting one or more features indicative of a quality of transiating the source intermediate linguistic structure comprises extracting a feature indicative of a number of identified mappings.

2. The method of claim 1 wherein extracting features comprises:

extracting a feature indicative of a size of the source string.

3. The method of claim 1 wherein extracting one or more features indicative of a quality of translating the source intermediate linguistic structure comprises:

extracting a feature indicative of a size of identified mappings.

4. The method of claim 1 wherein extracting one or more features indicative of a quality of translating the source intermediate linguistic structure comprises:

extracting a feature indicative of a frequency with which mappings where identified.

5. The method of claim 1 wherein each mapping has an associated confidence score and wherein extracting one or more features indicative of a quality of translating the source intermediate linguistic structure comprises:

extracting a feature indicative of the confidence scores associated with the identified mappings.

6. The method of claim 1 wherein translating the source intermediate linguistic structure comprises:

translating portions of the source intermediate linguistic structure using another translation device other than the mapping database.

7. The method of claim 1 wherein extracting features comprises:

extracting a feature indicative of the translation device used.

8. The method of claim 7 wherein extracting features comprises:

extracting a feature indicative of an amount of the source string translated with another translation device.

9. The method of claim 7 wherein extracting features comprises:

extracting a feature indicative of a type of words translated with another translation device.

10. The method of claim 1 wherein extracting features comprises:

calculating a perplexity of the target string with a statistical language model.

11. A method of training a machine translation computing device to generate confidence scores indicative of a quality of a translation result, comprising:

translating a source string with a machine translation computing device to generate a target string;

extracting features from the machine translator, the features being indicative of performance of translation steps in the machine translator;

obtaining a trusted entity-assigned translation score indicative of a trusted entity-assigned translation quality of the target string;

identifying a relationship between a subset of the extracted features and the trusted entity- assigned translation score; and

wherein translating includes parsing the source string into a source intermediate linguistic structure indicative of a meaning of the source string;

wherein extracting features comprises extracting one or more features indicative of a quality of parsing;

wherein translating includes translating the source intermediate linguistic structure into a target intermediate linguistic structure;

wherein extracting features comprises extracting one or more features indicative of a quality of translating the source intermediate linguistic structure into the target intermediate linguistic structure;

wherein translating the source intermediate linguistic structure comprises identifying mappings, in a mapping database, that map portions of the source intermediate linguistic structure to portions of the target intermediate linguistic structure;

wherein translating the source intermediate linguistic structure further comprises translating portions of the source intermediate linguistic structure using another translation device other than the mapping database; and

wherein some words in the source intermediate linguistic structure remain untranslated, and wherein extracting one or more features indicative of a quality of translating the source intermediate linguistic structure comprises extracting a feature indicative of untranslated words in the source intermediate linguistic structure.

12. A method of training a machine translation computing device to generate confidence scores indicative of a quality of a translation result, comprising:

translating a source string with a machine translation computing device to generate a target string;

extracting features from the machine translator, indicative of performance of translation steps in the machine translator;

obtaining a trusted entity-assigned translation score indicative of a trusted entity-assigned translation quality of the target string;

identifying a relationship between a subset of the extracted features and the trusted entity- assigned translation score;

wherein translating includes parsing the source string into a source intermediate linguistic structure indicative of a meaning of the source string;

wherein translating includes translating the source intermediate linguistic structure to a target intermediate linguistic structure;

wherein extracting features comprises extracting features indicative of a quality of translating the source intermediate linguistic structure into the target intermediate linguistic structure; and

wherein some words in the source intermediate linguistic structure remain untranslated, and wherein extracting one or more features indicative of a quality of translating the source intermediate linguistic structure comprises extracting a feature indicative of untranslated words in the source intermediate linguistic structure.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034542/0001 →