IP Library › Granted Patent US 11,551,013
Granted Patent B1
US 11,551,013 · App. 16/806,705 · Granted Jan 10, 2023

Automated quality assessment of translations

Inventors: Prabhakar Gupta (Delhi, IN); Anil Kumar Nelakanti (Bangalore, IN)
Assignee: Amazon Technologies, Inc.
G06F40/58G06F16/7867
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,551,013
App. No.
16/806,705
Granted
Jan 10, 2023
Kind
B1
Abstract

Technologies are provided for automated quality assessment of translations. In some embodiments, quality of a translation can be assessed by generating a machine-learning (ML) model that classifies the translation as pertaining to one of three quality categories. A first quality category can include, for example, translations that are deemed satisfactory. A second quality category can include, for example, translations that are deemed subject to edition prior to being deemed satisfactory. A third quality category can include, for example, translations that are deemed unsatisfactory. The generated ML model can then be applied to the translation and a corresponding sentence in a source language in order to classify the translation as pertaining to one of the three categories.

Claims (73)

1. A method, comprising:

receiving, by a computing system comprising at least one processor, first unlabeled data defining video subtitles in a source natural language;

receiving, by the computing system, second unlabeled data defining translations of the video subtitles to a target natural language;

generating, by the computing system, using the first unlabeled data and the second unlabeled data, first labeled data defining first translations corresponding to a first quality category, wherein the first quality category includes satisfactory translations;

generating, by the computing system, using the first unlabeled data and the second unlabeled data, second labeled data defining second translations corresponding to a second quality category, wherein the second quality category includes unacceptable translations subject to edition prior to being deemed satisfactory;

generating, by the computing system, using the first unlabeled data and the second unlabeled data, third labeled data defining third translations corresponding to a third quality category; wherein the third quality category includes unsatisfactory translations;

training, by the computing system, using the first labeled data, the second labeled data, and the third labeled data, a machine-learning model to classify a translation of a video subtitle in the source natural language to the target natural language as pertaining to the first quality category, the second quality category, or the third quality category;

receiving, by the computing system, first string data defining a particular video subtitle in the source natural language;

applying a first bidirectional long short-term memory (LSTM) encoder to the first string data;

receiving, by the computing system, second string data defining a translation of the particular video subtitle to the target natural language;

applying a second bidirectional (LSTM) encoder to the second string data; and

determining, by the computing system, a category of the translation of the particular subtitle to the target natural language by applying a convolutional neural network to the concatenated output from the first LTSM encoder and the second LTSM encoder, the category representing a quality assessment of the translation and corresponding to one of the first quality category, the second quality category, or the third quality category.

2. The method of claim 1 , wherein the source natural language is English and the target natural language is one of German, French, Spanish, Portuguese, and Italian.

3. The method of claim 1 , wherein the generating the first labeled data comprises,

training, using the first unlabeled data and the second unlabeled data, a statistical classifier to generate a score representing one of the first quality category, the second quality category, or the third quality category;

generating a defined score for a source-target-language data pair by applying the statistical classifier, the source-target-language data pair having a first datum from the first unlabeled data and a second datum from the second unlabeled data;

assigning, using the defined score, a label indicative of the first quality category to the target-target-language data pair.

4. The method of claim 3 , wherein the generating the first labeled data further comprises,

selecting a first video subtitle from the video subtitles in the source natural language; and

generating a translation of the first video subtitle by applying a neural machine translation model to the first video subtitle; and

assigning a second label indicative of the first quality category to the translation of the first video subtitle.

5. The method of claim 1 , wherein the generating the second labeled data further comprises,

selecting a translation of a second particular video subtitle, the translation of the second particular video subtitle being deemed satisfactory relative to a human curated translation, and at least one of:

modifying the translation of the second particular video subtitle by adding a subtitle-caption to the second particular video subtitle, or

switching an ordering of a first defined word in the translation of the second particular video and a second defined word in the translation of the second particular video.

6. A method, comprising:

receiving, by a computing system comprising at least one processor, first unlabeled data defining text strings in a first natural language;

receiving, by the computing system, second unlabeled data defining translations of the text strings to a target natural language;

generating, by the computing system, using the first unlabeled data and the second unlabeled data, first labeled data defining first translations corresponding to a first quality category;

generating, by the computing system, using the first unlabeled data and the second unlabeled data, second labeled data defining second translations corresponding to a second quality category;

generating, by the computing system, using the first unlabeled data and the second unlabeled data, third labeled data defining third translations corresponding to a third quality category;

generating, by the computing system, using the first labeled data, the second labeled data, and the third labeled data, a machine-learning model to classify a translation of a text string in the first natural language to the target natural language as pertaining to the first quality category, the second quality category, or the third quality category;

receiving first string data defining a particular text string in the first natural language;

applying a first bidirectional long short-term memory (LSTM) encoder to the first string data;

receiving second string data defining a translation of the particular text string to the target natural language;

applying a second bidirectional (LSTM) encoder to the second string data; and

determining a category of the translation of the particular text string to the target natural language by applying a convolutional neural network to the concatenated output from the first LTSM encoder and the second LTSM encoder, the category representing a quality assessment of the translation and corresponding to one of the first quality category, the second quality category, or the third quality category.

7. The method of claim 6 , wherein one of the first quality category, the second quality category, or the third quality category includes paraphrased translations or translations subject to edition after being generated.

8. The method of claim 6 , wherein the generating the first labeled data comprises,

training, using the first unlabeled data and the second unlabeled data, a statistical classifier to generate a score representing one of the first quality category, the second quality category, or the third quality category;

generating a defined score for a source-target-language data pair by applying the statistical classifier, the source-target-language data pair having a first datum from the unlabeled data and a second datum from the second unlabeled data; and

assigning, using the defined score, a label indicative of the first quality category to the source-target-language data pair.

9. The method of claim 6 , wherein the generating the first labeled data further comprises, selecting a first video subtitle from the text strings in the first natural language; and generating a translation of the first text string by applying a neural machine translation model to the text string; and assigning a second label indicative of the first quality category to the translation of the first text string.

10. The method of claim 6 , wherein the generating the second labeled data further comprises,

selecting a translation of a particular text string, the translation deemed satisfactory relative to a human curated translation; and

modifying the translation by applying a defined rule, the modified translation subject to edition prior to being deemed satisfactory.

11. A computing system, comprising:

at least one processor; and

at least one memory device having computer-executable instructions stored thereon that, in response to execution by the at least one processor, cause the computing system to perform operation comprising:

receiving, by a computing system comprising at least one processor, first unlabeled data defining text strings in a first natural language;

receiving, by the computing system, second unlabeled data defining translations of the text strings to a target natural language;

generating, by the computing system, using the first unlabeled data and the second unlabeled data, first labeled data defining first translations corresponding to a first quality category;

generating, by the computing system, using the first unlabeled data and the second unlabeled data, second labeled data defining second translations corresponding to a second quality category;

generating, by the computing system, using the first unlabeled data and the second unlabeled data, third labeled data defining third translations corresponding to a third quality category;

generating, by the computing system, using the first labeled data, the second labeled data, and the third labeled data, a machine-learning model to classify a translation of a text string in the first natural language to the target natural language as pertaining to the first quality category, the second quality category, or the third quality category;

receiving first string data defining a particular text string in the first natural language;

applying a first bidirectional long short-term memory (LSTM) encoder to the first string data;

receiving second string data defining a translation of the particular text string to the target natural language;

applying a second bidirectional (LSTM) encoder to the second string data; and

determining, by the computing system, a category of the translation of the particular text string to the target natural language by applying a convolutional neural network to the concatenated output from the first LTSM encoder and the second LTSM encoder, the category representing a quality assessment of the translation and corresponding to one of the first quality category, the second quality category, or the third quality category.

12. The computing system of claim 11 , wherein one of the first quality category, the second quality category, or the third quality category includes paraphrased translations or translations subject to edition after being generated.

13. The computing system of claim 11 , wherein the generating the first labeled data comprises,

training, using the first unlabeled data and the second unlabeled data, a statistical classifier to generate a score representing one of the first quality category, the second quality category, or the third quality category;

generating a defined score for a source-target-language data pair by applying the statistical classifier, the source-target-language data pair having a first datum from the unlabeled data and a second datum from the second unlabeled data; and

assigning, using the defined score, a label indicative of the first quality category to the source-target-language data pair.

14. The computing system of claim 13 , wherein the generating the first labeled data further comprises,

selecting a first video subtitle from the text strings in the first natural language; and

generating a translation of the first text string by applying a neural machine translation model to the text string; and

assigning a second label indicative of the first quality category to the translation of the first text string.

15. The computing system of claim 11 , wherein the generating the second labeled data further comprises,

selecting a translation of a particular text string, the translation deemed satisfactory relative to a human curated translation; and

modifying the translation by applying a defined rule, the modified translation subject to edition prior to being deemed satisfactory.

16. The computing system of claim 15 , wherein the defined rule dictates to add a subtitle-caption to the translation of the particular text string.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2020
From: GUPTA, PRABHAKAR; NELAKANTI, ANIL KUMAR KUMAR
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 051993/0556 →
Cited By (4)
US 12,242,819 US 12,272,383 US 12,462,094 US 12,737,562