IP Library Granted Patent US 10,311,148
Granted Patent B2
US 10,311,148 · App. 15/652,949 · Granted Jun 4, 2019

Methods and systems for multi-engine machine translation

Inventors: James Peter Marciano (Amherst, NH); Dean S. Blodgett (Tewksbury, MA)
Assignee: Lionbridge Technologies, Inc.
G06F17/289G06F17/28G06F17/2836G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,311,148
App. No.
15/652,949
Granted
Jun 4, 2019
Kind
B2
Abstract

A method for machine translation, comprising receiving a source string in a source language, an indication of a target language, and user identification information. The method includes using the user identification information to identify at least one memory with user-specific translation data. The method includes tokenizing the source string, using at least one processor, to produce a tokenized source string comprising any unique temporary textual elements associated with corresponding target textual elements during tokenization. The method includes obtaining a translated string from a translator, the translated string being at least a partial translation of the tokenized source string and including all of the unique temporary textual elements. The method includes generating an output string using at least one processor, the generation comprising replacing all of the unique temporary textual elements in the translated string with the associated target textual elements.

Claims (59)

1. A method for machine translation, comprising:

providing a translation profile interface associated with an application that allows an authorized user to modify user-specific translation data;

receiving, by the translation profile interface, from the authorized user, an update to the user-specific translation data, the update including an instruction for customizing a translation;

receiving a source string in a source language, an indication of a target language, and user identification information associated with the user;

using the user identification information to identify at least one memory with the user-specific translation data;

tokenizing the source string, using at least one processor, to produce a tokenized source string comprising any unique temporary textual elements associated with corresponding target textual elements during tokenization;

obtaining a translated string from a translator, the translated string being at least a partial translation of the tokenized source string and including all of the unique temporary textual elements; and

generating an output string using at least one processor, the generation comprising replacing all of the unique temporary textual elements in the translated string with the associated target textual elements, the generated output string complying with the instruction for customizing the translation.

2. The method of claim 1 , wherein the user-specific translation data comprises at least one of: glossary data, translation memory data, or source-identification rule data, and the corresponding target textual elements are stored as part of at least one of the user-specific translation data and sequestered data associated with the user identification information.

3. The method of claim 2 , wherein the tokenizing comprises:

searching the source string to identify all portions of the source string that meet a search criterion;

replacing all the identified portions of the source string with corresponding unique temporary textual elements; and

associating each unique temporary textual element with a corresponding target textual element.

4. The method of claim 3 , wherein:

the glossary data is configured to associate strings in the source language with strings in the target language;

the translation memory data is configured to associate strings in the source language with strings in the target language; or

the source-identification rule data comprises regular expressions for identifying strings in the source language.

5. The method of claim 4 , wherein:

the strings associated by the glossary data are words or phrases,

the strings associated by the translation memory data are translation units comprising sentences and grammatically independent translation units,

the user-specific translation data is stored in at least one of: a database, a comma-separated values file, or a translation memory exchange file, and

the at least one temporary textual element is a unique textual or alphanumeric token.

6. The method of claim 3 , wherein searching the source string to identify all portions of the source string that meet a search criterion comprises:

searching to identify a match between a string in the glossary data and a portion of the source string;

searching to identify a match between a string in the translation memory data and a portion of the source string; or

applying at least one regular expression of the source-identification rule data to identify any portions of the source string as private information for sequestration.

7. The method of claim 6 , wherein searching to identify a match between a string in the glossary data and a portion of the source string ignores at least one of:

semantically insignificant differences between the string in the glossary data and portions of the source string,

the case of the alphabetical characters in the source string, or

hyphenation in the source-string.

8. The method of claim 3 , wherein the sequestered data comprises names, account numbers, addresses and telephone numbers.

9. The method of claim 3 , further comprising a normalization step that comprises:

searching the source string to identify all portions of the source string that meet a first search criterion specified in normalization data associated with the user identification information; and

replacing all identified portions of the source string with corresponding normalized strings specified in the normalization data.

10. The method of claim 3 , further comprising:

determining that one or more delimiters are used to indicate the presence of do-not-translate terms or phrases within the source string; and

identifying one or more substrings of the source string that are identified by the one or more delimiters as do-not-translate terms or phrases.

11. The method of claim 3 , wherein one or more of the unique temporary textual elements comprises a concatenation of a unique numerical identifier of a normalized length and a randomized numerical string of a normalized length.

12. The method of claim 1 , further comprising training a machine translation engine using bilingual or monolingual training material.

13. A system for machine translation, comprising:

a translation profile interface receiving, from an authorized user, an update to user-specific translation data, the update including an instruction for customizing a translation;

a memory associating a user identifier with user-specific translation data comprising at least one of: glossary data, translation memory data, or source-identification rule data;

a first application executing on a first one or more processors, the first application:

receiving a source string in a source language, an indication of a target language, and the user identifier; and

tokenizing the source string to produce a tokenized source string; and

a second application executing on a second one or more processors without access to the user-specific translation data, the second application receiving the tokenized source string and outputting a translated string comprising at least a partial translation of the tokenized source string,

wherein the first application receives the translated string from the second application and generates an output string, the generation comprising replacing all temporary textual elements in the translated string with associated target textual elements, the generated output string complying with the instruction for customizing the translation.

14. The system of claim 13 , further comprising a user interface for creating or updating one or more of glossary data, translation memory data, or source-identification rule data, the user interface comprising one or more interface elements for creating, updating and deleting regular expressions.

15. A system for machine translation, comprising:

a translation profile interface receiving, from an authorized user, an update to user-specific translation data, the update including an instruction for customizing a translation;

a memory associating a user identifier with user-specific translation data;

a first application executing on a first set of one or more processors, the first application:

receiving a source string in a source language, an indication of a target language, and the user identifier;

normalizing the source string to produce a normalized source string; and

tokenizing the normalized source string to produce a tokenized source string comprising all temporary textual elements associated with corresponding target textual elements during the tokenizing process; and

a second application executing on a second set of one or more processors, the second application:

communicating the tokenized source string to a translator application executing on a third set of one or more processors; and

obtaining a translated string from the translator, the translated string being at least a partial translation of the tokenized source string and comprising all the temporary textual elements inserted during the tokenizing process,

wherein the first application generates an output string, the generation comprising replacing each temporary textual element in the translated string with an associated target textual element, the generated output string complying with the instruction for customizing the translation.

Assignments (4)
SECURITY INTEREST Recorded Feb 27, 2026
From: LIONBRIDGE TECHNOLOGIES, LLC
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 073926/0472 →
CHANGE OF NAME Recorded Oct 27, 2021
From: LIONBRIDGE TECHNOLOGIES, INC.
To: LIONBRIDGE TECHNOLOGIES, LLC
Reel/Frame 057924/0375 →
SECURITY INTEREST Recorded Dec 27, 2019
From: LIONBRIDGE TECHNOLOGIES, INC.
To: KKR LOAN ADMINISTRATION SERVICES LLC
Reel/Frame 051376/0327 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2017
From: MARCIANO, JAMES PETER; BLODGETT, DEAN SCOTT
To: LIONBRIDGE TECHNOLOGIES, INC.
Reel/Frame 043732/0147 →
Continuity (4)
Continuation 14835783 · Aug 26, 2015
Continuation 13834331 · Mar 15, 2013
Provisional Application 61617341 · Mar 29, 2012
Related Publication 20170344537A1 · Nov 30, 2017