IP Library Granted Patent US 9,747,284
Granted Patent B2
US 9,747,284 · App. 14/835,783 · Granted Aug 29, 2017

Methods and systems for multi-engine machine translation

Inventors: James Peter Marciano (Amherst, NH); Dean S. Blodgett (Tewksbury, MA)
G06F17/289G06F17/28G06F17/2836G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,747,284
App. No.
14/835,783
Granted
Aug 29, 2017
Kind
B2
Abstract

A method for machine translation, comprising receiving a source string in a source language, an indication of a target language, and user identification information. The method includes using the user identification information to identify at least one memory with user-specific translation data. The method includes tokenizing the source string, using at least one processor, to produce a tokenized source string comprising any unique temporary textual elements associated with corresponding target textual elements during tokenization. The method includes obtaining a translated string from a translator, the translated string being at least a partial translation of the tokenized source string and including all of the unique temporary textual elements. The method includes generating an output string using at least one processor, the generation comprising replacing all of the unique temporary textual elements in the translated string with the associated target textual elements.

Claims (60)

1. A method for machine translation, comprising:

receiving a source string in a source language, an indication of a target language, and user identification information;

using the user identification information to identify at least one memory with user-specific translation data;

tokenizing the source string, using at least one processor, to produce a tokenized source string comprising at least one unique temporary textual elements associated with corresponding target textual elements during tokenization, wherein tokenizing further comprises:

searching the source string to identify at least one portion of the source string that meet a search criterion;

replacing the identified at least one portion of the source string with a corresponding unique temporary textual element; and

associating each unique temporary textual element with a corresponding target textual element;

obtaining a translated string from a translator, the translated string being at least a partial translation of the tokenized source string and including the at least one unique temporary textual elements; and

generating an output string using at least one processor, the generation comprising replacing the at least one unique temporary textual elements in the translated string with the associated target textual elements.

2. The method of claim 1 , wherein the user-specific translation data comprises at least one of: glossary data, translation memory data, or source-identification rule data, and the corresponding target textual elements are stored as part of the user-specific translation data or as part of sequestered data associated with the user identification information.

3. The method of claim 2 , wherein:

the glossary data is configured to associate strings in the source language with strings in the target language;

the translation memory data is configured to associate strings in the source language with strings in the target language; or

the source-identification rule data comprises regular expressions for identifying strings in the source language.

4. The method of claim 3 , wherein:

the strings associated by the glossary data are words or phrases,

the strings associated by the translation memory data are translation units comprising sentences and grammatically independent translation units,

the user-specific translation data is stored in at least one of: a database, a comma-separated values file, or a translation memory exchange file, and

the at least one temporary textual element is a unique textual or alphanumeric token.

5. The method of claim 1 , wherein searching the source string to identify all portions of the source string that meet a search criterion comprises:

searching to identify a match between a string in the glossary data and a portion of the source string;

searching to identify a match between a string in the translation memory data and a portion of the source string; or

applying at least one regular expression of the source-identification rule data to identify any portions of the source string as private information for sequestration.

6. The method of claim 5 , wherein searching to identify a match between a string in the glossary data and a portion of the source string ignores at least one of:

semantically insignificant differences between the string in the glossary data and portions of the source string,

the case of the alphabetical characters in the source string, or

hyphenation in the source-string.

7. The method of claim 5 , wherein the sequestered data comprises names, account numbers, addresses and telephone numbers.

8. The method of claim 1 , further comprising a normalization step that comprises:

searching the source string to identify all portions of the source string that meet a first search criterion specified in normalization data associated with the user identification information; and

replacing all identified portions of the source string with corresponding normalized strings specified in the normalization data.

9. The method of claim 1 , further comprising:

determining that one or more delimiters are used to indicate the presence of do-not-translate terms or phrases within the source string; and

identifying one or more substrings of the source string that are identified by the one or more delimiters as do-not-translate terms or phrases.

10. The method of claim 1 , wherein the at least one unique temporary textual elements comprises a concatenation of a unique numerical identifier of a normalized length and a randomized numerical string of a normalized length.

11. The method of claim 1 , further comprising training a machine translation engine using bilingual or monolingual training material.

12. A system for machine translation, comprising:

a memory associating a user identifier with user-specific translation data comprising at least one of: glossary data, translation memory data, or source-identification rule data;

a first application executing on a first one or more processors, the first application:

receiving a source string in a source language, an indication of a target language, and the user identifier; and

tokenizing the source string to produce a tokenized source string, wherein tokenizing further comprises:

searching the source string to identify at least one portion of the source string that meet a search criterion;

replacing the identified at least one portion of the source string with a corresponding unique temporary textual element; and

associating each unique temporary textual element with a corresponding target textual element; and

a second application executing on a second one or more processors without access to the user-specific translation data, the second application receiving the tokenized source string and outputting a translated string comprising at least a partial translation of the tokenized source string,

wherein the first application receives the translated string from the second application and generates an output string, the generation comprising replacing the unique temporary textual elements in the translated string with the associated target textual elements.

13. The system of claim 12 , further comprising a user interface for creating or updating one or more of glossary data, translation memory data, or source-identification rule data, the user interface comprising one or more interface elements for creating, updating and deleting regular expressions.

14. A system for machine translation, comprising:

a memory associating a user identifier with user-specific translation data;

a first application executing on a first set of one or more processors, the first application:

receiving a source string in a source language, an indication of a target language, and the user identifier;

normalizing the source string to produce a normalized source string; and

tokenizing the normalized source string to produce a tokenized source string comprising at least one temporary textual elements associated with a corresponding target textual elements during the tokenizing process, wherein tokenizing further comprises:

searching the source string to identify at least one portion of the source string that meet a search criterion;

replacing the identified at least one portion of the source string with a corresponding unique temporary textual element; and

associating each unique temporary textual element with a corresponding target textual element; and

a second application executing on a second set of one or more processors, the second application:

communicating the tokenized source string to a translator application executing on a third set of one or more processors; and

obtaining a translated string from the translator, the translated string being at least a partial translation of the tokenized source string and comprising the at least one temporary textual elements inserted during the tokenizing process,

wherein the first application generates an output string, the generation comprising replacing the at least one temporary textual element in the translated string with the associated target textual element.

Assignments (8)
SECURITY INTEREST Recorded Feb 27, 2026
From: LIONBRIDGE TECHNOLOGIES, LLC
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 073926/0472 →
CHANGE OF NAME Recorded Oct 27, 2021
From: LIONBRIDGE TECHNOLOGIES, INC.
To: LIONBRIDGE TECHNOLOGIES, LLC
Reel/Frame 057924/0375 →
SECURITY INTEREST Recorded Dec 27, 2019
From: LIONBRIDGE TECHNOLOGIES, INC.
To: KKR LOAN ADMINISTRATION SERVICES LLC
Reel/Frame 051376/0327 →
RELEASE OF SECURITY INTEREST IN PATENTS (FIRST LIEN) RECORDED AT R/F 041880/0033 Recorded Dec 27, 2019
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: LIONBRIDGE TECHNOLOGIES, INC.
Reel/Frame 051430/0754 →
RELEASE OF SECURITY INTEREST IN PATENTS (SECOND LIEN) RECORDED AT R/F 041880/0136 Recorded Dec 27, 2019
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: LIONBRIDGE TECHNOLOGIES, INC.
Reel/Frame 051431/0427 →
FIRST LIEN PATENT SECURITY AGREEMENT Recorded Mar 3, 2017
From: LIONBRIDGE TECHNOLOGIES, INC.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 041880/0033 →
SECOND LIEN PATENT SECURITY AGREEMENT Recorded Mar 3, 2017
From: LIONBRIDGE TECHNOLOGIES, INC.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 041880/0136 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2015
From: MARCIANO, JAMES PETER; BLODGETT, DEAN SCOTT
To: LIONBRIDGE TECHNOLOGIES, INC.
Reel/Frame 036428/0872 →
Continuity (3)
Continuation 13834331 · Mar 15, 2013
Provisional Application 61617341 · Mar 29, 2012
Related Publication 20150363394A1 · Dec 17, 2015