IP Library Granted Patent US 9,141,606
Granted Patent B2
US 9,141,606 · App. 13/834,331 · Granted Sep 22, 2015

Methods and systems for multi-engine machine translation

Inventors: James Peter Marciano (Amherst, NH); Dean Scott Blodgett (Tewksbury, MA)
Assignee: Lionbridge Technologies, Inc.
G06F17/2836G06F17/28G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,141,606
App. No.
13/834,331
Granted
Sep 22, 2015
Kind
B2
Abstract

Systems and methods for multi-engine machine translations are disclosed. Exemplary methods and systems involve normalizing and/or tokenizing a source string using user-specific translation data. The user-specific translation data may include glossary data, translation memory data, and rule data for use in customizing translations and sequestering sensitive data during the translation process. The disclosed methods and systems also involve using one or more machine translation engines to obtain a translation of the normalized and/or tokenized source string.

Claims (44)

1. A method for machine translation, comprising:

receiving a translation request comprising user identification information, an indication of a source language, an indication of a target language, and a source string in the source language;

decoding the translation request using at least one processor, wherein the source string is spliced into multiple source strings if the source string is determined to exceed a character limit set by the translator;

normalizing the source string to produce a normalized source string;

tokenizing the normalized source string to produce a tokenized source string;

communicating the tokenized source string to a translator;

obtaining a translated string from the translator, the translated string being at least a partial translation of the tokenized source string and comprising any temporary textual elements inserted during tokenization; and

generating an output string using at least one processor, the generation comprising replacing all temporary textual elements in the translated string with associated target textual elements.

2. The method of claim 1 , wherein the decoding comprises:

using the user identification information to identify at least one memory comprising user-specific translation data, the user-specific translation data comprising at least one of: glossary data, translation memory data, normalization data and source-identification rule data.

3. The method of claim 2 , wherein normalizing comprises:

searching the source string to identify all portions of the source string that meet a first search criterion; and

replacing all identified portions of the source string with corresponding normalized strings from the normalization data.

4. The method of claim 3 , wherein tokenizing comprises:

searching to identify all portions of the normalized source string that meet at least a second search criterion;

replacing all identified portions of the normalized source string with corresponding unique temporary textual elements; and

associating each temporary textual element with a target textual element stored as part of at least one of: the user-specific translation data and sequestered data associated with the user identification information.

5. The method of claim 4 , wherein searching to identify all portions of the normalized source string that meet at least the second search criterion comprises at least one of:

searching to identify a match between an item in the glossary data with the at least one portion of the normalized source string;

searching to identify a match between an item in the translation memory data with the at least one portion of the normalized source string; and

applying a regular expression stored as part of the source-identification rule data to identify any portions of the normalized source string as private information for sequestration.

6. The method of claim 5 , wherein:

a matched item in the glossary data is at least one of a word and a phrase in the source language;

a matched item in the translation memory data is a translation unit comprising sentences and grammatically independent translation units in the source language; and

a portion of the source string identified as private information for sequestration is stored in a protected memory as sequestered data associated with the user identification information.

7. The method of claim 3 , wherein searching to identify all portions of the source string that meet the first search criterion comprises at least one of:

searching to identify all matches between the portions of the source string and a string specified in the normalization data; and

applying at least one regular expression specified in the normalization data.

8. A method for invoking a grammatically-sensitive or basic tokenization process, the method comprising:

searching a source string, using a processor, to identify a substring that matches a data item associated with user-specific translation data;

performing a first determination regarding whether the data item has an associated grammatical flag;

based on a result of the first determination, performing at least one of:

a second determination regarding whether grammatically-sensitive token data associated with the user-specific translation data contains a grammatical flag that matches the grammatical flag associated with the data item,

and a basic tokenization process;

based on a result of the second determination, performing at least one of:

a third determination regarding whether there exists at least one grammatically-sensitive token associated with the matched grammatical flag that is not already present in the source string, and

a basic tokenization process; and

based on a result of the third determination, performing at least one of:

replacing the identified substring in the source string with the at least one grammatically-sensitive token associated with the matched grammatical flag, and

associating the at least one grammatically-sensitive token in a memory with the identified substring, and

a basic tokenization process.

9. The method of claim 8 , wherein:

grammatically-sensitive token data comprises information on a grammatical flag, an indication of a designated machine translation engine, and one or more grammatically-sensitive tokens, and

the grammatical flag associated with the data item encodes grammatical attributes comprising a part of speech, a gender and a number indicating grammatical singularity or plurality.

Assignments (10)
SECURITY INTEREST Recorded Feb 27, 2026
From: LIONBRIDGE TECHNOLOGIES, LLC
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 073926/0472 →
CHANGE OF NAME Recorded Oct 27, 2021
From: LIONBRIDGE TECHNOLOGIES, INC.
To: LIONBRIDGE TECHNOLOGIES, LLC
Reel/Frame 057924/0375 →
SECURITY INTEREST Recorded Dec 27, 2019
From: LIONBRIDGE TECHNOLOGIES, INC.
To: KKR LOAN ADMINISTRATION SERVICES LLC
Reel/Frame 051376/0327 →
RELEASE OF SECURITY INTEREST IN PATENTS (FIRST LIEN) RECORDED AT R/F 041880/0033 Recorded Dec 27, 2019
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: LIONBRIDGE TECHNOLOGIES, INC.
Reel/Frame 051430/0754 →
RELEASE OF SECURITY INTEREST IN PATENTS (SECOND LIEN) RECORDED AT R/F 041880/0136 Recorded Dec 27, 2019
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: LIONBRIDGE TECHNOLOGIES, INC.
Reel/Frame 051431/0427 →
RELEASE OF SECURITY INTEREST IN PATENTS RECORDED AT R/F 034631/0925 Recorded Mar 3, 2017
From: HSBC BANK USA, NATIONAL ASSOCIATION, AS ADMINISTRATIVE AGENT
To: LIONBRIDGE TECHNOLOGIES, INC.
Reel/Frame 041879/0908 →
FIRST LIEN PATENT SECURITY AGREEMENT Recorded Mar 3, 2017
From: LIONBRIDGE TECHNOLOGIES, INC.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 041880/0033 →
SECOND LIEN PATENT SECURITY AGREEMENT Recorded Mar 3, 2017
From: LIONBRIDGE TECHNOLOGIES, INC.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 041880/0136 →
SECURITY INTEREST Recorded Jan 5, 2015
From: LIONBRIDGE TECHNOLOGIES, INC.
To: HSBC BANK USA, NATIONAL ASSOCIATION, AS ADMINISTRATIVE AGENT
Reel/Frame 034631/0925 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2014
From: MARCIANO, JAMES PETER; BLODGETT, DEAN SCOTT
To: LIONBRIDGE TECHNOLOGIES, INC.
Reel/Frame 033407/0582 →
Continuity (2)
Provisional Application 61617341 · Mar 29, 2012
Related Publication 20130262080A1 · Oct 3, 2013