IP Library › Granted Patent US 10,452,787
Granted Patent B2
US 10,452,787 · App. 15/152,619 · Granted Oct 22, 2019

Techniques for automated document translation

Inventors: Stephen Condie (Redmond, WA); Charles Reid (Kirkland, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F17/289G06F17/218G06F17/2247G06F17/241
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,452,787
App. No.
15/152,619
Granted
Oct 22, 2019
Kind
B2
Abstract

Techniques for automated document translation are described. An apparatus may comprise a translatable content component, an intermediate component, and a translation management component. The translatable content component may be generally operative to extract translatable content from an original document, and to construct a translated document based on extracted translated content, the translated document comprising a translation of the original document from a first language to a second language. The intermediate component may be operative to create one or more intermediate documents from extracted translatable content, and to extract translated content from one or more translated intermediate documents. The translation management component operative to transmit the one or more intermediate documents to a translation service for translation from a first language to a second language and to receive one or more translated intermediate documents from the translation service. Other embodiments are described and claimed.

Claims (81)

1. A system comprising:

at least one processor; and

memory coupled to the at least one processor, the memory comprising computer executable instructions that, when executed by the at least one processor, performs a method for providing sequential two-handed touch typing, the method comprising:

extracting translatable content from an original document;

creating a plurality of intermediate documents from the extracted translatable content, wherein the plurality of intermediate documents includes the extracted translatable content;

transmitting the plurality of intermediate documents to a translation service for translation from a first language to a second language;

receiving one or more translated intermediate documents from the translation service;

extracting translated content from the one or more translated intermediate documents; and

constructing a translated document comprising the translated content.

2. The system of claim 1 , wherein extracting the translatable content from the original document further comprises:

identifying one or more paragraphs in the original document;

extracting text from the one or more paragraphs;

generating one or more style identifiers for the extracted text;

identifying one or more runs of text; and

generating one or more annotation identifiers for inline objects in the original document.

3. The system of claim 2 , wherein creating the plurality of intermediate documents from the extracted translatable content comprises:

creating paragraphs tags for each identified paragraph;

identifying a predominant style identifier for each paragraph;

associating each paragraph with its predominant style identifier;

identifying off-style runs in each paragraph;

creating style tags for each off-style run; and

creating annotation tags from the annotation identifiers.

4. The method of claim 3 , wherein one or more of the identified paragraphs are stored as stored data, the stored data comprising a set of runs of text having associated style identifiers, wherein the stored data is usable to reconstruct formatting for the one or more identified paragraphs without storing actual text formatting options for the one or more identified paragraphs.

5. The method of claim 3 , wherein extracting the translated content from the translated intermediate documents comprises:

identifying one or more translated paragraphs in the translated intermediate documents;

extracting translated text from the one or more translated paragraphs; associating the translated text of each translated paragraph with the associated predominant style identifier for the translated paragraph;

identifying translated off-style runs in each translated paragraph;

associating style identifiers with the text of each identified translated off-style run;

identifying annotations in the translated intermediate documents; and

associating annotation identifiers from the identified annotations with their place in the extracted translated text.

6. The method of claim 5 , wherein constructing the translated document based on the extracted translated content comprises:

replacing the text from the one or more paragraphs of the original document with the extracted translated text from the translated paragraphs of the translated document, wherein styles are assigned to the extracted translated text using the associated style identifiers, wherein the inline objects from the original document are placed in the translated document based on the annotation identifiers associated with the extracted translated text.

7. The method of claim 1 , further comprising:

selecting a translation parser from a plurality of translation parsers for the original document based on a document type of the original document;

extracting the translatable content from the original document using the selected translation parser; and

constructing the translated document based on the extracted translated content using the selected translation parser.

8. The method of claim 1 , comprising creating the plurality of intermediate documents from the extracted translatable content to accommodate a defined number of pages for the translation service.

9. A system, comprising:

a computing device;

a translatable content component operative on the computing device to extract translatable content from an original document;

an intermediate component operative on the computing device to create a plurality of intermediate documents from the extracted translatable content, wherein the plurality of intermediate documents includes the extracted translatable content;

a translation management component operative on the computing device to transmit the one or more intermediate documents to a translation service and receive one or more translated intermediate documents from the translation service;

the intermediate component further operative to extract translated content from the plurality of translated intermediate documents; and

the translatable content component further operative to construct a translated document based on the extracted translated content, the translated document comprising a translation of the original document from a first language to a second language.

10. The system of claim 9 , the translatable content component further operative to identify one or more paragraphs in the original document, extract text from the one or more paragraphs, generate one or more style identifiers for the extracted text, identify one or more runs of text and generate one or more annotation identifiers for inline objects in the original document.

11. The system of claim 10 , the intermediate component further operative to create paragraph tags for each identified paragraph, identify a predominant style identifier for each paragraph, associate each paragraph with the predominant style identifier, identify off-style runs in each paragraph, create style tags for each off-style run, and create annotation tags using the annotation identifiers.

12. The system of claim 11 , the intermediate component further operative to identify one or more translated paragraphs in the translated intermediate documents, extract translated text from the one or more translated paragraphs, associate the translated text of each translated paragraph with the associated predominant style identifier for the translated paragraph, identify translated off-style runs in each translated paragraph, associate style identifiers with the text of each identified translated off-style run, identify annotations in the translated intermediate documents, and associate annotation identifiers from the identified annotations with their place in the extracted translated text.

13. The system of claim 12 , wherein the translated document is constructed by replacing the text from the one or more paragraphs of the original document with the extracted translated text from the translated paragraphs of the translated document, wherein styles are assigned to the extracted translated text using the associated style identifiers, and wherein the inline objects from the original document are placed in the translated document based on the annotation identifiers associated with the extracted translated text.

14. The system of claim 9 , comprising:

a selection component operative to select a translation parser from a plurality of translation parsers for the original document based on a document type of the original document;

the translatable content component further operative to extract the translatable content from the original document using the selected translation parser; and

the translatable content component further operative to construct the translated document based on the extracted translated content using the selected translation parser.

15. The system of claim 9 , wherein the plurality of intermediate documents accommodate a defined number of pages for the translation service.

16. The system of claim 9 , wherein the plurality of intermediate documents are hypertext markup language (HTML) formatted.

17. A method, comprising:

extracting translatable content from an original document;

creating a plurality of intermediate documents from the extracted translatable content, wherein the plurality of intermediate documents includes the extracted translatable content;

transmitting the plurality of intermediate documents to a translation service for translation from a first language to a second language;

receiving one or more translated intermediate documents from the translation service;

extracting translated content from the one or more translated intermediate documents; and

constructing a translated document based on the extracted translated content, the translated document comprising a translation of the original document from the first language to the second language.

18. The method of claim 17 , wherein extracting the translatable content from the original document further comprises:

identifying one or more paragraphs in the original document;

extracting text from the one or more paragraphs;

generating one or more style identifiers for the extracted text;

identifying one or more runs of text; and

generating one or more annotation identifiers for inline objects in the original document.

19. The method of claim 18 , wherein creating the plurality of intermediate documents from the extracted translatable content comprises:

creating paragraphs tags for each identified paragraph;

identifying a predominant style identifier for each paragraph;

associating each paragraph with its predominant style identifier;

identifying off-style runs in each paragraph;

creating style tags for each off-style run; and

creating annotation tags from the annotation identifiers.

20. The method of claim 19 , wherein extracting the translated content from the translated intermediate documents comprises:

identifying one or more translated paragraphs in the translated intermediate documents;

extracting translated text from the one or more translated paragraphs; associating the translated text of each translated paragraph with the associated predominant style identifier for the translated paragraph;

identifying translated off-style runs in each translated paragraph;

associating style identifiers with the text of each identified translated off-style run;

identifying annotations in the translated intermediate documents; and

associating annotation identifiers from the identified annotations with their place in the extracted translated text.

Continuity (2)
Continuation 13288147 · Nov 3, 2011
Related Publication 20160328392A1 · Nov 10, 2016
Cited By (1)
US 12,585,893