IP Library Patent Application 13818137
Patent Application
App. No. 13/818,137

STATISTICAL MACHINE TRANSLATION METHOD USING DEPENDENCY FOREST

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
13/818,137
Abstract

The present invention relates to the use of a plurality of dependency trees in tree-based statistical machine translation and proposes a dependency forest to effectively process the plurality of dependency trees. The present invention can improve a translation capability by generating a translation rule and a dependency language model by using the dependency forest and applying the generated translation rule and dependency language model when a source language text is converted to a target language text.

Claims (32)

1 . A method of generating a translation rule comprising:

extracting a translation rule by using a dependency forest generated by combining a plurality of dependency trees.

2 . The method of claim 1 , wherein respective nodes of the dependency forest are connected by a hyperedge, and the hyperedge packs all dependants having a common head.

3 . The method of claim 2 , wherein the nodes are distinguished by a span.

4 . The method of claim 1 , wherein the dependency forest is aligned with a source sentence string and a translation rule is extracted from string-to-forest aligned corpus.

5 . The method of claim 2 , wherein a plurality of best well-formed structures for each node is maintained by searching for a well-formed structure for the node.

6 . The method of claim 5 , wherein the plurality of best well-formed structures is obtained by connecting fixed structures of dependants of the node.

7 . The system of claim 6 , wherein a translation rule is extracted when a dependency structure within the well-formed structure corresponds to a word alignment.

8 . A method of generating a translation rule comprising:

performing a dependency analysis for a bilingual corpus;

generating a dependency tree by the dependency analysis and generating a dependency forest by combining a plurality of dependency trees;

searching for a plurality of well-formed structures for each node within the dependency forest; and

extracting a translation rule when dependency structures within the plurality of well-formed structures correspond to a word alignment.

9 . The method of claim 8 , wherein the plurality of well-formed structures is k-best fixed and floating structures and obtained by manipulating a fixed structure of dependants of the node.

10 . A statistical machine translation method comprising:

translating a source language by using a translation rule and a dependency language model generated from a dependency forest generated by combining a plurality of dependency trees.

11 . The statistical machine translation method of claim 10 , wherein respective nodes of the dependency forest are connected by a hyperedge, and the hyperedge packs all dependants having a common head.

12 . The statistical machine translation method of 11, wherein all heads and dependants thereof are collected by listing all hyperedges of the dependency forest, and the dependency language model is generated from the collected information.

13 . An apparatus for generating a translation rule comprising:

a means that generates a dependency tree by performing a dependency analysis for a corpus of a pair of languages and generates a dependency forest by combining a plurality of dependency trees;

a means that searches for a plurality of well-formed structures for each node within the dependency forest; and

a means that extracts a translation rule when dependency structures within the plurality of well-formed structures correspond to a word alignment.

14 . The apparatus of claim 13 , wherein the plurality of well-formed structures is k-best fixed and floating structures, and obtained by controlling a fixed structure of dependants of the node.

15 . A statistical machine translation apparatus comprising:

a dependency parser that generates a dependency tree by performing a dependency analysis for a source sentence and a target sentence of a corpus of a pair of languages and generates a dependency forest for the source sentence and the target sentence by combining a plurality of dependency trees;

a translation rule extractor that extracts a translation rule by using the dependency forest;

a language model trainer that generates a dependency language model by using the dependency forest of the target sentence; and

a decoder that converts a source sentence text to a target sentence text by applying the translation rule and the dependency language model.

16 . The statistical machine translation apparatus of claim 15 , wherein the dependency parser generates the dependency forest by connecting nodes forming a plurality of dependency trees by a hyperedge, and the hyperedge packs all dependants having a common head.

17 . The statistical machine translation apparatus of claim 16 , wherein the translation rule extractor searches for a plurality of well-formed structures for each node within the dependency forest and extracts a translation rule when dependency structures within the plurality of well-formed structures correspond to a word alignment.

18 . The statistical machine translation apparatus of claim 16 , wherein the language model trainer collects all heads and dependants thereof by listing all hyperedges of the dependency forest and generates the dependency language model from the collected information.

19 . A computer-readable recording medium recording a program for executing a process of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2013
From: HWANG, YOUNG SOOK; KIM, SANG BUM; LIN, SHOUXUN; TU, ZHAOPENG; LIU, YANG; LIU, QUN; YIN, CHANG HAO
To: SK PLANET CO., LTD.
Reel/Frame 029847/0479 →