IP Library › Granted Patent US 12,493,755
Granted Patent B2
US 12,493,755 · App. 18/017,938 · Granted Dec 9, 2025

Method for training machine translation model for generating pseudo parallel translation data, method for obtaining pseudo parallel translation data, and method for training machine translation model

Inventors: Benjamin Marie (Koganei, JP); Atsushi Fujita (Koganei, JP)
Assignee: National Institute of Information and Communications Technology
G06F40/58G06N3/09
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,493,755
App. No.
18/017,938
Granted
Dec 9, 2025
Kind
B2
Abstract

Provided is a pseudo parallel translation data generation apparatus for generating pseudo parallel translation data for accurately performing machine translation in an adaptation target domain even when there exists no parallel translation data for the adaptation target domain. Using other-domains parallel translation data D 0 (L 1 -L 2 ), other-domains first language data D 0 (L 1 ), other-domains second language D 0 (L 2 ), adaptation target domain first language data D 0 (R 1 ), and adaptation target domain second language data D 0 (R 2 ), the pseudo parallel translation data generation apparatus 100 performs optimization processing for a cross-lingual language model including an input data embedding unit 2 and an XLM processing unit 3 , and performs parameter optimization processing for a pseudo parallel translation data generation NMT model including the input data embedding unit after the optimization processing and a machine translation processing unit 5 . Performing processing using the pseudo parallel translation data generation machine translation model obtained by the parameter optimization processing allows for obtaining pseudo parallel translation data for the adaptation target domain for which no parallel translation data sets exist.

Claims (51)

1 . A training and inferring method for a machine translation model, which is configured to perform training processing by setting parameters and includes an input data embedding portion and a machine translation processing portion, for generating pseudo parallel translation data, which is performed with a processor and a memory that the processor can access, the method comprising:

(i) an initialization step of, using the processor, obtaining optimal parameters for a cross-lingual language model (XLM), which is configured to perform training processing and includes the input data embedding portion and an XLM processing portion, and setting optimal parameters that have been set in the input data embedding portion of the XLM model in which the optimal parameters have been set to initial parameters for parameters of the input data embedding portion, using:

an other-domains parallel translation data set Dsetp (L1-L2) containing a plurality of pieces of parallel translation data comprising other-domains first language data, which is data for a domain other than an adaptation target domain that is a domain for which pseudo parallel translation data is to be generated, and other-domains second language data, which is translated data in a second language of the other-domains first language data,

an other-domains monolingual data set Dsetm (L 1 ) containing a plurality of pieces of first language data for a domain other than the adaptation target domain,

an other-domains monolingual data set Dsetm (L 2 ) containing a plurality of pieces of second language data for a domain other than the adaptation target domain,

an adaptation target domain monolingual data set Dsetm (R 1 ) containing a plurality of pieces of data of the first language of the adaptation target domain, and

an adaptation target domain monolingual data set Dsetm (R 2 ) containing a plurality of pieces of data of the second language of the adaptation target domain, and

wherein obtaining optimal parameters for the XML model includes performing:

(A) training processing with masking processing that sets masked data obtaining by masking a part of sequences of monolingual input data included in the other-domains monolingual data set Dsetm (L 1 ), the other-domains monolingual data set Dsetm (L 2 ), the adaptation target domain monolingual data set Dsetm (R 1 ), and the adaptation target domain monolingual data set Dsetm (R 2 ) as input data for the XLM model, sets the monolingual input data as correct data, and then performs training processing so that a loss between the correct data and data outputted from the XLM model is reduced; and

(B) training processing with supervised data that sets one of the first language data for the domain other than the adaptation target domain and the second language data for the domain other than the adaptation target domain of parallel translation data contained in the other-domains parallel translation data set Dsetp (L1-L2) as input data for the XLM model, sets the other of the first language data for the domain other than the adaptation target domain and the second language data for the domain other than the adaptation target domain of the parallel translation data as correct data, and then performs training processing so that loss between the correct data and data outputted from the XLM model is reduced;

(ii) an optimization step of, using the processor, obtaining optimum parameters for the machine translation model, which includes the input data embedding portion and the machine translation processing portion, for generating pseudo parallel translation data by performing training processing using at least one of:

(1) auto-encoding processing that sets correct data to the same data as input data and then performs training processing for the machine translation model for generating pseudo parallel translation data with the initial parameters set,

(2) zero-shot round-trip machine translation processing that inputs again output data from the machine translation model for generating pseudo parallel translation data for the input data into the machine translation model for generating pseudo parallel translation data, and then performs training processing for the machine translation model for generating pseudo parallel translation data so that the output from the machine translation model for generating pseudo parallel translation data becomes the same data as the input data, for the machine translation model for generating pseudo parallel translation data with the initial parameters set, and

(3) supervised machine translation processing that sets one of first language data and second language data included in an other-domains translation data set Dsetp (L1-L2) to an input into the machine translation model for generating pseudo parallel translation data and sets the other to correct data, and then performs training processing for the machine translation model for generating pseudo parallel translation data with the initial parameters set;

(iii) a first machine translation step of, using the processor, setting the machine translation model, which has been set with the optimum parameters, by the control signal so that the machine translation model outputs the adaptation target domain second language data, and then performing machine translation processing on first language data obtained from the other-domains parallel translation data set Dsetp (L1-L2) using the machine translation model for generating pseudo parallel translation data to obtain pseudo parallel translation data for the adaptation target domain second language, which is a machine translation processing resultant data for the other-domains first language data; and

(iv) a second machine translation step of, using the processor, setting the machine translation model, which has been set with the optimum parameters, by the control signal so that the machine translation model outputs the adaptation target domain first language data, and then performing machine translation processing on second language data obtained from the other-domains parallel translation data set Dsetp (L1-L2) using the machine translation model for generating pseudo parallel translation data to obtain pseudo parallel translation data for the adaptation target domain in the first language, which is machine translation processing resultant data for the other-domains second language data;

wherein the machine translation model is configured to output data of a specified type in accordance with a control signal and is set so that one of (1) the other-domains first language data, (2) the other-domains second language data, (3) adaptation target domain first language data which is the first language data of the adaptation target domain, and (4) adaptation target domain second language data which is the second language data of the adaptation target domain, as specified by the control signal, is outputted.

2 . A pseudo parallel translation data obtaining method for obtaining pseudo parallel translation data for an adaptation target domain using the machine translation model for generating pseudo parallel translation data obtained by training processing for the machine translation model for generating pseudo parallel translation data according to claim 1 , the pseudo parallel translation data obtaining method, which is performed by the processor and the memory that the processor can access, comprising:

the first machine translation step of, using the processor, performing machine translation processing on first language data obtained from the other-domains parallel translation data set Dsetp (L1-L2) using the machine translation model for generating pseudo parallel translation data with a type of an output set to the adaptation target domain second language to obtain pseudo parallel translation data for the adaptation target domain second language, which is a machine translation processing resultant data for the other-domains first language data;

the second machine translation step of, using the processor, performing machine translation processing on second language data obtained from the other-domains parallel translation data set Dsetp (L1-L2) using the machine translation model for generating pseudo parallel translation data with a type of an output set to the adaptation target domain first language to obtain pseudo parallel translation data for the adaptation target domain in the first language, which is machine translation processing resultant data for the other-domains second language data; and

a pseudo parallel translation data obtaining step of, using the processor, pairing second language pseudo translated data of the adaptation target domain obtained in the first machine translation step with first language pseudo translated data of the adaptation target domain obtained in the second machine translation step to obtain pseudo parallel translation data of the adaptation target domain.

3 . A pseudo parallel translation data obtaining method for obtaining pseudo parallel translation data for an adaptation target domain using the machine translation model for generating pseudo parallel translation data obtained by training processing for the machine translation model for generating pseudo parallel translation data according to claim 1 , the pseudo parallel translation data obtaining method, which is performed with the processor and the memory that the processor can access, comprising:

a monolingual language data machine translation step of, using the processor, setting the machine translation model by the control signal so that the machine translation model outputs the adaptation target domain first language data or the adaptation target domain second language data, and then performing machine translation processing on first language data obtained from the adaptation target domain monolingual language data set Dsetm (R 1 ) or second language data obtained from the adaptation target domain monolingual language data set Dsetm (R 2 ) using the machine translation model for generating pseudo parallel translation data to obtain pseudo parallel translation data for the adaptation target domain second language, which is a machine translation processing resultant data for the adaptation target domain first language data or pseudo parallel translation data for the adaptation target domain first language data, which is a machine translation processing resultant data for the adaptation target domain second language data; and

a pseudo parallel translation data obtaining step of, using the processor, pairing first language data of the adaptation target domain that has been set to the input into the machine translation model for generating pseudo parallel translation data with the adaptation target domain second language pseudo translated data obtained in the monolingual language data machine translation step, or pairing second language data of the adaptation target domain that has been set to the input into the machine translation model for generating pseudo parallel translation data with the adaptation target domain first language pseudo translated data obtained in the monolingual language data machine translation step, thereby obtaining pseudo parallel translation data of the adaptation target domain.

4 . The pseudo parallel translation data obtaining method according to claim 2 , further comprising:

a filter processing step of, using the processor, obtaining a confidence degree indicating reliability of a result of machine translation processing for each sentence pair of the adaptation target domain pseudo parallel translation data obtained in the pseudo parallel translation data obtaining step, selecting only pseudo parallel translation data including sentence pairs whose confidence degrees are greater than or equal to a predetermined value, and outputting the selected data.

5 . A method for training a machine translation model, which is configured to perform training processing by setting parameters, and obtaining data in a second language by performing machine translation on data in a first language for an adaptation target domain, which is performed with a processor and a memory that the processor can access, the method comprising:

a machine translation model training step of, using the processor, training the machine translation model using:

pseudo translation data of the adaptation target domain obtained by the pseudo parallel translation data obtaining method according to claim 2 , and an other-domains parallel translation data set Dsetp (L1-L2) containing a plurality of pieces of translation data including other-domains first language data that is data in the first language for a domain other than the adaptation target domain, and other-domains second language data that is translated data in the second language for the other-domains first language data,

thereby obtaining optimal parameters for the machine translation model, and setting the optimal parameters to the machine translation model to obtain a trained machine translation model,

wherein training the machine translation model includes performing:

(A) processing that sets the other-domains first language data of the parallel translation data contained in the other-domains parallel translation data set Dsetp (L1-L2) as input data for the machine translation model, sets the other-domains second language data of the parallel translation data contained in the other-domains parallel translation data set Dsetp (L1-L2) as correct data, and then performs training processing so that loss between the correct data and data outputted from the machine translation model is reduced; and

(B) processing that sets the adaptation target domain first language data contained in the pseudo translation data as input data for the machine translation model, sets the adaptation target domain second data language data contained in the pseudo translation data as correct data, and then performs training processing so that loss between the correct data and data outputted from the machine translation model is reduced.

6 . A machine translation apparatus that performs machine translation processing using a trained machine translation model obtained by a training and inferring method for the machine translation model for generating pseudo parallel translation data, which is performed with a processor and a memory that the processor can access, the training and inferring method comprising:

(i) an initialization step of, using the processor, obtaining optimal parameters for a cross-lingual language model (XLM), which is configured to perform train processing and includes the input data embedding portion and an XLM processing portion, and setting optimal parameters that have been set in the input data embedding portion of the XLM model in which the optimal parameters have been set to initial parameters for parameters of the input data embedding portion, using:

an other-domains parallel translation data set Dsetp (L1-L2) containing a plurality of pieces of parallel translation data including other-domains first language data, which is data for a domain other than an adaptation target domain that is a domain for which pseudo parallel translation data is to be generated, and other-domains second language data, which is translated data in a second language of the other-domains first language data,

an other-domains monolingual data set Dsetm (L 1 ) containing a plurality of pieces of first language data for a domain other than the adaptation target domain,

an other-domains monolingual data set Dsetm (L 2 ) containing a plurality of pieces of second language data for a domain other than the adaptation target domain,

an adaptation target domain monolingual data set Dsetm (R 1 ) containing a plurality of pieces of data of the first language of the adaptation target domain, and

an adaptation target domain monolingual data set Dsetm (R 2 ) containing a plurality of pieces of data of the second language of the adaptation target domain,

wherein obtaining optimal parameters for the XML model includes performing:

(A) training processing with masking processing that sets masked data obtaining by masking a part of sequences of monolingual input data included in the other-domains monolingual data set Dsetm (L 1 ), the other-domains monolingual data set Dsetm (L 2 ), the adaptation target domain monolingual data set Dsetm (R 1 ), and the adaptation target domain monolingual data set Dsetm (R 2 ) as input data for the XLM model, sets the monolingual input data as correct data, and then performs training processing so that loss between the correct data and data outputted from the XLM model is reduced; and

B) training processing with supervised data that sets one of the first language data for the domain other than the adaptation target domain and the second language data for the domain other than the adaptation target domain of parallel translation data contained in the other-domains parallel translation data set Dsetp (L1-L2) as input data for the XLM model, sets the other of the first language data for the domain other than the adaptation target domain and the second language data for the domain other than the adaptation target domain of the parallel translation data as correct data, and then performs training processing so that loss between the correct data and data outputted from the XLM model is reduced;

(ii) an optimization step of, using the processor, obtaining optimum parameters for the machine translation model which includes the input data embedding portion and the machine translation processing portion, for generating pseudo parallel translation data by performing training processing using at least one of:

(1) auto-encoding processing that sets correct data to the same data as input data and then performs training processing for the machine translation model for generating pseudo parallel translation data with the initial parameters set,

(2) zero-shot round-trip machine translation processing that inputs again output data from the machine translation model for generating pseudo parallel translation data for the input data into the machine translation model for generating pseudo parallel translation data, and then performs training processing for the machine translation model for generating pseudo parallel translation data so that the output from the machine translation model for generating pseudo parallel translation data becomes the same data as the input data, for the machine translation model for generating pseudo parallel translation data with the initial parameters set, and

(3) supervised machine translation processing that sets one of first language data and second language data included in an other-domains translation data set Dsetp (L1-L2) to an input into the machine translation model for generating pseudo parallel translation data and sets the other to correct data, and then performs training processing for the machine translation model for generating pseudo parallel translation data with the initial parameters set;

(iii) a first machine translation step of, using the processor, setting the machine translation model, which has been set with the optimum parameters, by the control signal so that the machine translation model outputs the adaptation target domain second language data, and then performing machine translation processing on first language data obtained from the other-domains parallel translation data set Dsetp (L1-L2) using the machine translation model for generating pseudo parallel translation data to obtain pseudo parallel translation data for the adaptation target domain second language, which is a machine translation processing resultant data for the other-domains first language data; and

(iv) a second machine translation step of, using the processor, setting the machine translation model, which has been set with the optimum parameters, by the control signal so that the machine translation model outputs the adaptation target domain first language data, and then performing machine translation processing on second language data obtained from the other-domains parallel translation data set Dsetp (L1-L2) using the machine translation model for generating pseudo parallel translation data to obtain pseudo parallel translation data for the adaptation target domain in the first language, which is machine translation processing resultant data for the other-domains second language data;

wherein the machine translation model is configured to output data of a specified type in accordance with a control signal and is set so that one of ( 1 ) the other-domains first language data, ( 2 ) the other-domains second language data, ( 3 ) adaptation target domain first language data which is the first language data of the adaptation target domain, and ( 4 ) adaptation target domain second language data which is the second language data of the adaptation target domain, as specified by the control signal, is outputted.

7 . A machine translation apparatus that performs machine translation processing using a trained machine translation model obtained by the method for training a machine translation model according to claim 5 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2023
From: MARIE, BENJAMIN; FUJITA, ATSUSHI
To: NATIONAL INSTITUTE OF INFORMATION AND COMMUNICATIONS TECHNOLOGY
Reel/Frame 062482/0924 →
Priority Claims (1)
JP 2020-137323 · Aug 17, 2020 · national
Continuity (1)
Related Publication 20230274102A1 · Aug 31, 2023
References Cited (9)
US 20140067361A1 · Nikoulina · 2014 [cited by examiner]
US 20210027026A1 · Imamura · 2021 [cited by examiner]
JP 2018116324A · 2018 [cited by applicant]
JP 2020112915A · 2020 [cited by applicant]
Official Communication issued in International Patent Application No. PCT/JP2021/029060, mailed on Oct. 12, 2021. [cited by applicant]
Artetxe et al., “Unsupersized Statistical Machine Translation”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Oct. 31-Nov. 4, 2018, pp. 3632-3642. [cited by applicant]
Conneau et al., “Cross-lingual Language Model Pretraining”, 33rd Conference on Neural Information Processing Systems, 2019, 11 pages. [cited by applicant]
Morita et al., “A Study of Iterative Unsupervised Adaptation of Bidirectional Neural Machine Translation”, The Association for Natural Language Processing, Mar. 2019, pp. 1451-1454. [cited by applicant]
Vaswani et al., “Attention Is All You Need”, In Proceedings of the 30th Neural Information processing Systems Conference (NeurIPS), 2017, pp. 1-11. [cited by applicant]