IP Library › Granted Patent US 12,373,655
Granted Patent B2
US 12,373,655 · App. 18/122,316 · Granted Jul 29, 2025

Machine translation method and apparatus, device and storage medium

Inventors: Ruiqing Zhang (Beijing, CN); Hui Liu (Beijing, CN); Zhongjun He (Beijing, CN); Zhi Li (Beijing, CN); Hua Wu (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
G06F40/47G06F40/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,655
App. No.
18/122,316
Granted
Jul 29, 2025
Kind
B2
Abstract

A machine translation method includes: obtaining first target language text by performing first translation on source language text using an initial NMT model; identifying an untranslated part in the source language text based on the source language text and the first target language text; obtaining an adjusted NMT model by increasing an attention weight corresponding to the untranslated part in the initial NMT mode; and obtaining second target language text by performing second translation on the source language text using the adjusted NMT model.

Claims (65)

1. A machine translation method, comprising:

obtaining first target language text by performing first translation on source language text using an initial neural machine translation (NMT) model, wherein the source language text is a first sentence of a first language, and the first target language text is a second sentence of a second language different from the first language, which is obtained by translation of entire first sentence;

identifying an untranslated part in the source language text based on the source language text and the first target language text, comprising: inputting the source language text and the first target language text to a translation missing detection model; outputting by the translation missing detection model identification information being used for identifying whether a text unit in the source language text is untranslated; and identifying the untranslated part in the source language text based on the identification information;

obtaining an adjusted NMT model by increasing an attention weight corresponding to the untranslated part in the initial NMT model; and

obtaining second target language text by performing second translation on the source language text using the adjusted NMT model, wherein the second target language text is a third sentence of the second language,

wherein the translation missing detection model is obtained based on training data, the training data comprises input samples and a training missing tag, and the training data is generated by:

acquiring a first source language sample and a first target language sample translated from the first source language sample;

performing content augmentation on the first target language sample to obtain a second target language sample;

obtaining a second source language sample by reverse translation of the second target language sample;

taking the second source language sample and the first target language sample as the input samples; and

comparing the second source language sample with the first source language sample to determine the training missing tags.

2. The method of claim 1 , wherein the identification information comprises a missing tag corresponding to an untranslated text unit and non-missing tags corresponding to translated text units.

3. The method of claim 1 , wherein the training missing tag corresponds to a text unit included in the second source language sample but not included in the first source language sample.

4. The method of claim 1 , wherein performing content augmentation on the first target language sample to obtain the second target language sample comprises:

inputting the first target language sample to a generative pre-training model, and

outputting, by the generative pre-training model, the second target language sample with more words compared with the first target language sample.

5. The method of claim 1 , wherein performing content augmentation on the first target language sample to obtain the second target language sample comprises:

adding randomly a mask in the first target language sample, and inputting the first target language sample with the mask added into a masked language model; and

outputting, by the masked language model, the second target language sample with the mask being replaced with a text unit.

6. The method of claim 1 , wherein increasing the attention weight corresponding to the untranslated part in the initial NMT model comprises:

determining a maximum attention weight in the initial NMT model;

reducing the maximum attention weight, and determining a difference value of the maximum attention weight before and after the reduction; and

adding the difference value to the attention weight corresponding to the untranslated part.

7. The method of claim 6 , wherein in a case that the untranslated part comprises a plurality of text units, the adding the difference value to the attention weight corresponding to the untranslated part comprises:

determining a to-be-added value corresponding to each of the plural text units based on standard normal distribution and the difference value; and

adding the to-be-added value corresponding to each text unit to the attention weight corresponding to the text unit.

8. An electronic device, comprising:

at least one processor; and

a memory connected with the at least one processor communicatively;

wherein the memory stores instructions executable by the at least one processor to cause the at least one processor to perform a machine translation method comprising:

obtaining first target language text by performing first translation on source language text using an initial neural machine translation (NMT) model, wherein the source language text is a first sentence of a first language, and the first target language text is a second sentence of a second language different from the first language, which is obtained by translation of entire first sentence;

identifying an untranslated part in the source language text based on the source language text and the first target language text, comprising: inputting the source language text and the first target language text to a translation missing detection model; outputting by the translation missing detection model identification information being used for identifying whether a text unit in the source language text is untranslated; and identifying the untranslated part in the source language text based on the identification information;

obtaining an adjusted NMT model by increasing an attention weight corresponding to the untranslated part in the initial NMT model; and

obtaining second target language text by performing second translation on the source language text using the adjusted NMT model, wherein the second target language text is a third sentence of the second language,

wherein the translation missing detection model is obtained based on training data, the training data comprises input samples and a training missing tag, and the training data is generated by:

acquiring a first source language sample and a first target language sample translated from the first source language sample;

performing content augmentation on the first target language sample to obtain a second target language sample;

obtaining a second source language sample by reverse translation of the second target language sample;

taking the second source language sample and the first target language sample as the input samples; and

comparing the second source language sample with the first source language sample to determine the training missing tags.

9. The electronic device of claim 8 , wherein increasing the attention weight corresponding to the untranslated part in the initial NMT model comprises:

determining a maximum attention weight in the initial NMT model;

reducing the maximum attention weight, and determining a difference value of the maximum attention weight before and after the reduction; and

adding the difference value to the attention weight corresponding to the untranslated part.

10. The electronic device of claim 9 , wherein in a case that the untranslated part comprises a plurality of text units, the adding the difference value to the attention weight corresponding to the untranslated part comprises:

determining a to-be-added value corresponding to each of the plural text units based on standard normal distribution and the difference value; and

adding the to-be-added value corresponding to each text unit to the attention weight corresponding to the text unit.

11. A non-transitory computer readable storage medium storing computer instructions for causing a computer to perform a machine translation method comprising:

obtaining first target language text by performing first translation on source language text using an initial neural machine translation (NMT) model, wherein the source language text is a first sentence of a first language, and the first target language text is a second sentence of a second language different from the first language, which is obtained by translation of entire first sentence;

identifying an untranslated part in the source language text based on the source language text and the first target language text, comprising: inputting the source language text and the first target language text to a translation missing detection model; outputting by the translation missing detection model identification information being used for identifying whether a text unit in the source language text is untranslated; and identifying the untranslated part in the source language text based on the identification information;

obtaining an adjusted NMT model by increasing an attention weight corresponding to the untranslated part in the initial NMT model; and

obtaining second target language text by performing second translation on the source language text using the adjusted NMT model, wherein the second target language text is a third sentence of the second language,

wherein the translation missing detection model is obtained based on training data, the training data comprises input samples and a training missing tag, and the training data is generated by:

acquiring a first source language sample and a first target language sample translated from the first source language sample;

performing content augmentation on the first target language sample to obtain a second target language sample;

obtaining a second source language sample by reverse translation of the second target language sample;

taking the second source language sample and the first target language sample as the input samples; and

comparing the second source language sample with the first source language sample to determine the training missing tags.

12. The storage medium of claim 11 , wherein increasing the attention weight corresponding to the untranslated part in the initial NMT model comprises:

determining a maximum attention weight in the initial NMT model;

reducing the maximum attention weight, and determining a difference value of the maximum attention weight before and after the reduction; and

adding the difference value to the attention weight corresponding to the untranslated part.

13. The storage medium of claim 12 , wherein in a case that the untranslated part comprises a plurality of text units, the adding the difference value to the attention weight corresponding to the untranslated part comprises:

determining a to-be-added value corresponding to each of the plural text units based on standard normal distribution and the difference value; and

adding the to-be-added value corresponding to each text unit to the attention weight corresponding to the text unit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2023
From: ZHANG, RUIQING; LIU, HUI; HE, ZHONGJUN; LI, ZHI; WU, HUA
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 063004/0188 →
Priority Claims (1)
CN 202210465501.1 · Apr 26, 2022 · national
Continuity (1)
Related Publication 20230342561A1 · Oct 26, 2023
References Cited (19)
US 20180300317A1 · Bradbury · 2018 [cited by examiner]
US 20210312144A1 · Mizushima · 2021 [cited by examiner]
US 20220245362A1 · Nizar · 2022 [cited by examiner]
US 20230050134A1 · Biswas · 2023 [cited by examiner]
CN 108763222A · 2018 [cited by applicant]
CN 110443346A · 2019 [cited by applicant]
CN 111160036A · 2020 [cited by applicant]
CN 112926345A · 2021 [cited by applicant]
CN 113392657A · 2021 [cited by examiner]
CN 113408306A · 2021 [cited by applicant]
JP 2019096303A · 2019 [cited by applicant]
JP 2022114144A · 2022 [cited by applicant]
WO WO2019019916A1 · 2019 [cited by examiner]
WO WO2019107623A1 · 2019 [cited by examiner]
WO 2022118604A1 · 2022 [cited by applicant]
Lu, Yu et al, “Attention Calibration for Transformer in Neural Machine Translation”, Aug. 6, 2021, Association of Computational Linguistics, Proceedings of the 59th Annual Meeting of the Association for Computational Li… [cited by examiner]
P. Shah and V. Bakrola, “Neural Machine Translation System of Indic Languages—An Attention based Approach,” 2019 Second International Conference on Advanced Computational and Communication Paradigms (ICACCP), Gangtok, I… [cited by examiner]
Y. Lu, J. Zhang, J. Zeng, S. Wu and C. Zong, “Attention Analysis and Calibration for Transformer in Natural Language Generation,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 1927-193… [cited by examiner]
Notice of Reasons for Refusal of Japanese patent application No. 2023001658 issued Apr. 2, 2024, 5 pages. [cited by applicant]
Cited By (1)
US 12,688,362