IP Library › Granted Patent US 12,248,758
Granted Patent B2
US 12,248,758 · App. 17/619,282 · Granted Mar 11, 2025

Generation device and normalization model

Inventors: Toshimitsu Nakamura (Chiyoda-ku, JP); Noritaka Okamoto (Chiyoda-ku, JP); Wataru Uchida (Chiyoda-ku, JP); Yoshinori Isoda (Chiyoda-ku, JP)
Assignee: NTT DOCOMO, INC.
G06F40/47G06F40/279G06F40/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,758
App. No.
17/619,282
Granted
Mar 11, 2025
Kind
B2
Abstract

A generation device is a device that generates a translated sentence in a second language from an input sentence in a first language to be translated. The generation device includes: an acquisition unit that acquires the input sentence; a normalization unit that converts the input sentence into a normalized sentence that is a grammatically correct sentence in the first language; and a first translation unit that generates the translated sentence by translating the normalized sentence into the second language using first parallel translation data that is parallel translation data between the first language and the second language. The normalization unit generates the normalized sentence by using second parallel translation data that is parallel translation data between a third language and the first language. A data amount of the second parallel translation data is larger than a data amount of the first parallel translation data.

Claims (27)

1. A generation device that generates a translated sentence in a second language different from a first language from an input sentence in the first language to be translated, the generation device comprising:

processing circuitry configured to implement

an acquisition unit configured to acquire the input sentence obtained by converting a result of voice recognition of voices issued by a user into a text;

a normalization unit including a normalization model configured to receive the input sentence and to output a normalized sentence that is a grammatically correct sentence in the first language;

a first translation unit configured to generate the translated sentence by translating the normalized sentence into the second language using first parallel translation data that is parallel translation data between the first language and the second language; and

a generation unit configured to generate learning data,

wherein the generation unit includes:

a second translation unit configured to generate a translated sentence for learning by translating an original sentence for learning in the first language into a third language, which is different from the first language and the second language, using second parallel translation data that is parallel translation data between the third language and the first language; and

a third translation unit configured to generate a normalized sentence for learning that is a grammatically correct sentence in the first language by translating the translated sentence for learning into the first language using the second parallel translation data, and configured to output a combination of the normalized sentence for learning and the original sentence for learning as the learning data,

wherein the normalization model is a machine-translation model generated by executing machine-learning using the learning data, and

wherein a data amount of the second parallel translation data is larger than a data amount of the first parallel translation data.

2. The generation device according to claim 1 ,

wherein each of the second translation unit and the third translation unit is a machine-translation model generated by executing machine-learning using the second parallel translation data.

3. The generation device according to claim 1 ,

wherein the normalization unit further includes a detection unit configured to detect an error expression included in the normalized sentence,

wherein the detection unit adds designation of an exclusion expression based on the error expression to the input sentence and outputs the input sentence to which the designation of the exclusion expression is added to the normalization model, and

wherein the normalization model receives the input sentence to which the designation of the exclusion expression is added and outputs the normalized sentence not including the exclusion expression.

4. The generation device according to claim 2 ,

wherein the normalization unit further includes a detection unit configured to detect an error expression included in the normalized sentence,

wherein the detection unit adds designation of an exclusion expression based on the error expression to the input sentence and outputs the input sentence to which the designation of the exclusion expression is added to the normalization model, and

wherein the normalization model receives the input sentence to which the designation of the exclusion expression is added and outputs the normalized sentence not including the exclusion expression.

5. The generation device according to claim 3 ,

wherein the normalization model outputs a likelihood of each word constituting the normalized sentence together with the normalized sentence, and

wherein the detection unit detects the error expression based on the likelihood of each word.

6. The generation device according to claim 4 ,

wherein the normalization model outputs a likelihood of each word constituting the normalized sentence together with the normalized sentence, and

wherein the detection unit detects the error expression based on the likelihood of each word.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2021
From: NAKAMURA, TOSHIMITSU; OKAMOTO, NORITAKA; UCHIDA, WATARU; ISODA, YOSHINORI
To: NTT DOCOMO, INC.
Reel/Frame 058392/0506 →
Priority Claims (1)
JP 2019-111803 · Jun 17, 2019 · national
Continuity (1)
Related Publication 20220245363A1 · Aug 4, 2022
References Cited (10)
US 20130262077A1 · Fuji · 2013 [cited by examiner]
US 20140303960A1 · Orsini · 2014 [cited by examiner]
US 20180052829A1 · Lee · 2018 [cited by examiner]
US 20190236147A1 · Lee · 2019 [cited by examiner]
JP 2005108184A · 2005 [cited by applicant]
JP 201582204A · 2015 [cited by applicant]
International Preliminary Report on Patentability issued Dec. 30, 2021 in PCT/JP2020/016914, (submitting English translation only) 5 pages. [cited by applicant]
International Search Report mailed on Jun. 16, 2020 in PCT/JP2020/016914 filed on Apr. 17, 2020, 3 pages). [cited by applicant]
Mizumoto et al., “A Study on the Use of Multiple Methods in English Writing Error Correction”, IPSJ SIG Technical Reports [CD-ROM], Oct. 15, 2012, vol. 2012-NL-208, No. 8, pp. 1-7, 11 total pages (with partial English t… [cited by applicant]
Japanese Office Action issued Oct. 3, 2023 in Japanese Application 2021-527408, (with unedited computer-generated English translation), 6 pages. [cited by applicant]