IP Library Granted Patent US 12,505,315
Granted Patent B2
US 12,505,315 · App. 17/618,372 · Granted Dec 23, 2025

Information learning apparatus, information processing apparatus, information learning method, information processing method and program

Inventors: Makoto Morishita (Tokyo, JP); Jun Suzuki (Tokyo, JP); Masaaki Nagata (Tokyo, JP)
Assignee: NTT, Inc.
G06F40/58G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,315
App. No.
17/618,372
Granted
Dec 23, 2025
Kind
B2
Abstract

An information learning apparatus includes a memory and a processor configured to perform encoding first data in training data in which the first data related to a first series and second data which is correct data for the first data in a second series are associated with each other; decoding data generated in the encoding to generate third data related to the second series; fourth data related to the first series for data generated in the decoding; and learning, based on an error between the second data and the third data and an error between the first data and the fourth data, parameters used by the encoding, the decoding, and the generating, wherein the generating and the encoding share parameters.

Claims (34)

1 . An information learning apparatus for learning a model of a neural network that generates, from a first series, a second series and reconstructs the first sequence, the information learning apparatus comprising:

a memory and a processor, wherein

the memory stores parameters of the neural network,

the neural network comprises a first model and a second model,

the first model comprises a first self-attention layer, a first attention layer, a first feed-forward layer, and a first output layer,

the neural network performs:

encoding, based on the first self-attention layer and the first feed-forward layer, first data in training data, by further associating the first data of the first series and second data in the second series with each other, and the second data represents correct data of the first data;

decoding, based on the second model, data generated in the encoding to generate third data related to the second series;

reconstructing, based at least on the first self-attention layer, the first feed-forward layer, and the first output layer, the first series for data generated in the decoding as fourth data; and

learning, based on an error between the second data and the third data and an error between the first data and the fourth data, the parameters of the neural network used by the encoding, the decoding, and the reconstructing, and the learning further comprises updating the parameters of the neural network stored in the memory, wherein

the encoding and the reconstructing operations share parameters of the first self-attention layer and the first feed-forward layer of the neural network, as stored in the memory.

2 . The information learning apparatus according to claim 1 , wherein the neural network further performs:

encoding fifth data in the training data in which the fifth data related to the second series and sixth data which is correct data for the fifth data in the first series are associated with each other, wherein

the reconstructing further comprises decoding data generated in the encoding by the encoding of the fifth data to generate seventh data related to the first series,

the decoding further comprises generating eighth data related to the second series for data generated in the decoding by the generating,

the learning further comprises learning, based on an error between the sixth data and the seventh data and an error between the fifth data and the eighth data, parameters used by the generating, the encoding of the fifth data, and the decoding, and

the decoding and the encoding of the fifth data share at least a part of the parameters of the neural network.

3 . An information processing apparatus that generates an output sentence for an input sentence based on the parameters learned by the learning of the information learning apparatus according to claim 1 , the

the neural network further performs:

encoding the input sentence based on the parameters of the neural network; and

decoding data generated in the encoding based on the parameters to generate the output sentence.

4 . An information learning method performed by a computer including a memory and a processor, the method comprising:

encoding, based on a first self-attention layer of a first model of a neural network and a first feed-forward layer of the first model of the neural network, first data in training data, by further associating the first data of a first series and second data in a second series with each other, and the second data represents correct data of the first data;

decoding, based on a second model of the neural network, data generated in the encoding to generate third data related to the second series;

reconstructing, based at least on the first self-attention layer, the first feed-forward layer, and the first output layer, the first series for data generated in the decoding as fourth data; and

learning, based on an error between the second data and the third data and an error between the first data and the fourth data, parameters of the neural network used by the encoding, the decoding, and the reconstructing, wherein the learning further comprises updating the parameters of the neural network stored in the memory, and

the encoding and the reconstructing operations share parameters of the first self-attention layer and the first feed-forward layer of the neural network, as stored in the memory.

5 . An information processing method of generating an output sentence for an input sentence based on the parameters of the neural network learned by the learning of the information learning apparatus according to claim 1 , the information processing method performed by a computer including a memory and a processor, the method comprising:

encoding the input sentence based on the parameters of the neural network; and

decoding data generated in the encoding based on the parameters of the neural network to generate the output sentence.

6 . A non-transitory computer-readable recording medium having computer-readable instructions stored thereon, which when executed, cause a computer to function as the information learning apparatus according to claim 1 .

7 . A non-transitory computer-readable recording medium having computer-readable instructions stored thereon, which when executed, cause a computer to function as the information processing apparatus according to claim 3 .

8 . The information learning apparatus according to claim 1 , wherein the second series comprises text.

9 . The information learning apparatus according to claim 1 , wherein the neural network performs text-to-speech conversion, document summarization, document parsing, and/or response sentence generation to a given query.

Assignments (2)
CHANGE OF NAME Recorded Oct 22, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 073184/0535 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2021
From: MORISHITA, MAKOTO; SUZUKI, JUN; NAGATA, MASAAKI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 058364/0847 →
Continuity (1)
Related Publication 20220366155A1 · Nov 17, 2022
References Cited (15)
Bojar, Ondřej, et al. “UFAL submissions to the IWSLT 2016 MT track.” Proceedings of the 13th International Conference on Spoken Language Translation. 2016. (Year: 2016). [cited by examiner]
Morishita et al., “Neural Machine Translation Integrating Interactive Reproductive Learning”, In Proceedings of the 25th Annual Meeting of the Association for Natural Language Processing (NLP019), pp. 1387 to 1390, 2019. [cited by applicant]
Tu et al., “Neural Machine Translation With Reconstruction”, In Proceedings of AAAI, pp. 1 to 7, 2016. [cited by applicant]
Lample et al., “Unsupervised Machine Translation Using Monolingual Corpora Only”, In Proceedings of ICLR, pp. 1 to 14, 2018. [cited by applicant]
Tu et al., “Modeling Coverage for Neural Machine Translation”, In Proceedings of ACL, pp. 76 to 85, 2016. [cited by applicant]
Tu et al., “Neural Machine Translation With Reconstruction”, In Proceedings of AAAI, pp. 3097 to 3103, 2017. [cited by applicant]
Vaswani et al., “Attention Is All You Need”, In Proceedings of NIPS, pp. 6000 to 6010, 2017. [cited by applicant]
Sennrich et al., “Neural Machine Translation of Rare Words with Subword Units”, In Proceedings of ACL, pp. 1715 to 1725, 2016. [cited by applicant]
Xia et al., “Dual Learning for Machine Translation”, In Proceedings of NIPS, pp. 820 to 828, 2016. [cited by applicant]
Xia et al., “Model-Level Dual Learning”, In Proceedings of ICML, pp. 5383 to 5392, 2018. [cited by applicant]
Cettolo et al., “WIT3: Web Inventory of Transcribed and Translated Talks”, In Proceedings of EAMT, pp. 261 to 268, 2012. [cited by applicant]
Gehring et al., “Convolutional Sequence to Sequence Learning”, In Proceedings of ICML, pp. 1243 to 1252, 2017. [cited by applicant]
Press et al., “Using the Output Embedding to Improve Language Models”, In Proceedings of EACL, pp. 157 to 163, 2017. [cited by applicant]
Papineni et al., “BLEU: A Method for Automatic Evaluation of Machine Translation”, In Proceedings of ACL, pp. 311 to 318, 2002. [cited by applicant]
Koehn., “Statistical Significance Tests for Machine Translation Evaluation”, In Proceedings of EMNLP, pp. 388 to 395, 2004. [cited by applicant]