IP Library Granted Patent US 12,462,113
Granted Patent B2
US 12,462,113 · App. 17/610,589 · Granted Nov 4, 2025

Embedding an unknown word in an input sequence

Inventors: Makoto Morishita (Tokyo, JP); Jun Suzuki (Tokyo, JP); Sho Takase (Tokyo, JP); Hidetaka Kamigaito (Tokyo, JP); Masaaki Nagata (Tokyo, JP)
G06F40/44G06F40/30G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,113
App. No.
17/610,589
Granted
Nov 4, 2025
Kind
B2
Abstract

An information learning apparatus includes a memory and a processor configured to perform generating, for each of processing units constituting an input sequence included in training data, a third embedded vector based on a first embedded vector for the processing unit and a second embedded vector corresponding to an unknown word; executing a process based on a learning target parameter, with the third embedded vector generated for each of the processing units as an input; and learning, for a processing result by the executing, the parameter based on an error of an output corresponding to the input sequence in the training data.

Claims (15)

1 . An information learning apparatus comprising:

a memory and a processor configured to:

generate, for each of processing units constituting an input sequence included in training data, a third embedded vector by adding a second embedded vector corresponding to an unknown word to a first embedded vector of each processing unit regardless of whether said each processing unit represents the unknown word or not;

execute, using training data, a process based on a learning target parameter of a sequence-to-sequence model to generate an inference result, with the third embedded vector generated for said each processing unit as an input; and

learn the learning target parameter of the sequence-to-sequence model, wherein the learning target parameter is based on an error between the inference result of the execution and a ground truth output corresponding to the input sequence in the training data, wherein a learnt neural network with the learnt learning target parameter converts an input statement into one or more output words, and the input statement comprises the unknown word.

2 . A non-transitory computer-readable recording medium having computer-readable instructions stored thereon, which when executed, cause a computer including a memory and a processor to execute respective operations in the information processing apparatus according to claim 1 .

3 . An information processing apparatus comprising:

a memory and a processor configured to perform

generating, for each of processing units constituting an input sequence, a third embedded vector by adding a second embedded vector corresponding to an unknown word to a first embedded vector for each processing unit regardless of whether said each processing unit represents the unknown word or not; and

executing, using training data, a process based on a learned parameter of a sequence-to-sequence model to generate an inference result, with the third embedded vector generated for said each processing unit as an input, wherein the learned parameter is based on an error between the inference result of a previous execution of the process and a ground truth output corresponding to the input sequence in training data, a learnt neural network with the learnt learning target parameter converts an input statement into one or more output words as the inference result, and the input statement comprises the unknown word.

4 . A non-transitory computer-readable recording medium having computer-readable instructions stored thereon, when executed, cause a computer including a memory and a processor to execute respective operations in the information processing apparatus according to claim 3 .

5 . An information learning method executed by a computer including a memory and a processor the method comprising:

generating, for each of processing units constituting an input sequence included in training data, a third embedded vector by adding a second embedded vector corresponding to an unknown word to a first embedded vector of each processing unit regardless of whether said each processing unit represents the unknown word or not;

executing, using training data, a process based on a learning target parameter of a sequence-to-sequence model to generate an inference result, with the third embedded vector generated for said each processing unit as an input; and

learning the learning target parameter of the sequence-to-sequence model, wherein the learning target parameter is based on an error between the inference result of the execution and a ground truth output corresponding to the input sequence in the training data, wherein a learnt neural network with the learnt learning target parameter converts an input statement into one or more output words, and the input statement comprises the unknown word.

Assignments (2)
CHANGE OF NAME Recorded Oct 22, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 073184/0535 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 11, 2021
From: MORISHITA, MAKOTO; SUZUKI, JUN; TAKASE, SHO; KAMIGAITO, HIDETAKA; NAGATA, MASAAKI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 058088/0396 →
Continuity (1)
Related Publication 20220215182A1 · Jul 7, 2022
References Cited (12)
US 10346721B2 · Albright · 2019 [cited by examiner]
US 10679148B2 · Chen · 2020 [cited by examiner]
US 11604956B2 · Bradbury · 2023 [cited by examiner]
US 20170323203A1 · Matusov · 2017 [cited by examiner]
US 20180121787A1 · Hashimoto et al. · 2018 [cited by applicant]
US 20190130273A1 · Keskar · 2019 [cited by examiner]
Zhang et al. “Subword-augmented Embedding for Cloze Reading Comprehension”. arXiv:1806.09103v1 [cs.CL] Jun. 24, 2018. (Year: 2018). [cited by examiner]
Sennrich, R., Haddow, B., and Birch, A. (2016) “Neural Machine Translation of Rare Words with Subword Units,” in Proceedings of ACL, pp. 1715-1725. [cited by applicant]
Luong, M.-T., Pham, H., and Manning, C. D. (2015) “Effective Approaches to Attention-based Neural Machine Translation,” in Proceedings of EMNLP. [cited by applicant]
Bahdanau, D., Cho, K., and Bengio, Y. (2015) “Neural Machine Translation by Jointly L earning to Align and Translate,” in Proceedings of ICLR. [cited by applicant]
Makoto Morishita, Jun Suzuki, Masaaki Nagata (2018) “Improving Neural Machine Translation by Incorporating Hierarchical Subword Features” The 27th International Conference on Computational Linguistics (COLING). [cited by applicant]
Jun Suzuki, Sho Takase, Hidetaka Kamigaito, Makoto Morishita, Masaaki Nagata (2018) “An Empirical Study of Building a Strong Baseline for Constituency Parsing” The 56th Annual Meeting of the Association for Computationa… [cited by applicant]