IP Library Granted Patent US 12,190,870
Granted Patent B2
US 12,190,870 · App. 16/966,056 · Granted Jan 7, 2025

Learning device, learning method, and learning program

Inventors: Atsunori Ogawa (Musashino, JP); Marc Delcroix (Musashino, JP); Shigeki Karita (Musashino, JP); Tomohiro Nakatani (Musashino, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G10L15/1815G06F18/2113G06F18/214G06F18/217G06N3/08G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,870
App. No.
16/966,056
Granted
Jan 7, 2025
Kind
B2
Abstract

A learning device includes a memory, and processing circuitry coupled to the memory and configured to receive an input of a plurality of series for learning having known accuracy, and learn a model represented by a neural network, the model being capable of determining accuracy levels of two series when given feature amounts of the two series among the plurality of series.

Claims (31)

1. A learning device comprising:

a memory; and

processing circuitry coupled to the memory and configured to:

receive an input of a plurality of series for learning having known accuracy, and

learn a model represented by a neural network, the model determining accuracy levels of two series of the plurality of series based on a one-to-one comparison between the two series when given feature amounts of the two series among the plurality of series, one of the two series of the plurality of series being a series estimated to have a highest accuracy of the plurality of series after repeated one-to-one comparisons among the plurality of series.

2. The learning device according to claim 1 , wherein the model converts two hypotheses into hidden state vectors using a recurrent neural network, and outputs a first posterior probability indicating that sequence of the accuracy levels of the two series is correct and a second posterior probability indicating that the sequence of the accuracy levels of the two series is incorrect on a basis of the hidden state vectors using a neural network.

3. The learning device according to claim 1 , wherein the processing circuitry is further configured to assign a correct label to the model to be learned when a series having higher accuracy among the two series is placed in a higher rank than another series, and assign an error label to the model to be learned when the series having the higher accuracy among the two series is placed in a lower rank than an other series.

4. The learning device according to claim 3 , wherein the processing circuitry is further configured to:

receive an input of N-best hypotheses for learning having known speech recognition accuracy, and

assign a correct label to a model to be learned when a hypothesis having higher speech recognition accuracy among two hypotheses of the N-best hypotheses is placed in a higher rank than another hypothesis, and assign the error label to the model to be learned when the hypothesis having the higher speech recognition accuracy among the two hypotheses is placed in a lower rank than an other hypotheses.

5. The learning device according to claim 4 , wherein one hypothesis among the two hypotheses is an Oracle hypothesis having highest speech recognition accuracy.

6. The learning device according to claim 5 , wherein another hypothesis among the two hypotheses includes at least any of a first hypothesis having second-highest speech recognition accuracy after the Oracle hypothesis, a second hypothesis having a highest speech recognition score in the N-best hypotheses, a third hypothesis having lowest speech recognition accuracy, and a fourth hypothesis having a lowest speech recognition score in the N-best hypotheses.

7. The learning device according to claim 6 , wherein the other hypothesis among the two hypotheses includes a prescribed number of hypotheses and the first to the fourth hypotheses, the prescribed number of hypotheses being extracted according to a prescribed rule from hypotheses excluding the Oracle hypothesis, the first hypothesis, the second hypothesis, the third hypothesis, and the fourth hypothesis from the N-best hypotheses.

8. A learning method that is performed by a learning device, the learning method comprising:

receiving an input of a plurality of series for learning having known accuracy; and

learning a model represented by a neural network, a model determining accuracy levels of two series of the plurality of series based on a one-to-one comparison between the two series when given feature amounts of the two series among the plurality of series, one of the two series of the plurality of series being a series estimated to have a highest accuracy of the plurality of series after repeated one-to-one comparisons among the plurality of series.

9. A non-transitory computer-readable recording medium storing therein a learning program that causes a computer to execute a process comprising:

receiving an input of a plurality of series for learning having known accuracy; and

learning a model represented by a neural network, the model determining accuracy levels of two series of the plurality of series based on a one-to-one comparison between two series when given feature amounts of the two series among the plurality of series, one of the two series of the plurality of series being a series estimated to have a highest accuracy of the plurality of series after repeated one-to-one comparisons among the plurality of series.

10. The learning device according to claim 1 , wherein the model is a recurrent neural network.

11. The learning device according to claim 10 , wherein the processing circuitry

maintains a series of the two of series having a highest accuracy of the two series and selects another series from a plurality of series to form a pair of another of series, and

determines an accuracy level of each series in a pair of an other two.

12. The method according to claim 8 , wherein the model is a recurrent neural network.

13. The method according to claim 12 , further comprising:

maintaining a series of the two of series having a highest accuracy of the two series and selects another series from a plurality of series to form a pair of another of series; and

determining an accuracy level of each series in the pair of an other two series.

14. The non-transitory computer-readable medium according to claim 9 , wherein the model is a recurrent neural network.

15. The non-transitory computer-readable medium according to claim 14 , further comprising:

maintaining a series of the two of series having a highest accuracy of the two series and selects another series from a plurality of series to form a pair of another of series; and

determining an accuracy level of each series in the pair of an other two series.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2021
From: OGAWA, ATSUNORI; DELCROIX, MARC; KARITA, SHIGEKI; NAKATANI, TOMOHIRO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 055092/0307 →
Priority Claims (1)
JP 2018-017224 · Feb 2, 2018 · national
Continuity (1)
Related Publication 20200365143A1 · Nov 19, 2020
References Cited (52)
US 5255347A · Matsuba · 1993 [cited by examiner]
US 7149687B1 · Gorin · 2006 [cited by examiner]
US 9015093B1 · Commons · 2015 [cited by examiner]
US 10388272B1 · Thomson · 2019 [cited by examiner]
US 10573312B1 · Thomson · 2020 [cited by examiner]
US 10762903B1 · Kahan · 2020 [cited by examiner]
US 10911596B1 · Do · 2021 [cited by examiner]
US 10971153B2 · Thomson · 2021 [cited by examiner]
US 11017778B1 · Thomson · 2021 [cited by examiner]
US 11132988B1 · Steedman Henderson · 2021 [cited by examiner]
US 11145293B2 · Prabhavalkar · 2021 [cited by examiner]
US 11145312B2 · Thomson · 2021 [cited by examiner]
US 11170761B2 · Thomson · 2021 [cited by examiner]
US 11204964B1 · Coope · 2021 [cited by examiner]
US 11295739B2 · Li · 2022 [cited by examiner]
US 11367432B2 · Peyser · 2022 [cited by examiner]
US 11380330B2 · Kahan · 2022 [cited by examiner]
US 11537661B2 · Coope · 2022 [cited by examiner]
US 11551663B1 · Bissell · 2023 [cited by examiner]
US 11594221B2 · Thomson · 2023 [cited by examiner]
US 11646019B2 · Prabhavalkar · 2023 [cited by examiner]
US 11837222B2 · Ogawa · 2023 [cited by examiner]
US 20040186714A1 · Baker · 2004 [cited by examiner]
US 20150324686A1 · Julian · 2015 [cited by examiner]
US 20160034446A1 · Yamamoto · 2016 [cited by examiner]
US 20180330718A1 · Hori · 2018 [cited by examiner]
US 20200027444A1 · Prabhavalkar · 2020 [cited by examiner]
US 20200066271A1 · Li · 2020 [cited by examiner]
US 20200175961A1 · Thomson · 2020 [cited by examiner]
US 20200349922A1 · Peyser · 2020 [cited by examiner]
US 20200365143A1 · Ogawa · 2020 [cited by examiner]
US 20210027785A1 · Kahan · 2021 [cited by examiner]
US 20210035564A1 · Ogawa · 2021 [cited by examiner]
US 20220005465A1 · Prabhavalkar · 2022 [cited by examiner]
US 20220093093A1 · Krishnan · 2022 [cited by examiner]
US 20220093094A1 · Krishnan · 2022 [cited by examiner]
US 20220093101A1 · Krishnan · 2022 [cited by examiner]
US 20220107979A1 · Coope · 2022 [cited by examiner]
US 20220122587A1 · Thomson · 2022 [cited by examiner]
US 20220199084A1 · Li · 2022 [cited by examiner]
US 20220262356A1 · Ogawa · 2022 [cited by examiner]
JP 7280382B2 · 2023 [cited by examiner]
WO WO2020117504A1 · 2020 [cited by examiner]
WO WO2020226777A1 · 2020 [cited by examiner]
WO WO2022060970A1 · 2022 [cited by examiner]
Shimaoka, S., et al., “Learning of the word vector in the automatic encoder,” Proceedings of the 19th annual meeting of the Association for Natural Language Processing, 10 pages, Mar. 2013. [cited by applicant]
International Search Report and Written Opinion mailed on Apr. 9, 2019 for PCT/JP2019/003734 filed on Feb. 1, 2019, 9 pages including English Translation of the International Search Report. [cited by applicant]
Iwatate, M., et al., “Japanese Dependency Parsing Using a Tournament Model,” vol. 15, No. 5, Oct. 2008, pp. 169-185. [cited by applicant]
Oba, T., et al., “Round-Robin Duel Discriminative Language Models,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 20, No. 4, May 2012, pp. 1244-1255. [cited by applicant]
Ogawa, A. and Hori, T., “Error detection and accuracy estimation in automatic speech recognition using deep bidirectional recurrent neural networks,” Speech Communication, vol. 89, May 2017, pp. 70-83. [cited by applicant]
Ogawa, A., “Rescoring of N-Best speech recognition hypotheses using an encoder-classifier model that performs one-to-one hypothesis comparison,” Lecture proceedings of 2018 spring research conference of the Acoustical S… [cited by applicant]
Shimaoka, S., et al., “Learning of the word vector in the automatic encoder,” Proceedings of the 19th annual meeting of the Association for Natural Language Processing, 10 pages. [cited by applicant]