IP Library Granted Patent US 11,380,301
Granted Patent B2
US 11,380,301 · App. 16/970,798 · Granted Jul 5, 2022

Learning apparatus, speech recognition rank estimating apparatus, methods thereof, and program

Inventors: Tomohiro Tanaka (Yokosuka, JP); Ryo Masumura (Yokosuka, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G10L15/01G06N3/08G10L15/16G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,380,301
App. No.
16/970,798
Granted
Jul 5, 2022
Kind
B2
Abstract

A learning apparatus comprises a learning part that learns an error correction model by a set of a speech recognition result candidate and a correct text of speech recognition for given audio data, wherein the speech recognition result candidate includes a speech recognition result candidate which is different from the correct text, and the error correction model is a model that receives a word sequence of the speech recognition result candidate as input and outputs an error correction score indicating likelihood of the word sequence of the speech recognition result candidate in consideration of a speech recognition error.

Claims (28)

1. A learning apparatus comprising:

processing circuitry configured to

learn an error correction model by a set of a speech recognition result candidate and a correct text of speech recognition for given audio data,

wherein the speech recognition result candidate includes a speech recognition result candidate which is different from the correct text, and the error correction model is a model that receives a word sequence of the speech recognition result candidate as input and outputs an error correction score indicating likelihood of the word sequence of the speech recognition result candidate in consideration of a speech recognition error, and

the processing circuitry learns the error correction model using a set of distributed representations of speech recognition result candidates and word distributed representation sequences of the correct text.

2. The learning apparatus according to claim 1 ,

wherein the set of the speech recognition result candidate and the correct text used for learning the error correction model is composed of a plurality of speech recognition result candidates and one correct text.

3. A non-transitory computer-readable recording medium that records a program for causing a computer to function as the learning apparatus according to claim 1 or 2 .

4. A speech recognition rank estimating apparatus, using an error correction model learned by a learning apparatus, comprising:

processing circuitry configured to

input word distributed representation sequences for word sequences of speech recognition result candidates and distributed representations of the speech recognition result candidates into the error correction model, and finds error correction scores, which are outputs of the error correction model, for respective word sequences of the speech recognition result candidates; and

rank the speech recognition result candidates using the error correction scores,

wherein the learning apparatus learns an error correction model by a set of a speech recognition result candidate and a correct text of speech recognition for given audio data, and the speech recognition result candidate includes a speech recognition result candidate which is different from the correct text, and the error correction model is a model that receives a word sequence of the speech recognition result candidate as input and outputs an error correction score indicating likelihood of the word sequence of the speech recognition result candidate in consideration of a speech recognition error.

5. The speech recognition rank estimating apparatus according to claim 4 ,

wherein the processing circuitry ranks the speech recognition result candidates using scores calculated by weighting and adding speech recognition scores for respective word sequences of the speech recognition result candidates and the error correction scores.

6. A non-transitory computer-readable recording medium that records a program for causing a computer to function as the speech recognition rank estimating apparatus according to claim 5 .

7. A learning method comprising:

learning an error correction model by a set of a speech recognition result candidate and a correct text of speech recognition for given audio data,

wherein the speech recognition result candidate includes a speech recognition result candidate which is different from the correct text, and the error correction model is a model that receives a word sequence of the speech recognition result candidate as input and outputs an error correction score indicating likelihood of the word sequence of the speech recognition result candidate in consideration of a speech recognition error, and

the learning further includes learning the error correction model using a set of distributed representations of speech recognition result candidates and word distributed representation sequences of the correct text.

8. The learning method according to claim 7 ,

wherein the set of the speech recognition result candidate and the correct text used for learning the error correction model is composed of a plurality of speech recognition result candidates and one correct text.

9. A method comprising:

learning an error correction model by a set of a speech recognition result candidate and a correct text of speech recognition for given audio data,

wherein the speech recognition result candidate includes a speech recognition result candidate which is different from the correct text, and the error correction model is a model that receives a word sequence of the speech recognition result candidate as input and outputs an error correction score indicating likelihood of the word sequence of the speech recognition result candidate in consideration of a speech recognition error,

the method further comprising:

inputting word distributed representation sequences for word sequences of speech recognition result candidates and distributed representations of the speech recognition result candidates into the error correction model, and finding error correction scores, which are outputs of the error correction model, for respective word sequences of the speech recognition result candidates; and

ranking the speech recognition result candidates using the error correction scores.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2020
From: TANAKA, TOMOHIRO; MASUMURA, RYO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 053838/0942 →
Priority Claims (1)
JP JP2018-029076 · Feb 21, 2018 · national
Continuity (1)
Related Publication 20210090552A1 · Mar 25, 2021