IP Library Granted Patent US 11,990,136
Granted Patent B2
US 11,990,136 · App. 17/428,276 · Granted May 21, 2024

Speech recognition device, search device, speech recognition method, search method, and program

Inventors: Tetsuo Amakasu (Tokyo, JP); Kaname Kasahara (Tokyo, JP); Takafumi Hikichi (Tokyo, JP); Masayuki Sugizaki (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G10L15/32G06F16/245G06N3/04G10L15/02G10L15/04G10L15/142G10L15/16G10L15/22G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,990,136
App. No.
17/428,276
Granted
May 21, 2024
Kind
B2
Abstract

It is intended to acquire a highly accurate speech recognition result for a subject of a conversation, while inhibiting an increase in the amount of calculation. A speech recognition device ( 10 ) according to the present invention includes a first speech recognition unit ( 11 ) that performs speech recognition processing using a first method on speech data of a conversation made by a plurality of speakers and outputs a speech recognition result for each of respective uttered speech segments of the plurality of speakers, a determination unit ( 13 ) that determines a subject segment based on a result of the speech recognition processing by the first speech recognition unit 11 , and a second speech recognition unit ( 14 ) that performs speech recognition processing using a second method higher in accuracy than the first method on the speech data in the segment determined to be the subject segment and outputs a speech recognition result as a subject text.

Claims (17)

1. A speech recognition device a processor configured to execute operations comprising:

performing first speech recognition processing using a first method on speech data of a conversation made by a plurality of speakers and outputs a speech recognition result for each of respective uttered speech segments of the plurality of speakers;

determining, on the basis of a result of the first speech recognition processing, a subject segment of the conversation, wherein the subject segment represents a segment of the speech data including a part of the conversation with utterances about a subject; and

performing second speech recognition processing using a second method higher in accuracy than the first method on speech data in the segment determined to be the subject segment by the determiner and outputs a speech recognition result as a subject text.

2. The speech recognition device according to claim 1 , wherein the determining further comprises determining a domain of the subject based on a keyword included in the uttered speech or the subject segment and a synonym of the keyword, and the second speech recognition processing uses, as the second method, a speech recognition method that uses a speech recognition model in accordance with the domain of the subject of the determined subject segment or a speech recognition method using a speech recognition model in accordance with an acoustic feature of the uttered speech or with a degree of reliability of the first speech recognition processing.

3. A retrieval device comprising a processor configured to execute operations comprising:

retrieving, on the basis of the retrieval query, a retrieval index in which the subject text output from the second speech recognition processing according to claim 1 is associated with an identifier of the conversation including the uttered speech and outputting the identifier associated with the subject text including the retrieval query or similar to the query.

4. The speech recognition device according to claim 1 , wherein the determining further comprises determining the subject segment based on a structure of a conversation by a plurality of speakers, and wherein the structure of a conversation includes a subject confirmation speech uttered to confirm the subject of the conversation.

5. A speech recognition method, comprising:

performing speech recognition processing using a first method on speech data of a conversation made by a plurality of speakers and outputting a speech recognition result for each of respective uttered speech segments of the plurality of speakers;

determining, on the basis of a result of the speech recognition processing using the first method, a subject segment of the conversation, wherein the subject segment represents, wherein the subject segment represents a segment of the speech data including a part of the conversation with utterances about a subject; and

performing, speech recognition processing using a second method higher in accuracy than the first method on the speech data in the segment determined to be the subject segment and outputting a speech recognition result as a subject text.

6. The speech recognition method according to claim 5 , wherein the first method includes a speech recognition method using a HMM (Hidden Markov Model) method or a HMM-DNN (Deep Neural Network) method, and the second method includes a speech recognition method using a CNN-NIN (Convolutional Neural Network and Network In Network) method.

7. The speech recognition method according to claim 5 , wherein the determining further comprises determining the subject segment based on a structure of a conversation by a plurality of speakers, and wherein the structure of a conversation includes a subject confirmation speech uttered to confirm the subject of the conversation.

8. The speech recognition method according to claim 5 , wherein the first method includes a speech recognition method using a HMM (Hidden Markov Model) method or a HMM-DNN (Deep Neural Network) method, and the second method includes a speech recognition method using a CNN-NIN (Convolutional Neural Network and Network In Network) method.

9. The speech recognition method according to claim 5 , wherein the determining further comprises determining a domain of the subject based on a keyword included in the uttered speech or the subject segment and a synonym of the keyword, and

the second speech recognition processing uses, as the second method, a speech recognition method that uses a speech recognition model in accordance with the domain of the subject of the determined subject segment or a speech recognition method using a speech recognition model in accordance with an acoustic feature of the uttered speech or with a degree of reliability of the first speech recognition processing.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2021
From: AMAKASU, TETSUO; KASAHARA, KANAME; HIKICHI, TAKAFUMI; SUGIZAKI, MASAYUKI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 057072/0811 →
Priority Claims (1)
JP 2019-019476 · Feb 6, 2019 · national
Continuity (1)
Related Publication 20220108699A1 · Apr 7, 2022