IP Library Granted Patent US 9,672,820
Granted Patent B2
US 9,672,820 · App. 14/490,722 · Granted Jun 6, 2017

Simultaneous speech processing apparatus and method

Inventors: Satoshi Kamatani (Kanagawa, JP); Akiko Sakamoto (Kanagawa, JP)
Assignee: Kabushiki Kaisha Toshiba
G10L15/18G06F17/277G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,672,820
App. No.
14/490,722
Granted
Jun 6, 2017
Kind
B2
Abstract

According to one embodiment, a simultaneous speech processing apparatus includes an acquisition unit, a speech recognition unit, a detection unit and an output unit. The acquisition unit acquires a speech signal. The speech recognition unit generates a decided character string and at least one candidate character string. The detection unit detects a first character string as a processing piece character string if the first character string included in the decided character string exists commonly in one or more combined character strings on dividing the one or more combined character strings by a boundary indicating a morphological position serving as a start position of a processing piece in natural language processing. The output unit outputs the processing piece character string.

Claims (48)

1. A simultaneous speech processing apparatus, comprising:

a storage to store, as processing piece information, the processing piece character string and time information of a speech signal corresponding to a speech section in which the processing piece character string is uttered, in association with each other; and

a processor programmed to:

acquire a speech signal;

generate a decided character string and at least one candidate character string, the decided character string being a character string corresponding to a speech section in which part of the speech signal having undergone speech recognition processing is converted into a character string, the at least one candidate character string being a character string corresponding to a speech section in which a character string as a conversion result is variable during speech recognition processing in a speech section succeeding the decided character string;

detect a first character string as a processing piece character string if the first character string included in the decided character string exists commonly in one or more combined character strings on dividing the one or more combined character strings by a boundary, the one or more combined character strings being obtained by connecting the decided character string and the at least one candidate character string, the boundary indicating a morphological position serving as a start position of a processing piece in natural language processing;

output the processing piece character string; and

if a first processing piece information as new processing piece information is added in the storage and a second processing piece information exists, connect processing piece character strings included in the second processing piece information and the first processing piece information in time series to generate a reprocessing piece character string, and update the processing piece information stored in the storage with the reprocessing piece character string and time information corresponding to the reprocessing piece character string, the second processing piece information preceding the first processing piece information and corresponding to a speech section in which the second processing piece information is uttered continuously within a time falling within a threshold value.

2. The apparatus according to claim 1 , wherein the processor is further programmed to update a previously acquired processing piece character string if positions of the boundary are different between the previously acquired processing piece character string and a processing piece character sting that a newly acquired processing piece character sting is added to the previously acquired processing piece character string.

3. The apparatus according to claim 1 , wherein the processor is programmed to

acquire time information about a time during which the processing piece character string is uttered, and

determine—with reference to the time information whether or not the second processing piece information exists.

4. The apparatus according to claim 1 , wherein the processor is programmed to

acquire a speech rate which indicates a rate at which a speaker speaks, and

determine—with reference to the speech rate whether or not the second processing piece information exists.

5. The apparatus according to claim 1 , wherein if the natural language processing is machine translation, the processing piece is a translation device suitable for simultaneously translating in parallel the speech signal.

6. The apparatus according to claim 1 , wherein if the natural language processing is speech interaction, the processing piece is a device for simultaneously outputting in parallel the speech signal as a speech interaction task.

7. A simultaneous speech processing method, comprising:

storing, as processing piece information, the processing piece character string and time information of a speech signal corresponding to a speech section in which the processing piece character string is uttered, in association with each other in the storage;

acquiring a speech signal;

generating a decided character string and at least one candidate character string, the decided character string being a character string corresponding to a speech section in which part of the speech signal having undergone speech recognition processing is converted into a character string, the at least one candidate character string being a character string corresponding to a speech section in which a character string as a conversion result is variable during speech recognition processing in a speech section succeeding the decided character string;

detecting a first character string as a processing piece character string if the first character string included in the decided character string exists commonly in one or more combined character strings on dividing the one or more combined character strings by a boundary, the one or more combined character strings being obtained by connecting the decided character string and the at least one candidate character string, the boundary indicating a morphological position serving as a start position of a processing piece in natural language processing;

outputting the processing piece character string; and

connecting, if first processing piece information as new processing piece information is added in the storage and second processing piece information exists, processing piece character strings included in the second processing piece information and the first processing piece information in time series to generate a reprocessing piece character string, and updating the processing piece information stored in the storage with the reprocessing piece character string and time information corresponding to the reprocessing piece character string, the second processing piece information preceding the first processing piece information and corresponding to a speech section in which the second processing piece information is uttered continuously within a time falling within a threshold value.

8. The method according to claim 7 , further comprising updating a previously acquired processing piece character string if a position of the boundary changes in accordance with a relationship between a newly acquired processing piece character string and the previously acquired processing piece character string.

9. The method according to claim 7 , wherein

the generating the decided character string acquires time information about a time during which the processing piece character string is uttered, and

the updating the processing piece information determines with reference to the time information whether or not the second processing piece information exists.

10. The method according to claim 7 , wherein

the generating the decided character string acquires a speech rate which indicates a rate at which a speaker speaks, and

the updating the processing piece information determines with reference to the speech rate whether or not the second processing piece information exists.

11. The method according to claim 7 , wherein if the natural language processing is machine translation, the processing piece is a translation device suitable for simultaneously translating in parallel the speech signal.

12. The method according to claim 7 , wherein if the natural language processing is speech interaction, the processing piece is a device for simultaneously outputting in parallel the speech signal as a speech interaction task.

13. A non-transitory computer readable medium including computer executable instructions, wherein the instructions, when executed by a processor, cause the processor to perform a method comprising:

storing, as processing piece information, the processing piece character string and time information of a speech signal corresponding to a speech section in which the processing piece character string is uttered, in association with each other in the storage;

acquiring a speech signal;

generating a decided character string and at least one candidate character string, the decided character string being a character string corresponding to a speech section in which part of the speech signal having undergone speech recognition processing is converted into a character string, the at least one candidate character string being a character string corresponding to a speech section in which a character string as a conversion result is variable during speech recognition processing in a speech section succeeding the decided character string;

detecting a first character string as a processing piece character string if the first character string included in the decided character string exists commonly in one or more combined character strings on dividing the one or more combined character strings by a boundary, the one or more combined character strings being obtained by connecting the decided character string and the at least one candidate character string, the boundary indicating a morphological position serving as a start position of a processing piece in natural language processing;

outputting the processing piece character string; and

connecting, if first processing piece information as new processing piece information is added in the storage and second processing piece information exists, processing piece character strings included in the second processing piece information and the first processing piece information in time series to generate a reprocessing piece character string, and updating the processing piece information stored in the storage with the reprocessing piece character string and time information corresponding to the reprocessing piece character string, the second processing piece information preceding the first processing piece information and corresponding to a speech section in which the second processing piece information is uttered continuously within a time falling within a threshold value.

14. The medium according to claim 13 , further comprising updating a previously acquired processing piece character string if a position of the boundary changes in accordance with a relationship between a newly acquired processing piece character string and the previously acquired processing piece character string.

15. The medium according to claim 13 , wherein

the generating the decided character string acquires time information about a time during which the processing piece character string is uttered, and

the updating the processing piece information determines with reference to the time information whether or not the second processing piece information exists.

16. The medium according to claim 13 , wherein

the generating the decided character string acquires a speech rate which indicates a rate at which a speaker speaks, and

the updating the processing piece information determines with reference to the speech rate whether or not the second processing piece information exists.

17. The medium according to claim 13 , wherein if the natural language processing is machine translation, the processing piece is a translation device suitable for simultaneously translating in parallel the speech signal.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2014
From: KAMATANI, SATOSHI; SAKAMOTO, AKIKO
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 034028/0640 →
Priority Claims (1)
JP 2013-194639 · Sep 19, 2013 · national
Continuity (1)
Related Publication 20150081272A1 · Mar 19, 2015