IP Library Granted Patent US 11,587,553
Granted Patent B2
US 11,587,553 · App. 16/968,126 · Granted Feb 21, 2023

Appropriate utterance estimate model learning apparatus, appropriate utterance judgement apparatus, appropriate utterance estimate model learning method, appropriate utterance judgement method, and program

Inventors: Takashi Nakamura (Tokyo, JP); Takaaki Fukutomi (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G10L15/063G10L15/05G10L15/16G10L15/22G10L15/28G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,587,553
App. No.
16/968,126
Granted
Feb 21, 2023
Kind
B2
Abstract

Provided is technology for assessing whether uttered speech detected from input speech is speech suited to a prescribed purpose. A method comprises detecting, from input speech including speech uttered by a speaker and noise, the uttered speech corresponding to the speech uttered by the speaker, extracting an acoustic feature of the uttered speech, generating, from the uttered speech, a speech recognition result set with a recognition score, generating, from the speech recognition result set with the recognition score, a speech recognition result word vector expression set and a speech recognition result part-of-speech vector expression set, generating a target utterance estimation model, providing, using the target utterance estimation model, a probability of the uttered speech being suited to the prescribed purpose, and outputting the uttered speech and the speech recognition result set with the recognition score, the the uttered speech suitable to the prescribed purpose.

Claims (64)

1. A computer-implemented method for generating aspects of utterance, the method comprising:

receiving an input speech, the input speech comprising a speech uttered by a speaker and a noise;

detecting, based the input speech, an utterance, the utterance corresponding to the speech uttered by the speaker;

extracting an acoustic feature from the uttered speech;

generating a set of speech recognition results with recognition scores based on the uttered speech;

generating a set of speech-recognition-result word vector expressions and a set of speech-recognition-result part-of-speech vector expressions based on the set of speech recognition results with recognition scores;

generating a target utterance estimation model based on the extracted acoustic feature, the generated set of speech recognition results with recognition scores, the generated set of speech-recognition-result word vector expressions, an utterance time length of the uttered speech, and the generated set of speech-recognition-result part-of-speech vector expressions;

providing, by the generated target utterance estimation model, a probability of the uttered speech detected from the input speech being an utterance suitable for a predetermined purpose, wherein the generated target utterance estimation model predicts at least a part of the uttered speech as a sudden noise based at least on a combination of the generated set of speech-recognition-result word vector expressions and the utterance time length of the uttered speech, wherein the predetermined purpose excludes the sudden noise; and

causing, based on the probability of the uttered speech detected from the input speech being the utterance suitable for the predetermined purpose, removal of the at least a part of the uttered speech as the sudden noise.

2. The method of claim 1 , wherein the target utterance estimation model is based at least on a sequence of a combination of:

a recognition score of a word based at least on the set of speech recognition results with recognition scores,

a word vector of the word based at least on the set of speech-recognition-result word vector expressions,

a part-of-speech vector of the word based on the set of speech-recognition-result part-of-speech vector expressions, and

an acoustic feature of the word based on the acoustic feature of the uttered speech.

3. The method of claim 1 , the method further comprising:

rejecting the input speech as a background noise based on the probability of the uttered speech from the input speech being the utterance suitable for a predetermined purpose, wherein the predetermined purpose includes a spoken dialogue.

4. The method of claim 1 , wherein the target utterance estimation model is a model learned by a neural network, the neural network processing time-series data.

5. The method of claim 1 , the method further comprising:

receiving, by the target utterance estimation model, a correct answer of the input speech for training the target utterance estimation model, the correct answer being the utterance in a spoken dialogue.

6. The method of claim 1 , wherein each of the recognition scores comprises a numerical value based on one or more of a confidence score of speech recognition, an acoustic score indicating a similarity between the acoustic feature of the input speech and a feature based on an acoustic model, and a language score indicating a degree of matching between the speech recognition results and a language model.

7. The method of claim 1 , wherein the set of speech-recognition-result word vector expressions comprises a vector generated for each word in the set of speech recognition results with a space between adjacent words based on a morphological analysis, and wherein the set of speech-recognition-result part-of-speech vector expressions comprises a vector generated for each part-of-speech for words in the set of speech recognition results.

8. A system comprising:

a processor; and

a memory storing computer executable instructions that when executed by the processor cause the system to:

receive an input speech, the input speech comprising a speech uttered by a speaker and a noise;

detect, based the input speech, an utterance, the utterance corresponding to the speech uttered by the speaker;

extract an acoustic feature from the uttered speech;

generate a set of speech recognition results with recognition scores based on the uttered speech;

generate a set of speech-recognition-result word vector expressions and a set of speech-recognition-result part-of-speech vector expressions based on the set of speech recognition results with recognition scores;

generate a target utterance estimation model based on the extracted acoustic feature, the generated set of speech recognition results with recognition scores, the generated set of speech-recognition-result word vector expressions, an utterance time length of the uttered speech, and the generated set of speech-recognition-result part-of-speech vector expressions;

provide, by the generated target utterance estimation model, a probability of the uttered speech detected from the input speech being an utterance suitable for a predetermined purpose, wherein the generated target utterance estimation model predicts at least a part of the uttered speech as a sudden noise based at least on a combination of the generated set of speech-recognition-result word vector expressions and the utterance time length of the uttered speech, wherein the predetermined purpose excludes the sudden noise; and

causing, based on the probability of the uttered speech detected from the input speech being the utterance suitable for the predetermined purpose, removal of the at least a part of the uttered speech as the sudden noise.

9. The system of claim 8 , wherein the target utterance estimation model is based at least on a sequence of a combination of:

a recognition score of a word based at least on the set of speech recognition results with recognition scores,

a word vector of the word based at least on the set of speech-recognition-result word vector expressions,

a part-of-speech vector of the word based on the set of speech-recognition-result part-of-speech vector expressions, and

an acoustic feature of the word based on the acoustic feature of the uttered speech.

10. The system of claim 8 , the computer-executable instructions when executed further causing the system to:

reject the input speech as a background noise based on the probability of the uttered speech from the input speech being the utterance suitable for a predetermined purpose, wherein the predetermined purpose includes a spoken dialogue.

11. The system of claim 8 , wherein the target utterance estimation model is a model learned by a neural network, the neural network processing time-series data.

12. The system of claim 8 , the computer-executable instructions when executed further causing the system to:

receive, by the target utterance estimation model, a correct answer of the input speech for training the target utterance estimation model, the correct answer being the utterance in a spoken dialogue.

13. The system of claim 8 , wherein each of the recognition scores comprises a numerical value based on one or more of a confidence score of speech recognition, an acoustic score indicating a similarity between the acoustic feature of the input speech and a feature based on an acoustic model, and a language score indicating a degree of matching between the speech recognition results and a language model.

14. The system of claim 8 , wherein the set of speech-recognition-result word vector expressions comprises a vector generated for each word in the set of speech recognition results with a space between adjacent words based on a morphological analysis, and wherein the set of speech-recognition-result part-of-speech vector expressions comprises a vector generated for each part-of-speech for words in the set of speech recognition results.

15. A computer-readable non-transitory recording medium storing computer-executable instructions that when executed by a processor cause a computer system to:

receive an input speech, the input speech comprising a speech uttered by a speaker and a noise;

detect, based the input speech, an utterance, the utterance corresponding to the speech uttered by the speaker;

extract an acoustic feature from the uttered speech;

generate a set of speech recognition results with recognition scores based on the uttered speech;

generate a set of speech-recognition-result word vector expressions and a set of speech-recognition-result part-of-speech vector expressions based on the set of speech recognition results with recognition scores;

generate a target utterance estimation model based on the extracted acoustic feature, the generated set of speech recognition results with recognition scores, the generated set of speech-recognition-result word vector expressions, an utterance time length of the uttered speech, and the generated set of speech-recognition-result part-of-speech vector expressions;

provide, by the generated target utterance estimation model, a probability of the uttered speech detected from the input speech being an utterance suitable for a predetermined purpose, wherein the generated target utterance estimation model predicts at least a part of the uttered speech as a sudden noise based at least on a combination of the generated set of speech-recognition-result word vector expressions and the utterance time length of the uttered speech, wherein the predetermined purpose excludes the sudden noise; and

causing, based on the probability of the uttered speech detected from the input speech being the utterance suitable for the predetermined purpose, removal of the at least a part of the uttered speech as the sudden noise.

16. The computer-readable non-transitory recording medium of claim 15 , wherein the target utterance estimation model is based at least on a sequence of a combination of:

a recognition score of a word based at least on the set of speech recognition results with recognition scores,

a word vector of the word based at least on the set of speech-recognition-result word vector expressions,

a part-of-speech vector of the word based on the set of speech-recognition-result part-of-speech vector expressions, and

an acoustic feature of the word based on the acoustic feature of the uttered speech.

17. The computer-readable non-transitory recording medium of claim 15 , the computer-executable instructions when executed further causing the system to:

reject the input speech as a background noise based on the probability of the uttered speech from the input speech being the utterance suitable for a predetermined purpose, wherein the predetermined purpose includes a spoken dialogue.

18. The computer-readable non-transitory recording medium of claim 15 , wherein the target utterance estimation model is a model learned by a neural network, the neural network processing time-series data.

19. The computer-readable non-transitory recording medium of claim 15 , the computer-executable instructions when executed further causing the system to:

receive, by the target utterance estimation model, a correct answer of the input speech for training the target utterance estimation model, the correct answer being the utterance in a spoken dialogue.

20. The computer-readable non-transitory recording medium of claim 15 , wherein the set of speech-recognition-result word vector expressions comprises a vector generated for each word in the set of speech recognition results with a space between adjacent words based on a morphological analysis, and wherein the set of speech-recognition-result part-of-speech vector expressions comprises a vector generated for each part-of-speech for words in the set of speech recognition results.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2022
From: NAKAMURA, TAKASHI; FUKUTOMI, TAKAAKI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 061306/0449 →
Priority Claims (1)
JP JP2018-020773 · Feb 8, 2018 · national
Continuity (1)
Related Publication 20210035558A1 · Feb 4, 2021