Apparatus and method for speech recognition using user prompt
View Patent ↗An apparatus for speech recognition includes a user prompt playback unit configured to change a user prompt that induces utterance for speech recognition into a sound source and to play the sound source. The apparatus further includes a microphone detection signal extraction unit configured to extract a microphone detection signal when a user's speech is received as a speech signal through a microphone. The apparatus also includes a user prompt removal unit configured to remove the user prompt from the microphone detection signal. The apparatus further includes a speech recognition unit configured to recognize a speech based on a value from which the user prompt is removed. The apparatus also includes a response output unit configured to output a response related to the speech based on a result of recognizing the speech.
1 . An apparatus for speech recognition, the apparatus comprising:
a user prompt playback unit configured to change a user prompt that induces utterance for speech recognition into a sound source and play the sound source;
a microphone detection signal extraction unit configured to extract a microphone detection signal when a speech of a user is received as a speech signal through a microphone;
a user prompt removal unit configured to remove the user prompt from the microphone detection signal;
a speech recognition unit configured to recognize a speech based on a value from which the user prompt is removed;
a response output unit configured to output a response related to the speech based on a result of recognizing the speech;
a selection domain extraction unit configured to extract a domain value (Rdn) and a speech recognition result value from the result of recognizing the speech;
a text-to-speech (TTS) result acquisition unit configured to acquire a result value (Rt) output as text from the speech recognition result value;
a keyword extraction unit configured to extract a keyword from the result value (Rt) and transmit the extracted keyword to a server;
a sound source acquisition unit configured to acquire a sound source (Rdn(x)) corresponding to the keyword extracted from the server;
a mixing unit configured to mix a sound source (Rdn(x)) corresponding to the keyword extracted from the server, the user prompt (Mp), and the result value (Rt); and
a result output unit configured to output a mixed result,
wherein the sound source acquisition unit is further configured to acquire, from a database, a background sound corresponding to a domain value (Rdn) output from the speech recognition result,
wherein the microphone detection signal includes a user prompt (Mp), a noise (En) introduced into a vehicle, and speech (Vs) contents, and
wherein the speech (Vs) contents include speech data identified as beginning-of-speech (BOS) based on an input signal differing from a predetermined sound source pattern being detected and include speech data identified as end-of-speech (EOS) based on only the predetermined sound source pattern being detected.
2 . The apparatus of claim 1 , wherein the noise (En) introduced into the vehicle includes at least one of a wind noise or a road noise.
3 . The apparatus of claim 1 , wherein the sound source corresponding to the domain is pre-stored in the database.
4 . The apparatus of claim 1 , wherein the background sound is acquired by mapping a user-selected sound or a sound stored or recorded using a user terminal to an identifier of the database.
5 . A method for speech recognition, the method comprising:
changing a user prompt that induces utterance for speech recognition into a sound source and playing the sound source;
extracting a microphone detection signal when a speech of a user is received as a speech signal through a microphone;
removing the user prompt from the microphone detection signal;
recognizing a speech based on a value from which the user prompt is removed;
outputting a response related to the speech based on a result of recognizing the speech;
extracting a domain value (Rdn) and a speech recognition result value from the result of recognizing the speech;
acquiring a result value (Rt) output as text from the speech recognition result value;
extracting a keyword from the result value (Rt) output as text and transmitting the extracted keyword to a server;
acquiring a sound source (Rdn(x)) corresponding to the keyword extracted from the server;
mixing the sound source Rdn(x) corresponding to the keyword extracted from the server, the user prompt (Mp), and the result value (Rt) output as text; and
outputting a mixed result,
wherein acquiring the sound source Rdn(x) further includes acquiring, from a database, a background sound corresponding to a domain value (Rdn) output from the speech recognition result,
wherein the microphone detection signal includes a user prompt (Mp), a noise (En) introduced into a vehicle, and speech (Vs) contents, and
wherein the speech (Vs) contents include speech data identified as beginning-of-speech (BOS) based on an input signal differing from a predetermined sound source pattern being detected and include speech data identified as end-of-speech (EOS) based on only the predetermined sound source pattern being detected.
6 . The method of claim 5 , wherein the noise (En) introduced into the vehicle includes at least one of wind noise or road noise.
7 . The method of claim 5 , wherein the sound source corresponding to the domain is pre-stored in the database.
8 . The method of claim 5 , wherein the background sound is acquired by mapping a user-selected sound or a sound stored or recorded using a user terminal to an identifier of the database.
9 . A non-transitory computer-readable recording medium, which stores a computer program including computer-executable instructions configured to be executable by an apparatus for speech recognition including a processor, to cause the processor to execute:
a function including changing a user prompt that induces utterance for speech recognition into a sound source and playing the sound source;
a function including extracting a microphone detection signal when a speech of a user is received as a speech signal through a microphone;
a function including removing the user prompt from the microphone detection signal;
a function including recognizing a speech based on a value from which the user prompt is removed;
a function including outputting a response related to the speech based on a result of recognizing the speech;
a function including extracting a domain value (Rdn) and a speech recognition result value from the result of recognizing the speech;
a function including acquiring a result value (Rt) output as text from the speech recognition result value;
a function including extracting a keyword from the result value (Rt) and transmit the extracted keyword to a server;
a function including acquiring a sound source (Rdn(x)) corresponding to the keyword extracted from the server;
a function including mixing a sound source (Rdn(x)) corresponding to the keyword extracted from the server, the user prompt (Mp), and the result value (Rt); and
a function including outputting a mixed result,
wherein the function of acquiring a sound source (Rdn(x)) further comprises acquiring, from a database, a background sound corresponding to a domain value (Rdn) output from the speech recognition result,
wherein the microphone detection signal includes a user prompt (Mp), a noise (En) introduced into a vehicle, and speech (Vs) contents, and
wherein the speech (Vs) contents include speech data identified as beginning-of-speech (BOS) based on an input signal differing from a predetermined sound source pattern being detected and include speech data identified as end-of-speech (EOS) based on only the predetermined sound source pattern being detected.
10 . The computer-readable recording medium of claim 9 , wherein the noise (En) introduced into the vehicle includes at least one of wind noise or road noise.
11 . The computer-readable recording medium of claim 9 , wherein the background sound is acquired by mapping a user-selected sound or a sound stored or recorded using a user terminal to an identifier of the database.