Artificial intelligence apparatus and method for recognizing speech of user
An artificial intelligence apparatus for recognizing speech of a user includes a microphone and a processor configured to acquire, via the microphone, first speech data including speech of a user, generate a first speech recognition result corresponding to the first speech data, perform control corresponding to the generated first speech recognition result, generate an alternative speech recognition result corresponding to the first speech data if negative feedback is acquired from the user, and perform control corresponding to the generated alternative speech recognition result.
1. An artificial intelligence apparatus for recognizing speech of a user, the artificial intelligence apparatus comprising:
a microphone; and
a processor configured to:
acquire, via the microphone, first speech data including speech of the user,
generate a first speech recognition result corresponding to the first speech data,
perform control corresponding to the generated first speech recognition result,
generate an alternative speech recognition result corresponding to the first speech data if negative feedback is acquired from the user,
perform control corresponding to the generated alternative speech recognition result,
calculate word-by-word reliability corresponding to each section in the first speech data,
convert the first speech data into first text by selecting words having highest reliability for each section, and
generate the first speech recognition result based on the first text.
2. The artificial intelligence apparatus of claim 1 , wherein the processor is configured to:
correct the word-by-word reliability corresponding to each section in the first speech data,
convert the first speech data into second text by selecting words having highest corrected reliability for each section, and
generate the alternative speech recognition result based on the second text.
3. The artificial intelligence apparatus of claim 2 , wherein the processor is configured to:
extract a named entity and a verb phrase from the first text,
determine respective domains of the extracted named entity and the extracted verb phrase, and
correct the word-by-word reliability based on the determined domains.
4. The artificial intelligence apparatus of claim 3 , wherein the processor is configured to:
determine a domain weight for each of the domains, and
correct the word-by-word reliability based on the domain weight.
5. The artificial intelligence apparatus of claim 4 , wherein the processor is configured to:
determine a dominant domain based on the determined domains,
calculate a distance from each of the determined domain to the dominant domain, and
determine a domain weight as decreasing as the calculated distance of a domain increases.
6. The artificial intelligence apparatus of claim 1 , further comprising a camera,
wherein the processor is configured to:
acquire image data via the camera,
generate an image recognition result corresponding to the image data, and
determine whether negative feedback is included in the image recognition result.
7. The artificial intelligence apparatus of claim 6 , the processor is configured to generate the image recognition result by recognizing an expression or a gesture of the user from the image data, and
wherein the negative feedback includes a frowning expression or a hand waving gesture.
8. The artificial intelligence apparatus of claim 1 , wherein the processor is configured to:
acquire second speech data via the microphone,
generate a second speech recognition result corresponding to the second speech data, and
determine whether negative feedback is included in the second speech recognition result.
9. The artificial intelligence apparatus of claim 8 , wherein the negative feedback includes negative evaluation of or negative reaction to control corresponding to the first speech recognition result.
10. A method of recognizing speech of a user, the method comprising:
acquiring, via a microphone, first speech data including speech of the user,
generating a first speech recognition result corresponding to the first speech data,
performing control corresponding to the generated first speech recognition result,
generating an alternative speech recognition result corresponding to the first speech data if negative feedback is acquired from the user, and
performing control corresponding to the generated alternative speech recognition result,
wherein the generating the first speech recognition result corresponding to the first speech data includes:
calculating word-by-word reliability corresponding to each section in the first speech data,
converting the first speech data into first text by selecting words having highest reliability for each section, and
generating the first speech recognition result based on the first text.
11. A non-transitory computer readable medium having recorded thereon a program for performing a method of recognizing speech of a user, the method comprising:
acquiring, via a microphone, first speech data including speech of the user,
generating a first speech recognition result corresponding to the first speech data,
performing control corresponding to the generated first speech recognition result,
generating an alternative speech recognition result corresponding to the first speech data if negative feedback is acquired from the user, and
performing control corresponding to the generated alternative speech recognition result,
wherein the generating the first speech recognition result corresponding to the first speech data includes:
calculating word-by-word reliability corresponding to each section in the first speech data,
converting the first speech data into first text by selecting words having highest reliability for each section, and
generating the first speech recognition result based on the first text.