Method and apparatus for providing voice recognition service
View Patent ↗A method and apparatus for providing a voice recognition service can include separating a user utterance from noise and converting the user utterance into a text to generate content of the user utterance, extracting, from the content of the user utterance, call words information, domain information, end service name information, and operations information, generating a corrected user command by correcting a user command, and generating response information by using the corrected user command.
1 . An apparatus for providing a voice recognition service, comprising:
at least one processor; and
a memory storing instructions that, when executed by the at least one processor, cause the apparatus to:
separate a user utterance, in which call word information for specifying one of a plurality of voice assistants is omitted, from noise;
convert the user utterance into text to generate content of the user utterance;
extract, from the content of the user utterance, extracted information including at least one of domain information, end service name information, and operations information;
generate a corrected user command including the call word information, the domain information, the end service name information, and the operations information, based on usage counts of combinations of each voice assistant of the plurality of voice assistants and the extracted information;
convert the corrected user command into a voice command by generating speech through a text-to-speech (TTS) process;
generate a control signal for a vehicle in response to the voice command; and
control the vehicle based on the control signal.
2 . The apparatus of claim 1 , wherein execution of the instructions by the at least one processor further causes the apparatus to:
classify an intention of the user utterance, by using a natural language understanding engine.
3 . The apparatus of claim 1 , wherein execution of the instructions by the at least one processor further causes the apparatus to:
determine the call word information by using a voice assistant usage history of a voice assistant, in response to the content of the user utterance not including the call word information but including the end service name information.
4 . The apparatus of claim 1 , wherein execution of the instructions by the at least one processor further causes the apparatus to:
determine the call word information and the end service name information by using a voice assistant usage history of a voice assistant, an end service usage history of an end service, and a recently utilized service usage history of a recently utilized service, in response to the content of the user utterance not including the call word information and the end service name information.
5 . The apparatus of claim 1 , wherein execution of the instructions by the at least one processor further causes the apparatus to:
determine the call word information and the end service name information by using an end service based on the domain information, in response to the content of the user utterance including the domain information not being supported by an invoked voice assistant.
6 . The apparatus of claim 1 , wherein execution of the instructions by the at least one processor further causes the apparatus to:
determine the call word information and the end service name information by using a voice assistant activated earliest among the plurality of voice assistants, in response to the content of the user utterance including the domain information that has no usage history.
7 . The apparatus of claim 1 , wherein:
execution of the instructions by the at least one processor further causes the apparatus to store the usage counts by using a source database and a virtual database copy,
the source database stores all existing usage histories, and
the virtual database copy is a copy of the source database.
8 . A method of voice recognition service, the method comprising:
separating a user utterance, in which call word information for specifying one of a plurality of voice assistants is omitted, from noise;
converting the user utterance into text to generate content of the user utterance;
extracting, from the content of the user utterance, extracted information including at least one of domain information, end service name information, and operations information;
generating a corrected user command including the call word information, the domain information, the end service name information, and the operations information, based on usage counts of combinations of each voice assistant of the plurality of voice assistants and the extracted information;
converting the corrected user command into a voice command by generating speech through a text-to-speech process;
generating a control signal for a vehicle in response to the voice command; and
controlling the vehicle based on the control signal.
9 . The method of claim 8 , wherein extracting, from the content of the user utterance, the extracted information including at least one of the domain information, the end service name information, and the operations information comprises:
classifying an intention of the user utterance by using a natural language understanding engine to extract the domain information, the end service name information, and the operations information.
10 . The method of claim 8 , wherein generating the corrected user command comprises;
determining the call word information by using a voice assistant usage history of a voice assistant to generate the corrected user command, in response to the content of the user utterance not including the call word information but including the end service name information.
11 . The method of claim 8 , wherein generating the corrected user command comprises:
determining the call word information and the end service name information by using a voice assistant usage history of a voice assistant, an end service usage history of an end service, and a recently utilized service usage history of a recently utilized service to generate the corrected user command, in response to the content of the user utterance not including the call word information and the end service name information.
12 . The method of claim 8 , wherein generating the corrected user command comprises:
determining the call word information and the end service name information by using an end service based on the domain information to generate the corrected user command, in response to the content of the user utterance including the domain information not being supported by an invoked voice assistant.
13 . The method of claim 8 , wherein generating the corrected user command comprises:
determining the call word information and the end service name information by using a voice assistant activated earliest among the plurality of voice assistants to generate the corrected user command, in response to the content of the user utterance including the domain information having no usage history.
14 . A non-transitory computer-readable medium storing a computer program including computer-executable instructions for causing, when executed by a computer, the computer to perform steps of:
separating a user utterance, in which call word information for specifying one of a plurality of voice assistants is omitted, from noise;
converting the user utterance into text to generate content of the user utterance;
extracting, from the content of the user utterance, extracted information including at least one of domain information, end service name information, and operations information;
generating a corrected user command including the call word information, the domain information, the end service name information, and the operations information, based on usage counts of combinations of each voice assistant plurality of voice assistants and the extracted information;
converting the corrected user command into a voice command by generating speech through a text-to-speech process;
generating a control signal for a vehicle in response to the voice command; and
controlling the vehicle based on the control signal.
15 . The apparatus of claim 1 , wherein execution of the instructions by the at least one processor further causes the apparatus to:
operate a source database and a virtual database copy to manage the usage counts of combinations,
wherein:
the virtual database copy is generated by copying the source database when the vehicle is started, and
the source database is updated by using the virtual database copy when the vehicle is turned off.
16 . The apparatus of claim 1 , wherein execution of the instructions by the at least one processor further causes the apparatus to:
update the usage counts of combinations when a predetermined amount of time has passed after a service corresponding to the corrected user command is provided to a user.