Electronic apparatus and controlling method thereof
An electronic apparatus is provided. The electronic apparatus includes a communication interface with communication circuitry, a memory configured to store at least one instruction and a processor, and the processor is configured to receive a first audio recognized as a wake up word by an external device from the external device, determine whether the first audio corresponds to the wake up word by analyzing the first audio, based on determining that the first audio does not correspond to the wake up word, obtain a neural network model for detecting a wake up word misrecognition based on the first audio, and transmit information regarding the neural network model to the external device.
1 . An apparatus comprising:
a communication interface with communication circuitry;
a memory configured to store at least one instruction; and
a processor,
wherein the processor is configured to:
receive, from an external device, a first audio, recognized as a wake up word by the external device,
receive, from the external device, second audio captured subsequent to the first audio, wherein the second audio includes a user voice subsequent to an operation performed by the external device based on recognition of the first audio as the wake up word,
analyze the first audio and the second audio so as to determine whether the first audio corresponds to the wake up word,
based on determining that the first audio does not correspond to the wake up word, obtain a neural network model trained to identify audio that corresponds to a misrecognized wake up word, and
transmit information regarding the neural network model to the external device so as to enable the external device to input the first audio to the neural network model and determine whether the first audio corresponds to the misrecognized wake up word based on output from the neural network model.
2 . The apparatus of claim 1 , wherein the processor is further configured to, based on a text corresponding to the first audio not being detected, determine that the first audio does not correspond to the wake up word.
3 . The apparatus of claim 1 , wherein the processor is further configured to:
obtain a text corresponding to the first audio; and
based on a similarity between the text corresponding to the first audio and the wake up word being less than a predetermined value, determine that the first audio does not correspond to the wake up word.
4 . The apparatus of claim 1 , wherein the processor is further configured to:
obtain a text corresponding to the second audio; and
based on the text corresponding to the second audio not having a predetermined sentence structure, determine that the first audio does not correspond to the wake up word.
5 . The apparatus of in claim 1 ,
wherein the second audio includes a user voice regarding an operation performed as the external device recognizes the first audio as the wake up word, and
wherein the processor is further configured to determine whether the first audio corresponds to the wake up word by analyzing the user voice.
6 . The apparatus of claim 1 , wherein the processor is further configured to determine whether the first audio corresponds to the wake up word based on a user feedback input through a user interface (UI) provided by the external device.
7 . The apparatus of claim 1 , wherein the processor is further configured to:
based on determining that the first audio does not correspond to the wake up word, store the first audio in the memory;
identify a plurality of third audios forming a cluster from among the first audio stored in the memory; and
train the neural network model based on the plurality of third audios.
8 . A method of controlling an electronic apparatus, the method comprising:
receiving, from an external device, a first audio recognized as a wake up word by the external device;
receiving, from the external device, a second audio captured subsequent to the first audio, wherein the second audio includes a user voice subsequent to an operation performed by the external device based on recognition of the first audio as the wake up word;
analyzing the first audio and the second audio to determine whether the first audio corresponds to the wake up word;
based on determining that the first audio does not correspond to the wake up word, obtaining a neural network model trained to identify audio that corresponds to a misrecognized wake up word; and
transmitting information regarding the neural network model to the external device.
9 . The method of claim 8 , wherein the determining of whether the first audio corresponds to the wake up word comprises determining, based on a text corresponding to the first audio not being detected, that the first audio does not correspond to the wake up word.
10 . The method of claim 8 , wherein the determining of whether the first audio corresponds to the wake up word comprises:
obtaining a text corresponding to the first audio; and
based on a similarity between the text corresponding to the first audio and the wake up word being less than a predetermined value, determining that the first audio does not correspond to the wake up word.
11 . The method of claim 8 , wherein the determining of whether the first audio corresponds to the wake up word comprises:
obtaining a text corresponding to the second audio; and
based on the text corresponding to the second audio not having a predetermined sentence structure, determining that the first audio does not correspond to the wake up word.
12 . The method of claim 8 ,
wherein the second audio includes a user voice regarding an operation performed as the external device recognizes the first audio as the wake up word, and
wherein the determining of whether the first audio corresponds to the wake up word comprises analyzing the user voice.
13 . The method of claim 8 , further comprising:
based on the first audio corresponding to the wake up word, obtaining a response corresponding to the second audio; and
transmitting information regarding the obtained response to the external device.
14 . The method of claim 8 , further comprising:
determining whether the first audio corresponds to the wake up word based on a user feedback input through a user interface (UI) provided by the external device.
15 . The method of claim 8 , further comprising determining whether the first audio corresponds to the wake up word by determining whether the first audio corresponds to a misrecognition word using the neural network model.
16 . The method of claim 8 , wherein the information regarding the neural network model comprises at least one of parameters regarding the neural network model or a message requesting to download the neural network model.