Electronic apparatus, control method thereof and electronic system
An electronic apparatus, including a processor connected with a microphone, a memory and a communication interface, and configured to: based on receiving a user voice through the microphone, acquire an operation result by inputting the user voice into the first neural network model, and identify at least one device corresponding to the user voice by inputting the operation result into the second neural network model, and control the communication interface to transmit the operation result to the at least one device, wherein the first neural network model is configured to, after only some layers of a third neural network model trained to identify a text from a voice are additionally trained, include only the additionally trained some layers, and wherein the second neural network model is trained to identify a device corresponding to a voice.
1 . An electronic apparatus comprising:
a microphone;
a communication interface; and
at least one processor; and
a memory configured to store a first neural network model, a second neural network model, and instructions which, when executed by the at least one processor, cause the electronic apparatus to:
based on receiving a user voice during a plurality of predetermined intervals through the microphone, perform a device identification operation during each predetermined interval from among the plurality of predetermined intervals while the user voice is being received, wherein the device identification operation comprises:
acquiring an operation result by inputting the user voice into the first neural network model,
identifying at least one device corresponding to the user voice by inputting the operation result into the second neural network model, and
controlling the communication interface to transmit the operation result to the at least one device,
wherein the first neural network model includes only some layers of a third neural network model,
wherein the some layers are obtained by:
training the third neural network model comprising the some layers and remaining layers to identify a text from a voice,
fixing weight values of the remaining layers, and
additionally training the some layers based on a plurality of sample user voices corresponding to the electronic apparatus and a plurality of sample texts corresponding to the plurality of sample user voices,
wherein, when the some layers are additionally trained, the remaining layers are not additionally trained, and
wherein the second neural network model is trained to identify a device corresponding to the voice.
2 . The electronic apparatus of claim 1 , wherein the instructions are further configured to cause the electronic apparatus to:
acquire the operation result in a predetermined time unit by inputting the user voice into the first neural network model in the predetermined time unit,
identify the at least one device in the predetermined time unit by inputting the operation result acquired in the predetermined time unit into the second neural network model, and
control the communication interface to transmit the operation result acquired in the predetermined time unit to the identified at least one device in the predetermined time unit.
3 . The electronic apparatus of claim 1 , wherein the memory is further configured to store information about a plurality of devices and information about a plurality of projection layers, and
wherein the instructions further cause the electronic apparatus to:
identify information about a second dimension that can be processed at the at least one device based on the information about the plurality of devices,
based on the operation result having a first dimension different from the second dimension, change the operation result to have the second dimension based on projection layers corresponding to the first dimension and the second dimension among the plurality of projection layers, and
control the communication interface to transmit the changed operation result having the second dimension to the at least one device.
4 . The electronic apparatus of claim 1 , wherein the memory is further configured to store information about a plurality of devices, and
wherein the instructions further cause the electronic apparatus to:
based on identifying that a voice recognition function is not provided in the at least one device based on the information about the plurality of devices, acquire the text corresponding to the user voice by inputting the operation result into the remaining layers, and
control the communication interface to transmit the acquired text to the at least one device.
5 . The electronic apparatus of claim 1 , wherein the instructions further cause the electronic apparatus to:
acquire scores for a plurality of devices by inputting the operation result into the second neural network model, and
control the communication interface to transmit the operation result to devices having scores greater than or equal to a threshold value among the acquired scores.
6 . The electronic apparatus of claim 1 ,
wherein the memory is further configured to store information about a plurality of projection layers and the remaining layers of the third neural network model, and
wherein the instructions further cause the electronic apparatus to:
based on receiving a first response from the at least one device after transmitting the operation result to the at least one device, control the communication interface to transmit a subsequent operation result to the at least one device, and
based on receiving a second response from the at least one device after transmitting the operation result to the at least one device, process the operation result with one of the plurality of projection layers or input the operation result into the remaining layers.
7 . The electronic apparatus of claim 6 , wherein the instructions further cause the electronic apparatus to:
based on the second response including information about a second dimension that can be processed at the at least one device, change a first dimension of the operation result based on projection layers corresponding to the first dimension and the second dimension among the plurality of projection layers, and control the communication interface to transmit the changed operation result to the at least one device, and
based on the second response including information that operation information cannot be processed, acquire the text corresponding to the user voice by inputting the operation result into the remaining layers, and control the communication interface to transmit the acquired text to the at least one device.
8 . The electronic apparatus of claim 1 , wherein the at least one device is configured to:
acquire the text corresponding to the user voice by inputting the operation result into a fourth neural network model stored in the at least one device, and
perform an operation corresponding to the acquired text, and
wherein the fourth neural network model is configured to fix weight values of the some layers, and after the remaining layers of the third neural network model are additionally trained based on the plurality of sample user voices corresponding to the at least one device and the plurality of sample texts corresponding to the plurality of sample user voices, include only the additionally trained remaining layers.
9 . A control method of an electronic apparatus, the method comprising:
based on receiving a user voice during a plurality of predetermined intervals, performing a device identification operation during each predetermined interval from among the plurality of predetermined intervals while the user voice is being received, wherein the device identification operation comprises:
acquiring an operation result by inputting the user voice into a first neural network model;
identifying at least one device corresponding to the user voice by inputting the operation result into a second neural network model; and
transmitting the operation result to the at least one device,
wherein the first neural network model includes only some layers of a third neural network model,
wherein the some layers are obtained by:
training the third neural network model comprising the some layers and remaining layers to identify a text from a voice,
fixing weight values of the remaining layers, and
additionally training the some layers based on a plurality of sample user voices corresponding to the electronic apparatus and a plurality of sample texts corresponding to the plurality of sample user voices,
wherein, when the some layers are additionally trained, the remaining layers are not additionally trained, and
wherein the second neural network model is trained to identify a device corresponding to the voice.
10 . The control method of claim 9 ,
wherein the acquiring comprises: acquiring the operation result in a predetermined time unit by inputting the user voice into the first neural network model in the predetermined time unit,
wherein the identifying comprises identifying the at least one device in the predetermined time unit by inputting the operation result acquired in the predetermined time unit into the second neural network model, and
wherein the transmitting comprises transmitting the operation result acquired in the predetermined time unit to the identified at least one device in the predetermined time unit.
11 . The control method of claim 9 , further comprising:
identifying information about a second dimension that can be processed at the at least one device based on information about a plurality of devices; and
based on the operation result being having a first dimension different from the second dimension, changing the operation result to have the second dimension based on projection layers corresponding to the first dimension and the second dimension among a plurality of projection layers,
wherein the transmitting comprises transmitting the changed operation result having the second dimension to the at least one device.
12 . The control method of claim 9 , further comprising:
based on identifying that a voice recognition function is not provided in the at least one device based on information about a plurality of devices, acquiring the text corresponding to the user voice by inputting the operation result into the remaining layers of the third neural network model,
wherein the transmitting comprises transmitting the acquired text to the at least one device.
13 . The control method of claim 9 , wherein the identifying comprises acquiring scores for a plurality of devices by inputting the operation result into the second neural network model, and
wherein the transmitting comprises transmitting the operation result to devices having scores greater than or equal to a threshold value among the acquired scores.
14 . The control method of claim 9 , further comprising:
based on receiving a first response from the at least one device after transmitting the operation result to the at least one device, transmitting a subsequent operation result to the at least one device; and
based on receiving a second response from the at least one device after transmitting the operation result to the at least one device, processing the operation result with one of a plurality of projection layers or inputting the operation result into the remaining layers of the third neural network model.
15 . The electronic apparatus of claim 1 , wherein the instructions further cause the electronic apparatus to determine whether the at least one device has a capacity to process the operation result, and
wherein based on the electronic apparatus determining that the at least one device does not have the capacity to process the operation result, the instructions further cause the electronic apparatus to obtain an additional operation result by inputting the operation result into the remaining layers of the third neural network model, and control the communication interface to transmit the additional operation result to the at least one device.
16 . The electronic apparatus of claim 1 , wherein each of the plurality of predetermined intervals is equal to approximately 25 milliseconds.