Artificial intelligence apparatus and method for estimating sound source localization thereof
An artificial intelligence (AI) apparatus including a memory and a processor configured to estimate a sound source localization based on at least one of image information, sound source information, and sensor information stored in the memory. The processor is configured to pre-process at least one of the image information, the sound source information, or the sensor information to generate test data, input the test data into a pre-trained AI model to estimate the sound source localization, calculate a sound source localization estimation evaluation score of the AI model for the test data, classify the test data into validation data based on the calculated sound source localization estimation evaluation score, change the AI model based on the classified validation data, and input the test data into the changed AI model to update the AI model.
1 . An artificial intelligence apparatus comprising:
a memory;
an input sensor; and
a processor configured to:
receive input sensor information related to a target object via the input sensor;
obtain object data from the input sensor information;
group the object data to generate a grouped data set corresponding to the target object;
obtain a first estimate of a position of the target object based on the object data;
acquire identification information of at least one external device positioned around the target object at the first estimated position;
receive additional input sensor information related to the target object from the at least one external device positioned around the target object and extract additional object data from the additional input sensor information;
group the additional object data in the grouped data set corresponding to the target object; and
obtain at least one additional estimate of the position of the target object based on the additional object data,
wherein the processor is further configured to:
generate test data set comprising at least one of the object data or the additional object data;
input the test data into a pre-trained artificial intelligence model to estimate a sound source localization and provide a sound source localization estimation result information including a position, action, and moving direction of the target object;
calculate a sound source localization estimation evaluation score of the artificial intelligence model for the test data;
classify the test data into validation data based on the calculated sound source localization estimation evaluation score;
retrain the artificial intelligence model based on the validation data; and
input the test data into the retrained artificial intelligence model to update the artificial intelligence model.
2 . The artificial intelligence apparatus according to claim 1 , wherein the input sensor is one of a plurality of input sensors configured to obtain at least image information, sound information, or sensor based information of the target object.
3 . The artificial intelligence apparatus according to claim 2 , wherein the processor is further configured to:
perform pre-processing so that object image data of the target object is extracted from the image information;
perform pre-processing so that object sound data corresponding to the target object is extracted from the sound information; and
perform pre-processing so that object sensor data corresponding to the target object is extracted from the sensor based information.
4 . The artificial intelligence apparatus according to claim 3 , wherein
the processor is further configured to:
generate a grouped data set for each of a plurality of target objects based on at least object image data, object sound data, or object sensor data extracted for each target object using image information, sound information, or sensor based information of each target object received from the plurality of input sensors.
5 . The artificial intelligence apparatus according to claim 4 , wherein a grouped data set for a particular target object comprises object sound data based on sound information of the particular target object received from a plurality of devices disposed around the particular target object, which corresponds to object image data of the particular target object grouped in the grouped data set.
6 . The artificial intelligence apparatus according to claim 4 , wherein a grouped data set for a particular target object comprises object sensor data based on sensor based information of the particular target object received from a plurality of devices disposed around the particular target object, which corresponds to object image data of the particular target object grouped in the grouped data set.
7 . The artificial intelligence apparatus according to claim 4 , wherein the processor is further configured to generate test data set comprising at least one of the object image data, the object sound data, or the object sensor data for each target object.
8 . The artificial intelligence apparatus according to claim 1 , wherein, based on there being a plurality of target objects, the processor is further configured to provide sound source localization estimation result information comprising a position, action, and moving direction of each target object.
9 . The artificial intelligence apparatus according to claim 1 , wherein the processor is further configured to:
analyze behavior of the target object in an indoor space based on the sound source localization estimation result information; and
provide at least one of a control service of a device disposed in the indoor space, a recommendation information service, or a notification information transmission service to an external server and an external terminal corresponding to an action of the target object.
10 . The artificial intelligence apparatus according to claim 1 , wherein the processor is further configured to:
calculate the sound source localization estimation evaluation score of the artificial intelligence model for the test data based on the sound source localization estimation result; and
match the sound source localization estimation evaluation score with the corresponding test data.
11 . The artificial intelligence apparatus according to claim 1 , wherein the processor is further configured to:
classify the test data into the validation data based on the sound source localization estimation evaluation score being equal to or greater than a preset reference score; and
disregard the test data corresponding to the sound source localization estimation evaluation score based on the sound source localization estimation evaluation score being less than the preset reference score.
12 . The artificial intelligence apparatus according to claim 1 , wherein the processor is further configured to:
change the artificial intelligence model by inputting the validation data into the artificial intelligence model to retrain the artificial intelligence model.
13 . The artificial intelligence apparatus according to claim 1 , wherein the processor is further configured to update the artificial intelligence model by:
inputting new test data into the retrained artificial intelligence model to update the artificial intelligence model; and
re-performing estimation of the sound source localization.
14 . A method for estimating a sound source localization of an artificial intelligence apparatus, the method comprising:
receiving input sensor information related to a target object via the input sensor;
obtaining object data from the input sensor information;
grouping the object data to generate a grouped data set corresponding to the target object;
obtaining a first estimate of a position of the target object based on the object data;
acquiring identification information of at least one external device positioned around the target object at the first estimated position;
receiving additional input sensor information related to the target object from the at least one external device positioned around the target object and extract additional object data from the additional input sensor information;
grouping the additional object data in the grouped data set corresponding to the target object; and
obtaining at least one additional estimate of the position of the target object based on the additional object data,
the method further comprising:
generating test data set comprising at least one of the object data or the additional object data;
inputting the test data into a pre-trained artificial intelligence model to estimate a sound source localization and provide a sound source localization estimation result information including a position, action, and moving direction of the target object;
calculating a sound source localization estimation evaluation score of the artificial intelligence model for the test data;
classifying the test data into validation data based on the calculated sound source localization estimation evaluation score;
retraining the artificial intelligence model based on the validation data; and
inputting the test data into the retrained artificial intelligence model to update the artificial intelligence model.