Speaker recognition device and operating method thereof
An object of the present disclosure is to ensure speaker recognition performance even if voice uttered by a voice input device different from a voice input device of voice uttered in a speaker registration process.
1 . A speaker recognition device comprising:
a communication interface configured to communicate with an electronic device;
a memory configured to store a speaker profile including a first voice feature set including voice features of a first voice signal and a first voice input channel matched to the first voice signal; and
a processor configured to receive a second voice signal corresponding to a voice command uttered by a speaker and information on a second voice input channel indicating a device obtaining the voice command, from an electronic device,
obtain a similarity between the first voice feature set and a second voice feature set including voice features of the second voice signal, based on that the second voice input channel is not the first voice input channel stored in the memory,
transmit a notification indicating that the speaker is identified, to the electronic device, based on the obtained similarity being equal to or greater than a first similarity,
transmit a notification for checking that the speaker is matched to a pre-registered speaker, to the electronic device, based on the similarity being less than the first similarity and equal to or greater than a second similarity less than the first similarity, and
transmit a notification for registering a new speaker to the electronic device, based on the similarity being less than the second similarity.
2 . The speaker recognition device of claim 1 , wherein the processor compares a first embedding vector indicating the first voice feature set with a second embedding vector indicating the second voice feature set to obtain the similarity.
3 . The speaker recognition device of claim 1 , wherein each of the first voice input channel and the second voice input channel includes any one of a remote control device, the electronic device, and a mobile device, each including a microphone.
4 . A speaker recognition device comprising:
a communication interface configured to communicate with an electronic device;
a memory configured to store a speaker profile including a first voice feature set including voice features of a first voice signal and a first voice input channel matched to the first voice signal; and
a processor configured to receive a second voice signal corresponding to a voice command uttered by a speaker and information on a second voice input channel indicating a device obtaining the voice command, from an electronic device,
obtain a similarity between the first voice feature set and a second voice feature set including voice features of the second voice signal, based on the second voice input channel being different from the first voice input channel stored in the memory,
transmit a notification indicating that the speaker is identified, to the electronic device, based on the obtained similarity being equal to or greater than a first similarity, and
transmit a notification for registering the second voice input channel to the electronic device, based on the similarity being less than the first similarity and equal to or greater than a second similarity less than the first similarity.
5 . The speaker recognition device of claim 4 , wherein the processor is configured to obtain additional voice features of an additional voice signal corresponding to an additional voice command obtained through the second voice input channel, from the electronic device, and update the information on the second voice input channel and the additional voice features, to the speaker profile.
6 . An operating method of a speaker recognition device, the method comprising:
storing a speaker profile including a first voice feature set including voice features of a first voice signal and a first voice input channel matched to the first voice signal;
receiving a second voice signal corresponding to a voice command uttered by a speaker and information on a second voice input channel indicating a device obtaining the voice command, from an electronic device;
obtaining a similarity between the first voice feature set and a second voice feature set including voice features of the second voice signal, based on that the second voice input channel is not the stored first voice input channel; and
transmitting a notification indicating that the speaker is identified, to the electronic device, based on the obtained similarity being equal to or greater than a first similarity,
transmitting a notification for checking that the speaker is matched to a pre-registered speaker, to the electronic device, based on the similarity being less than the first similarity and equal to or greater than a second similarity less than the first similarity, and
transmitting a notification for registering a new speaker to the electronic device, based on the similarity being less than the second similarity.
7 . The method of claim 6 , wherein the obtaining of the similarity includes comparing a first embedding vector indicating the first voice feature set with a second embedding vector indicating the second voice feature set to obtain the similarity.
8 . The method of claim 6 , further comprising, based on that the similarity is less than the first similarity and is equal to or greater than a second similarity less than the first similarity, transmitting a notification for registering the second voice input channel to the electronic device.
9 . The method of claim 8 , further comprising:
obtaining additional voice features of an additional voice signal corresponding to an additional voice command obtained through the second voice input channel, from the electronic device; and
updating the information on the second voice input channel and the additional voice features, to the speaker profile.