IP Library › Granted Patent US 12,131,738
Granted Patent B2
US 12,131,738 · App. 17/944,401 · Granted Oct 29, 2024

Electronic apparatus and method for controlling thereof

Inventors: Jaeyoung Roh (Suwon-si, KR); Hejung Yang (Suwon-si, KR); Hojun Jin (Suwon-si, KR); Donghan Jang (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L15/22G10L15/02G10L15/26G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,131,738
App. No.
17/944,401
Granted
Oct 29, 2024
Kind
B2
Abstract

An electronic apparatus is disclosed. The electronic apparatus may include a microphone; a communication interface; a memory configured to store at least one instruction; and a processor configured to execute the at least one instruction to: obtain a user voice input for registering a wake-up voice input via the microphone; input the user voice input into a trained neural network model to obtain a first feature vector corresponding to text included in the user voice input; receive a verification data set determined based on information related to the text included in the user voice input from an external server via the communication interface; input a verification voice input included in the verification data set into the trained neural network model to obtain a second feature vector corresponding to the verification voice input; and identify whether to register the user voice input as the wake-up voice input based on a similarity between the first feature vector and the second feature vector.

Claims (49)

1. An electronic apparatus comprising:

a microphone;

a communication interface;

a memory configured to store at least one instruction; and

a processor configured to execute the at least one instruction to:

obtain a user voice input for registering a wake-up voice input via the microphone;

input the user voice input into a trained neural network model to obtain a first feature vector corresponding to text included in the user voice input;

receive a verification data set determined based on information related to the text included in the user voice input from an external server via the communication interface;

input a verification voice input included in the verification data set into the trained neural network model to obtain a second feature vector corresponding to the verification voice input; and

identify whether to register the user voice input as the wake-up voice input based on a similarity between the first feature vector and the second feature vector.

2. The apparatus of claim 1 , wherein the processor is further configured to:

recognize the user voice input to obtain the information related to the text included in the user voice input; and

transmit the information related to the text included in the user voice input to the external server via the communication interface,

wherein the external server is configured to obtain, using the information related to the text, the verification voice input based on a first phoneme sequence of a verification voice text that includes a number of common phonemes with a second phoneme sequence of the text included in the user voice input.

3. The apparatus of claim 2 , wherein the verification voice input is voice data corresponding to the verification voice text having the number of common phonemes that is equal to or greater than a threshold value.

4. The apparatus of claim 2 , wherein the processor is further configured to:

based on the similarity between the first feature vector and the second feature vector being less than the threshold value, input another verification voice input included in the verification data set into the trained neural network model to obtain a third feature vector corresponding to the another verification voice input;

compare another similarity between the first feature vector and the third feature vector; and

based on the other similarity between the first feature vector and the third feature vector being equal to or greater than the threshold value, provide a guide message requesting an additional user voice input for registering a wake-up voice input.

5. The apparatus of claim 4 , wherein the processor is further configured to;

based on a plurality of similarities between feature vectors corresponding to all verification voice inputs included in the verification data set and the first feature vector being less than the threshold value, register the user voice input as the wake-up voice input.

6. The apparatus of claim 1 , wherein the processor is further configured to:

input the user voice input into a voice recognition model to obtain the text included in the user voice; and

based on at least one of a length and a duplication of phonemes of the text included in the user voice input, identify whether to register the text included in the user voice input as the text of the wake-up voice input.

7. The apparatus of claim 6 , wherein the processor is further configured to:

based on a number of phonemes of the text included in the user voice input being less than a first threshold value or the number of phonemes of the text included in the user voice input being duplicated by greater than a second threshold value, provide a guide message requesting an utterance of an additional user voice input including another text for registering a wake-up word.

8. The apparatus of claim 7 , wherein the guide message is configured to include a message for recommending the another text determined based on usage history information of the electronic apparatus as the text of the wake-up voice input.

9. The apparatus of claim 1 , wherein the processor is further configured to:

input the user voice input into a trained voice identification model to obtain a feature value indicating whether the user voice input is a user voice input uttering a specific text; and

identify whether to register the user voice input as the wake-up voice input based on the feature value.

10. The apparatus of claim 9 , wherein the processor is further configured to:

based on the feature value being less than the threshold value, provide the guide message requesting the additional user voice input for registering the wake-up voice input.

11. A method of controlling an electronic apparatus, the method comprising:

obtaining a user voice input for registering a wake-up voice input;

inputting the user voice input into a trained neural network model to obtain a first feature vector corresponding to text included in the user voice input;

receiving a verification data set determined based on information related to the text included in the user voice input from an external server;

inputting a verification voice input included in the verification data set into the trained neural network model to obtain a second feature vector corresponding to the verification voice input; and

identifying whether to register the user voice input as the wake-up voice input based on a similarity between the first feature vector and the second feature vector.

12. The method of claim 11 , further comprising:

recognizing the user voice input to obtain the information related to the text included in the user voice input; and

transmitting the information related to the text included in the user voice input to the external server,

wherein the server is configured to obtain the verification voice based on a first phoneme sequence of a verification voice text that includes a number of common phonemes with a second phoneme sequence of the text included in the user voice input.

13. The method of claim 12 , wherein the verification voice input is voice data corresponding to the verification voice text having the number of common phonemes that is equal to or greater than a threshold value.

14. The method of claim 12 , further comprising:

based on the similarity between the first feature vector and the second feature vector being less than the threshold value, inputting another verification voice input included in the verification data set into the trained neural network model to obtain a third feature vector corresponding to the another verification voice input;

comparing another similarity between the first feature vector and the third feature vector; and

based on the other similarity between the first feature vector and the third feature vector being equal to or greater than the threshold value, providing a guide message requesting an additional user voice input for registering a wake-up voice input.

15. The method of claim 14 , further comprising:

based on a plurality of similarities between feature vectors corresponding to all verification voice inputs included in the verification data set and the first feature vector being less than the threshold value, registering the user voice input as the wake-up voice input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2022
From: ROH, JAEYOUNG; YANG, HEJUNG; JIN, HOJUN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061433/0380 →
Priority Claims (1)
KR 10-2021-0000983 · Jan 5, 2021 · national
Continuity (2)
Continuation PCTKR2021012883 · Sep 17, 2021
Related Publication 20230017927A1 · Jan 19, 2023