IP Library › Granted Patent US 12,367,868
Granted Patent B2
US 12,367,868 · App. 18/085,936 · Granted Jul 22, 2025

Electronic apparatus and controlling method thereof

Inventor: Seongkyu Mun (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L15/16G10L21/0208G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,868
App. No.
18/085,936
Granted
Jul 22, 2025
Kind
B2
Abstract

An electronic device includes a memory storing first vector information obtained from a pre-registered user voice, and a processor configured to obtain, based on a user voice being received, second vector information of a first filtered user voice by inputting the received user voice and the first vector information stored in the memory to a trained first neural network model, obtain second filtered user voice information by inputting the second vector information of the first filtered user voice and the received user voice to a trained second neural network model, and perform voice recognition based on the second filtered user voice information.

Claims (42)

1. An electronic apparatus, comprising:

a memory storing first vector information obtained from a pre-registered user voice; and

a processor configured to:

based on a user voice being received, obtain second vector information of a first filtered user voice by inputting the received user voice and the first vector information stored in the memory to a trained first neural network model,

obtain second filtered user voice information by inputting the second vector information of the first filtered user voice and the received user voice to a trained second neural network model, and

perform voice recognition based on the second filtered user voice information.

2. The electronic apparatus of claim 1 , wherein the trained first neural network model is trained to output third vector information of user voice data included in user voice data for learning, based on the first vector information stored in the memory and the user voice data for learning being input, and

wherein a parameter of the trained first neural network model is updated based on a loss value obtained by comparing the output third vector information on the user voice data and ground truth data.

3. The electronic apparatus of claim 1 , wherein the trained second neural network model is trained to output, based on the first vector information stored in the memory and user voice data for learning being input, information on user voice data included in the user voice data for learning, and

wherein a parameter of the trained second neural network model is updated based on a loss value obtained by comparing the output information on the user voice data and ground truth data.

4. The electronic apparatus of claim 1 , wherein the trained first neural network model is configured to use the first vector information stored in the memory and data for learning as input data, and use third vector information on user voice data included in the data for learning as output data, and

wherein the data for learning comprises the user voice data and noise data.

5. The electronic apparatus of claim 1 , wherein the trained second neural network model is configured to use the first vector information stored in the memory and user voice data for learning as input data, and use information on user voice data included in the user voice data for learning as output data, and

wherein the data for learning comprises the user voice data and noise data.

6. The electronic apparatus of claim 1 , wherein the second vector information of the first filtered user voice comprises third vector information of the user voice in which noise data included in the received user voice is filtered.

7. The electronic apparatus of claim 1 , wherein the processor is further configured to:

obtain frequency information by Fast Fourier Transform (FFT) converting the pre-registered user voice,

obtain the first vector information of the pre-registered user voice based on the frequency information, and

store the obtained first vector information in the memory.

8. The electronic apparatus of claim 1 , wherein the pre-registered user voice comprises a trigger voice for activating a voice recognition mode of the electronic apparatus.

9. A method of an electronic apparatus, the method comprising:

based on a user voice being received, obtaining second vector information of a first filtered user voice by inputting the received user voice and first vector information obtained from a pre-registered user voice to a trained first neural network model;

obtaining second filtered user voice information by inputting the second vector information of the first filtered user voice and the received user voice to a trained second neural network model; and

performing voice recognition based on the second filtered user voice information.

10. The method of claim 9 , wherein the trained first neural network model is trained to output third vector information on user voice data included in user voice data for learning, based on the first vector information and the user voice data for learning being input, and

wherein a parameter of the trained first neural network model is updated based on a loss value obtained by comparing the output third vector information on the user voice data and ground truth data.

11. The method of claim 9 , wherein the trained second neural network model is configured to output, based on the first vector information and user voice data for learning being input, information on user voice data included in the user voice data for learning, and

wherein a parameter of the trained second neural network model is updated based on a loss value obtained by comparing the output information on the user voice data and ground truth data.

12. The method of claim 9 , wherein the trained first neural network model is configured to use the first vector information and data for learning as input data, and use third vector information on user voice data included in the data for learning as output data, and

wherein the data for learning comprises the user voice data and noise data.

13. The method of claim 9 , wherein the trained second neural network model is configured to use the first vector information stored and user voice data for learning as input data, and use information on user voice data included the user voice data for learning as output data, and

wherein the user voice data for learning comprises the user voice data and noise data.

14. The method of claim 9 , wherein the second vector information of the first filtered user voice comprises third vector information of the user voice in which noise data included in the received user voice is filtered.

15. The method of claim 9 , further comprising:

obtaining frequency information by Fast Fourier Transform (FFT) converting the pre-registered user voice;

obtaining the first vector information of the pre-registered user voice based on the frequency information; and

storing the obtained vector information.

16. The method of claim 9 , wherein the pre-registered user voice is a trigger voice for activating a voice recognition mode of the electronic apparatus.

17. A non-transitory computer readable storage medium configured to store a computer instruction to perform an operation by an electronic apparatus based on being executed by a processor of the electronic apparatus, the operation comprising:

obtaining, based on a user voice being received, vector information of a first filtered user voice by inputting the received user voice and vector information obtained from a pre-registered user voice in a trained first neural network model;

obtaining second filtered user voice information by inputting the vector information of first filtered user voice and the received user voice in a trained second neural network model; and

performing voice recognition based on the second filtered user voice information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2022
From: MUN, SEONGKYU
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 062204/0064 →
Priority Claims (1)
KR 10-2021-0184130 · Dec 21, 2021 · national
Continuity (1)
Related Publication 20230197065A1 · Jun 22, 2023
References Cited (14)
US 10511908B1 · Fisher · 2019 [cited by examiner]
US 10720165B2 · Guo et al. · 2020 [cited by applicant]
US 20070055516A1 · Kakino · 2007 [cited by examiner]
US 20220114593A1 · Johnson · 2022 [cited by examiner]
US 20220114594A1 · Nunes · 2022 [cited by examiner]
US 20230197065A1 · Mun · 2023 [cited by examiner]
CN 110047490A · 2019 [cited by examiner]
KR 102002903B1 · 2019 [cited by applicant]
KR 1020220053456A · 2022 [cited by applicant]
KR 1020220169242A · 2022 [cited by applicant]
WO 2019145708A1 · 2019 [cited by applicant]
WO 2021071489A1 · 2021 [cited by applicant]
S. Jain, P. Jha and R. Suresh, “Design and implementation of an Automatic Speaker recognition system using neural and fuzzy logic in Matlab,” 2013 International Conference on Signal Processing and Communications on Sign… [cited by examiner]
S. Jain, P. Jna and R. Suresh, “Design and implementation of an Automatic Speaker recognition system using neural and fuzzy logic in Matlab,” 2013 International Conference on Signal Processing and Communication (ICSC), … [cited by examiner]