IP Library Granted Patent US 11,468,892
Granted Patent B2
US 11,468,892 · App. 17/004,474 · Granted Oct 11, 2022

Electronic apparatus and method for controlling electronic apparatus

Inventors: Hyeontaek Lim (Suwon-si, KR); Sejin Kwak (Suwon-si, KR); Youngjin Kim (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G10L15/22G06F3/167G06V20/10G06V40/10G10L13/00G10L15/183G10L15/1815G10L15/24G10L21/0232G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,468,892
App. No.
17/004,474
Granted
Oct 11, 2022
Kind
B2
Abstract

An electronic apparatus and a control method thereof are provided. The electronic apparatus includes a microphone, a camera, a memory storing an instruction, and a processor configured to control the electronic apparatus coupled with the microphone, the camera and the memory, and the processor is configured to, by executing the instruction, obtain a user image by photographing a user through the camera, obtain the user information based on the user image, and based on a user speech being input from the user through the microphone, recognize the user speech by using a speech recognition model corresponding to the user information among a plurality of speech recognition models.

Claims (53)

1. An electronic apparatus, comprising:

a microphone;

a camera;

a memory storing an instruction and a plurality of wake-up models; and

a processor configured to control the electronic apparatus and being coupled with the microphone, the camera and the memory,

wherein the processor, by executing the instruction, is further configured to:

obtain a user image by photographing a user through the camera,

obtain user information based on the user image,

identify a language model and an acoustic model corresponding to the user information,

identify a natural language understanding model of a plurality of natural language understanding models based on the user information,

obtain environment information based on information on an object other than the user comprised in the user image by analyzing the user image,

based on a user speech being input from the user through the microphone, recognize the user speech by using a speech recognition model, among a plurality of speech recognition models, corresponding to the user information,

obtain text data on the user speech by using the identified language model and the acoustic model,

perform a natural language understanding on the obtained text data through the identified natural language understanding model,

perform preprocessing on a speech signal comprising the user speech based on the environment information and the user information,

obtain response information on the user speech based on a result of the natural language understanding,

output a response speech on the user by inputting the response information to a text to speech (TTS) model corresponding to the user speech of a plurality of TTS models, and

identify an output method on the response information based on the user information,

wherein the plurality of wake-up models are neural network models capable of recognizing wake-up words, and trained according to user characteristics included in the user information,

wherein the processor, by executing the instruction, is further configured to identify a wake-up word in the user speech based on a wake-up model, of the plurality of wake-up models, corresponding to the user information,

wherein each of the plurality of speech recognition models comprise the language model and the acoustic model, and

wherein the output method on the response information comprises a method of outputting the response information through a display and a method of outputting the response information through a speaker.

2. The electronic apparatus of claim 1 , wherein the processor, by executing the instruction, is further configured to:

perform preprocessing by identifying a preprocessing filter for removing noise comprised in the speech signal based on the environment information, and

perform preprocessing by identifying a para meter for enhancing the user speech comprised in the speech signal based on the user information.

3. The electronic apparatus of claim 1 , wherein the processor, by executing the instruction, is further configured to obtain the user information by inputting the obtained user image to an object recognition model trained to obtain information on an object comprised in the obtained user image.

4. The electronic apparatus of claim 1 ,

wherein the memory stores a user image of a registered user matched with registered user information, and

wherein the processor, by executing the instruction, is further configured to obtain the user information by comparing the obtained user image with the user image of the registered user.

5. A control method of an electronic apparatus, the method comprising:

obtaining a user image by photographing a user through a camera;

obtaining user information based on the user image;

identifying a language model and an acoustic model corresponding to the user information;

identifying a natural language understanding model of a plurality of natural language understanding models based on the user information;

obtaining environment information based on information on an object other than the user comprised in the user image by analyzing the user image;

based on a user speech being input from the user through a microphone, recognizing the user speech by using a speech recognition model, among a plurality of speech recognition models, corresponding to the user information;

obtaining text data on the user speech by using the identified language model and the acoustic model;

performing a natural language understanding on the obtained text data through the identified natural language understanding model;

performing preprocessing on a speech signal comprising the user speech based on the environment information and the user information;

obtaining response information on the user speech based on a result of the natural language understanding;

outputting a response speech on the user by inputting the response information to a text to speech (TTS) model corresponding to the user speech of a plurality of TTS models; and

identifying an output method on the response information based on the user information,

wherein the electronic apparatus comprises a plurality of wake-up models, the plurality of wake-up models being neural network models capable of recognizing wake-up words, and being trained according to user characteristics included in the user information,

wherein the method further comprises identifying a wake-up word in the user speech based on a wake-up model, of the plurality of wake-up models, corresponding to the user information,

wherein each of the plurality of speech recognition models comprise the language model and the acoustic model, and

wherein the output method on the response information comprises a method of outputting the response information through a display and a method of outputting the response information through a speaker.

6. The control method of claim 5 , wherein the performing of the preprocessing comprises:

performing preprocessing by identifying a preprocessing filter for removing noise comprised in the speech signal based on the environment information; and

performing preprocessing by identifying a parameter for enhancing the user speech comprised in the speech signal based on the user information.

7. The control method of claim 5 , wherein the obtaining of the user information comprises obtaining the user information by inputting the obtained user image to an object recognition model trained to obtain information on an object comprised in the obtained user image.

8. The control method of claim 5 ,

wherein the electronic apparatus matches and stores a user image of a registered user with registered user information, and

wherein the obtaining the user information comprises obtaining the user information by comparing the obtained user image with the user image of the registered user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2020
From: LIM, HYEONTAEK; KWAK, SEJIN; KIM, YOUNGJIN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 053616/0021 →
Priority Claims (1)
KR 10-2019-0125172 · Oct 10, 2019 · national
Continuity (1)
Related Publication 20210110821A1 · Apr 15, 2021
Cited By (1)
US 12,373,027