IP Library Granted Patent US 11,087,755
Granted Patent B2
US 11,087,755 · App. 16/327,646 · Granted Aug 10, 2021

Electronic device for voice recognition, and control method therefor

Inventor: Myung-suk Song (Seoul, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L15/22G06F3/167G10L15/20G10L21/02G10L25/87G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,087,755
App. No.
16/327,646
Granted
Aug 10, 2021
Kind
B2
Abstract

An electronic device and a control method therefor are provided. The electronic device comprises: a plurality of microphones for receiving audio signals in which voice signals are included; a communicator for receiving state information according to information about a connection with the electronic device; and a processor for determining the noise environment around the electronic device based on one or more of audio amplitude information that is received by the plurality of microphones and that is output by an external device, and state information of the external device, and for processing the voice signal based on the determined noise environment so as to perform voice recognition.

Claims (39)

1. An electronic device comprising:

a plurality of microphones;

a communicator; and

a processor configured to:

identify a noise environment around the electronic device based on at least one from among amplitude information of an audio signal output by an external device and received by the plurality of microphones and state information of the external device received through the communicator, the audio signal comprising a voice signal; and

perform a function corresponding to the voice signal based on the identified noise environment,

wherein the processor is further configured to:

based on an amplitude of the audio signal that is output by the external device being greater than or equal to a predetermined value corresponding to the state information, identify the noise environment around the electronic device as a first state, and

based on the amplitude of the audio signal that is output by the external device being less than the predetermined value or the external device not being connected, identify the noise environment around the electronic device as a second state.

2. The electronic device of claim 1 , further comprising:

a memory,

wherein the processor is further configured to:

store initialized state information of the external device at a time when the external device is initially connected in the memory, and

based on the external device being connected to the electronic device, update and identify the noise environment around the electronic device on a basis of pre-stored state information.

3. The electronic device of claim 1 , wherein the processor is further configured to:

detect a voice signal section composed of a plurality of successive frames based on the received state information of the external device from the audio signal and a noise environment around the electronic device, identify, from the detected voice section, an input direction of the audio signal based on the received state information of the external device and the noise environment around the electronic device, perform beamforming based on the received state information of the external device and the noise environment around the electronic device from the input direction information of the audio signal, and process the voice signal.

4. The electronic device of claim 3 , wherein the processor is further configured to:

in the first state, set a length of hang-over which identifies subsequent frames after the detected voice section as a voice to a first length, and in the second state, set the length of hang-over as a second length which is longer than the first length, and detect the voice section.

5. The electronic device of claim 3 , wherein the processor is further configured to

detect the voice section by, in the first state, applying a first high weighted value to a frame which is identified as a section that is not considered the voice section from the audio signal, and in the second state, applying a second high weighted value to a frame which is identified as the voice section from the audio signal.

6. The electronic device of claim 3 , wherein the processor is further configured to

identify an input direction of the audio signal by setting, in the first state, an input angle search range in a direction in which the audio signal is inputtable to a first range that is a left direction and a right direction generated in a previous frame of the detected voice section, and setting, in the second state, the input angle search range in a direction in which the audio signal is inputtable to a second range that is a left direction and a right direction generated in a previous frame of the detected voice section, the second range being wider than the first range.

7. The electronic device of claim 3 , wherein the processor is further configured to:

in the first state, fix a target direction to estimate the input direction of the audio signal, wherein an audio signal of a frame that is received after the detected voice section is amplified with respect to the fixed target direction, and in the second state, set the target direction to estimate the input direction of the audio signal to all directions and identify the input direction of the audio signal in all input angle ranges.

8. A method for processing a voice signal by an electronic device comprising a plurality of microphones, the method comprising:

identifying a noise environment around the electronic device based on at least one from among amplitude information of an audio signal output by an external device and received by the plurality of microphones and state information of the external device received from the external device, the audio signal comprising a voice signal; and

performing a function corresponding to the voice signal based on the identified noise environment,

wherein the identifying the noise environment further comprises:

based on an amplitude of the audio signal that is output by the external device being greater than or equal to a predetermined value corresponding to the state information, identifying the noise environment around the electronic device as a first state, and

based on the amplitude of the audio signal that is output by the external device being less than the predetermined value or the external device not being connected, identifying the noise environment around the electronic device as a second state.

9. The method of claim 8 , wherein the receiving the state information further comprises:

storing initialized state information of an external device at a time when the external device is initially connected in the memory; and

based on the external device being connected to the electronic device, updating and identifying the noise environment around the electronic device on a basis of pre-stored state information.

10. The method of claim 8 , wherein the processing the voice signal further comprises:

detecting a voice signal section composed of a plurality of successive frames based on the received state information of the external device from the audio signal and the noise environment around the electronic device; identifying, from the detected voice section, an input direction of the audio signal based on the received state information of the external device and the noise environment around the electronic device; and

perform beamforming based on the received state information of the external device and the noise environment around the electronic device from the input direction information of the audio signal, and processing the voice signal.

11. The method of claim 10 , wherein the detecting the voice section comprises, in the mode state, setting a length of hang-over which identifies subsequent frames after the detected voice section as a voice to a first length, and in the second state, setting the length of hang-over as a second length which is longer than the first length, and detecting the voice section.

12. The method of claim 10 , wherein the detecting the voice section comprises

detecting voice section by, in the first state, applying a first high weighted value to a frame which is identified as a section that is not considered the voice section from the audio signal, and in the second state, applying a second high weighted value to a frame which is identified as the voice section from the audio signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2019
From: SONG, MYUNG-SUK
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 048413/0936 →
Priority Claims (1)
KR 10-2016-0109481 · Aug 26, 2016 · national
Continuity (1)
Related Publication 20190221210A1 · Jul 18, 2019