IP Library Granted Patent US 11,508,378
Granted Patent B2
US 11,508,378 · App. 16/661,658 · Granted Nov 22, 2022

Electronic device and method for controlling the same

Inventors: Kwangyoun Kim (Suwon-si, KR); Kyungmin Lee (Suwon-si, KR); Youngho Han (Suwon-si, KR); Sungsoo Kim (Suwon-si, KR); Sichen Jin (Suwon-si, KR); Jisun Park (Suwon-si, KR); Yeaseul Song (Suwon-si, KR); Jaewon Lee (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L17/00G10L25/51G10L25/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,508,378
App. No.
16/661,658
Granted
Nov 22, 2022
Kind
B2
Abstract

An electronic device is provided. The electronic device includes a microphone to receive audio, a communicator, a memory configured to store computer-executable instructions, and a processor configured to execute the computer-executable instructions. The processor is configured to determine whether the received audio includes a predetermined trigger word; based on determining that the predetermined trigger word is included in the received audio; activate a speech recognition function of the electronic device; detect a movement of a user while the speech recognition function is activated; and based on detecting the movement of the user, transmit a control signal, to a second electronic device to activate a speech recognition function of the second electronic device.

Claims (78)

1. An electronic device comprising:

a microphone to receive audio;

a communicator;

a memory configured to store computer-executable instructions; and

a processor configured to execute the computer-executable instructions to:

determine whether the received audio includes a predetermined trigger word spoken by a user,

based on determining that the predetermined trigger word is included in the received audio, activate a speech recognition function of the electronic device,

detect a movement of the user while the speech recognition function is activated, and

based on detecting the movement of the user, transmit a control signal, to a second electronic device to activate a speech recognition function of the second electronic device,

wherein the processor is further configured to:

compare a signal-to-noise ratio (SNR) of each of a plurality of frames constituting the received audio,

determine a frame among the plurality of frames having the SNR greater than or equal to a predetermined value,

recognize the predetermined trigger word based on the determined frame,

obtain first speech recognition information by performing speech recognition on the received audio,

receive second speech recognition information through the communicator from the second electronic device receiving the control signal,

obtain a final recognition result based on the first speech recognition information and the second speech recognition information,

obtain the final recognition result by applying a language model to the second speech recognition information when the second speech recognition information received from the second electronic device is information indicating that an acoustic model is applied and the language model is not applied, and

obtain the final recognition result by applying the acoustic model and the language model to the second speech recognition information when the second speech recognition information received from the second electronic device is information indicating that the acoustic model and the language model are not applied.

2. The electronic device according to claim 1 , wherein the processor is further configured to detect the movement of the user based on the received audio obtained through the microphone after the speech recognition function is activated.

3. The electronic device according to claim 1 , wherein the memory stores information on a plurality of electronic devices that receive the audio, and

wherein the processor is further configured to:

based on the movement of the user, identify one of the plurality of electronic devices that is closest to the user, and

control the communicator to transmit the control signal to the identified electronic device.

4. The electronic device according to claim 1 , wherein the processor is further configured to:

obtain time information on a time at which the control signal is transmitted to the second electronic device, and

match the first speech recognition information and the second speech recognition information based on the obtained time information to obtain the final recognition result.

5. The electronic device according to claim 4 , wherein the obtained time information includes information on an absolute time at which the control signal is transmitted and information on a relative time at which the control signal is transmitted to the second electronic device based on a time at which the speech recognition function of the electronic device is activated.

6. The electronic device according to claim 1 , wherein the processor is further configured to control the communicator to transmit the control signal, to the second electronic device, for providing a feedback on the final recognition result of the electronic device.

7. The electronic device according to claim 1 , wherein the processor is further configured to activate the speech recognition function of the electronic device when a second control signal for activating the speech recognition function is received from the second electronic device.

8. The electronic device according to claim 7 , wherein the processor is further configured to:

receive user information from the second electronic device, and

identify the received audio corresponding to the user information among a plurality of audios received through the microphone after the speech recognition function is activated by the second control signal.

9. The electronic device according to claim 7 , wherein the processor is further configured to:

obtain speech recognition information by performing speech recognition on the received audio until an utterance of the user ends after the speech recognition function is activated by the second control signal, and

transmit the obtained speech recognition information to the second electronic device.

10. The electronic device according to claim 7 , wherein the processor is further configured to identify a first user and a second user based on the received audio among a plurality of audios.

11. A method for controlling an electronic device, the method comprising:

receiving audio through a microphone of the electronic device;

determining whether the received audio includes a predetermined trigger word spoken by a user;

based on determining that the predetermined trigger word is included in the received audio, activating a speech recognition function of the electronic device;

detecting a movement of the user moves the speech recognition function is activated; and

based on detecting the movement of the user, transmitting a control signal, to a second electronic device to activate a speech recognition function of the second electronic device,

wherein the method further comprises:

comparing a signal-to-noise ratio (SNR) of each of a plurality of frames constituting the received audio,

determining a frame among the plurality of frames having the SNR greater than or equal to a predetermined value,

recognizing the predetermined trigger word based on the determined frame,

obtaining first speech recognition information by performing speech recognition on the received audio,

receiving second speech recognition information through a communicator from the second electronic device receiving the control signal,

obtaining a final recognition result based on the first speech recognition information and the second speech recognition information,

applying a language model to the second speech recognition information when the second speech recognition information received from the second electronic device is information indicating that an acoustic model is applied and the language model is not applied, and

applying the acoustic model and the language model to the second speech recognition information when the second speech recognition information received from the second electronic device is information indicating that the acoustic model and the language model are not applied.

12. The method according to claim 11 , wherein in the detecting the movement of the user is based on the received audio obtained through the microphone after the speech recognition function is activated.

13. The method according to claim 11 , wherein the electronic device stores information on a plurality of electronic devices that receive the audio, and

wherein the method further comprises: based on the movement of the user, identifying one of the plurality of electronic devices that is closest to the user, and

transmitting the control signal to the identified electronic device.

14. The method according to claim 11 , further comprising:

obtaining time information on a time at which the control signal is transmitted to the second electronic device, and

matching the first speech recognition information and the second speech recognition information based on the obtained time information to obtain the final recognition result.

15. The method according to claim 14 , wherein the obtained time information includes information on an absolute time at which the control signal is transmitted and information on a relative time at which the control signal is transmitted to the second electronic device based on a time at which the speech recognition function of the electronic device is activated.

16. An electronic device comprising:

a communicator;

a memory configured to include at least one instruction; and

a processor configured to execute the at least one instruction,

wherein the processor is configured to:

receive a first audio signal of a user speech through the communicator from a first external device,

determine whether the first audio signal of the user speech includes a predetermined trigger word spoken by a user,

based on determining that the first audio signal of the user speech includes the predetermined trigger word, control the communicator to transmit a control signal, to a second external device, for receiving a second audio signal of the user speech from the second external device located in a movement direction of the user when a movement of the user is detected based on information included in the received first audio signal,

receive the second audio signal through the communicator from the second external device, and

match the received first audio signal and the received second audio signal to perform speech recognition on the user speech,

wherein the processor is further configured to:

compare a signal-to-noise ratio (SNR) of each of a plurality of frames constituting the first audio signal,

determine a frame among the plurality of frames having the SNR greater than or equal to a predetermined value,

recognize the predetermined trigger word based on the determined frame,

obtain first speech recognition information by performing speech recognition on the first audio signal,

receive second speech recognition information through the communicator from a second electronic device receiving the control signal,

obtain a final recognition result based on the first speech recognition information and the second speech recognition information,

obtain the final recognition result by applying a language model to the second speech recognition information when the second speech recognition information received from the second electronic device is information indicating that an acoustic model is applied and the language model is not applied, and

obtain the final recognition result by applying the acoustic model and the language model to the second speech recognition information when the second speech recognition information received from the second electronic device is information indicating that the acoustic model and the language model are not applied.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2019
From: KIM, KWANGYOUN; LEE, KYUNGMIN; HAN, YOUNGHO; KIM, SUNGSOO; JIN, SICHEN; PARK, JISUN; SONG, YEASEUL; LEE, JAEWON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 050809/0906 →
Priority Claims (2)
KR 10-2018-0126946 · Oct 23, 2018 · national
KR 10-2019-0030660 · Mar 18, 2019 · national
Continuity (1)
Related Publication 20200126565A1 · Apr 23, 2020
Cited By (3)
US 12,334,067 US 12,335,550 US 12,700,420