IP Library › Granted Patent US 11,967,322
Granted Patent B2
US 11,967,322 · App. 17/569,994 · Granted Apr 23, 2024

Server for identifying false wakeup and method for controlling the same

Inventors: Sunok Kim (Suwon-si, KR); Sunbeom Kwon (Suwon-si, KR); Soonhee Jo (Suwon-si, KR); Kiwan Eom (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G10L15/22G10L15/083G10L15/30G10L21/028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,967,322
App. No.
17/569,994
Granted
Apr 23, 2024
Kind
B2
Abstract

A server is provided. The server includes a communication circuitry, and at least one processor operatively connected with the communication circuitry. The at least one processor may be configured to, in response to traffic of a plurality of speeches to wake up a voice assistant feature, received within a preset period being a preset value or more, generate a plurality of clusters based on similarities between the plurality of speeches, and determine whether to respond to each of speeches included in each of the plurality of clusters based on similarities between the speeches included in each of the plurality of clusters.

Claims (44)

1. A server comprising:

a communication circuitry;

one or more processors operatively connected with the communication circuitry; and

memory storing one or more computer programs including computer-executable instructions that, when executed by the one or more processors, cause the server to:

in response to traffic of a plurality of speeches to wake up a voice assistant feature, received within a preset period being a preset value or more, generate a plurality of clusters based on similarities between the plurality of speeches, and

determine whether to respond to each of speeches included in each of the plurality of clusters based on similarities between the speeches included in each of the plurality of clusters,

wherein the one or more computer programs further comprise computer-executable instructions that, when executed by the one or more processors, cause the server to, as at least part of determining whether to respond:

in response to similarities between a plurality of speeches included in a first cluster among the plurality of clusters falling within a first range not less than a preset value, determine the plurality of speeches included in the first cluster as false wakeup speeches or sounds, and

terminate a process without responding to the plurality of speeches included in the first cluster.

2. The server of claim 1 , wherein the one or more computer programs further comprise computer-executable instructions that, when executed by the one or more processors, cause the server to:

identify respective feature points of the plurality of speeches, and

generate the plurality of clusters based on similarities between the respective feature points of the plurality of speeches.

3. The server of claim 1 , wherein the one or more computer programs further comprise computer-executable instructions that, when executed by the one or more processors, cause the server to:

in response to at least one speech whose length is a preset value or more being included in the plurality of speeches, obtain respective parts of the at least one speech, and

generate the plurality of clusters based on similarities between the respective parts of the at least one speech, for the at least one speech.

4. The server of claim 1 , wherein the one or more computer programs further comprise computer-executable instructions that, when executed by the one or more processors, cause the server to transfer the plurality of speeches for waking up the voice assistant feature to an automatic speech recognition (ASR) circuitry in response to the traffic being less than the preset value.

5. The server of claim 1 , wherein the one or more computer programs further comprise computer-executable instructions that, when executed by the one or more processors, cause the server to transmit a command to increase a threshold for wakeup recognition during a preset time to a plurality of terminal devices individually corresponding to the plurality of speeches included in the first cluster.

6. The server of claim 5 , wherein the one or more computer programs further comprise computer-executable instructions that, when executed by the one or more processors, cause the server to, in response to a false wakeup speech or sound being received from at least one terminal device among the plurality of terminal devices within the preset time, transmit a command to extend the preset time to the at least one terminal device.

7. The server of claim 1 , wherein the one or more computer programs further comprise computer-executable instructions that, when executed by the one or more processors, cause the server to transmit response data for identifying whether to wake up to a plurality of terminal devices individually corresponding to the plurality of speeches included in the first cluster in response to the similarities between the plurality of speeches included in the first cluster falling within a second range lower than the first range.

8. The server of claim 7 , wherein the one or more computer programs further comprise computer-executable instructions that, when executed by the one or more processors, cause the server to transfer the plurality of speeches to an automatic speech recognition (ASR) circuitry in response to the similarities between the plurality of speeches included in the first cluster falling within a third range lower than the second range.

9. The server of claim 1 , wherein the one or more computer programs further comprise computer-executable instructions that, when executed by the one or more processors, cause the server to store feature points for the plurality of speeches included in the first cluster, as feature points of false wakeup speeches or sounds, in the memory.

10. A method for controlling a server, the method comprising:

in response to traffic of a plurality of speeches to wake up a voice assistant feature, received within a preset period being a preset value or more, generating a plurality of clusters based on similarities between the plurality of speeches; and

determining whether to respond to each of speeches included in each of the plurality of clusters based on similarities between the speeches included in each of the plurality of clusters,

wherein the determining of whether to respond comprises:

in response to similarities between a plurality of speeches included in a first cluster among the plurality of clusters falling within a first range not less than a preset value, determining the plurality of speeches included in the first cluster as false wakeup speeches or sounds; and

terminating a process without responding to the plurality of speeches included in the first cluster.

11. The method of claim 10 , wherein the generating of the plurality of clusters comprises identifying respective feature points of the plurality of speeches and generating the plurality of clusters based on similarities between the respective feature points of the plurality of speeches.

12. The method of claim 10 , wherein the generating of the plurality of clusters comprises:

in response to at least one speech whose length is a preset value or more being included in the plurality of speeches, obtaining respective parts of the at least one speech; and

generating the plurality of clusters based on similarities between the respective parts of the at least one speech, for the at least one speech.

13. The method of claim 10 , further comprising storing feature points of the plurality of speeches included in the first cluster, as feature points of false wakeup speeches or sounds.

14. The method of claim 10 , further comprising transferring the plurality of speeches for waking up the voice assistant feature to an automatic speech recognition (ASR) circuitry in response to the traffic being less than the preset value.

15. The method of claim 10 , further comprising transmitting a command to increase a threshold for wakeup recognition during a preset time to a plurality of terminal devices individually corresponding to the plurality of speeches included in the first cluster.

16. The method of claim 15 , further comprising, in response to a false wakeup speech or sound being received from at least one terminal device among the plurality of terminal devices within the preset time, transmitting a command to extend the preset time to the at least one terminal device.

17. The method of claim 10 , wherein the determining of whether to respond comprises transmitting response data identifying whether to wake up to a plurality of terminal devices individually corresponding to the plurality of speeches included in the first cluster in response to the similarities between the plurality of speeches included in the first cluster falling within a second range lower than the first range.

18. The method of claim 17 , wherein the determining of whether to respond comprises transferring the plurality of speeches to an automatic speech recognition (ASR) circuitry in response to the similarities between the plurality of speeches included in the first cluster falling within a third range lower than the second range.

19. One or more non-transitory computer-readable storage media storing computer-executable instructions that, when executed by one or more processors of a server, cause the server to perform operations, the operations comprising:

in response to traffic of a plurality of speeches to wake up a voice assistant feature, received within a preset period being a preset value or more, generating a plurality of clusters based on similarities between the plurality of speeches; and

determining whether to respond to each of speeches included in each of the plurality of clusters based on similarities between the speeches included in each of the plurality of clusters,

wherein the determining of whether to respond comprises:

in response to similarities between a plurality of speeches included in a first cluster among the plurality of clusters falling within a first range not less than a preset value, determining the plurality of speeches included in the first cluster as false wakeup speeches or sounds; and

terminating a process without responding to the plurality of speeches included in the first cluster.

20. The one or more non-transitory computer-readable storage media of claim 19 , wherein the generating of the plurality of clusters comprises identifying respective feature points of the plurality of speeches and generating the plurality of clusters based on similarities between the respective feature points of the plurality of speeches.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2022
From: KIM, SUNOK; KWON, SUNBEOM; JO, SOONHEE; EOM, KIWAN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 058575/0685 →
Priority Claims (1)
KR 10-2021-0058837 · May 6, 2021 · national
Continuity (2)
Continuation PCTKR2021019212 · Dec 16, 2021
Related Publication 20220358918A1 · Nov 10, 2022