IP Library › Granted Patent US 12,536,993
Granted Patent B2
US 12,536,993 · App. 18/328,343 · Granted Jan 27, 2026

System and method for post-ASR false wake-up suppression

Inventors: Tapas Kanungo (Redmond, WA); Preeti Saraswat (Santa Clara, CA); Stephen Michael Walsh (Sunnyvale, CA)
Assignee: Samsung Electronics Co., Ltd.
G10L15/08G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,993
App. No.
18/328,343
Granted
Jan 27, 2026
Kind
B2
Abstract

A method includes obtaining a speech signal. The method also includes predicting a first likelihood of a wake word or phrase being spoken in the speech signal using a first machine learning model trained to receive the speech signal as input. The method further includes, responsive to the first likelihood exceeding a first threshold, performing automatic speech recognition on the speech signal to determine a textual representation of the speech signal. The method also includes predicting a second likelihood of the wake word or phrase being spoken in the speech signal using a second machine learning model trained to receive at least one of the textual representation, audio features associated with the speech signal, and context features associated with the electronic device. In addition, the method includes, responsive to the second likelihood exceeding a second threshold, generating instructions to perform an action requested in the speech signal.

Claims (70)

1 . A method comprising:

obtaining, by at least one processing device of an electronic device, a speech signal;

predicting, by the at least one processing device, a first likelihood of a wake word or phrase being spoken in the speech signal using a first machine learning model trained to receive the speech signal as input;

responsive to the first likelihood exceeding a first threshold, performing, by the at least one processing device, automatic speech recognition on the speech signal to determine a textual representation of the speech signal;

predicting, by the at least one processing device, a second likelihood of the wake word or phrase being spoken in the speech signal using a second machine learning model trained to receive at least one of the textual representation, audio features associated with the speech signal, and context features associated with the electronic device;

responsive to the second likelihood exceeding a second threshold, generating, by the at least one processing device, instructions to perform an action requested in the speech signal;

determining, by the at least one processing device, an accuracy of the first machine learning model used to predict the first likelihood using a plurality of previous speech signals; and

updating, by the at least one processing device, the second threshold based on the determined accuracy.

2 . The method of claim 1 , wherein predicting the second likelihood using the second machine learning model includes:

performing, by the at least one processing device, natural language processing on the textual representation using a deep transformer model included in the second machine learning model and outputting one or more results;

providing, by the at least one processing device to a multi-layer perceptron model of the second machine learning model, the one or more results of the natural language processing, the audio features, and the context features; and

receiving, by the at least one processing device, an output from the multi-layer perceptron model indicating whether or not to suppress a wake-up operation of a voice assistant.

3 . The method of claim 1 , wherein predicting the second likelihood using the second machine learning model includes:

performing, by the at least one processing device, natural language processing on the textual representation using a deep transformer model included in the second machine learning model and outputting one or more results;

providing, by the at least one processing device to a random forest classifier model of the second machine learning model, the one or more results of the natural language processing, the audio features, and the context features; and

receiving, by the at least one processing device, an output from the random forest classifier model indicating whether or not to suppress a wake-up operation of a voice assistant.

4 . The method of claim 1 , further comprising:

determining, by the at least one processing device, a background noise environment associated with the speech signal; and

setting the second threshold based on the determined background noise environment.

5 . The method of claim 1 , wherein the audio features include at least one of: a bag of words, a total word count, a total character count, a unique word count, a stop word count, an audio time, and a signal-to-noise ratio.

6 . The method of claim 1 , wherein the context features include at least one of: at least one user characteristic, a launch method, and foreground application information.

7 . The method of claim 4 , wherein:

determining the background noise environment associated with the speech signal includes determining that the background noise environment is associated with an environment in which use of a voice assistant is likely; and

setting the second threshold based on the determined background noise environment includes decreasing the second threshold.

8 . An electronic device comprising:

at least one processing device configured to:

obtain a speech signal;

predict a first likelihood of a wake word or phrase being spoken in the speech signal using a first machine learning model trained to receive the speech signal as input;

responsive to the first likelihood exceeding a first threshold, perform automatic speech recognition on the speech signal to determine a textual representation of the speech signal;

predict a second likelihood of the wake word or phrase being spoken in the speech signal using a second machine learning model trained to receive at least one of the textual representation, audio features associated with the speech signal, and context features associated with the electronic device;

responsive to the second likelihood exceeding a second threshold, generate instructions to perform an action requested in the speech signal;

determine an accuracy of the first machine learning model used to predict the first likelihood using a plurality of previous speech signals; and

update the second threshold based on the determined accuracy.

9 . The electronic device of claim 8 , wherein, to predict the second likelihood using the second machine learning model, the at least one processing device is configured to:

perform natural language processing on the textual representation using a deep transformer model included in the second machine learning model and outputting one or more results;

provide, to a multi-layer perceptron model of the second machine learning model, the one or more results of the natural language processing, the audio features, and the context features; and

receive an output from the multi-layer perceptron model indicating whether or not to suppress a wake-up operation of a voice assistant.

10 . The electronic device of claim 8 , wherein, to predict the second likelihood using the second machine learning model, the at least one processing device is configured to:

perform natural language processing on the textual representation using a deep transformer model included in the second machine learning model and outputting one or more results;

provide, to a random forest classifier model of the second machine learning model, the one or more results of the natural language processing, the audio features, and the context features; and

receive an output from the random forest classifier model indicating whether or not to suppress a wake-up operation of a voice assistant.

11 . The electronic device of claim 8 , wherein the at least one processing device is further configured to:

determine a background noise environment associated with the speech signal; and

set the second threshold based on the determined background noise environment.

12 . The electronic device of claim 8 , wherein the audio features include at least one of: a bag of words, a total word count, a total character count, a unique word count, a stop word count, an audio time, and a signal-to-noise ratio.

13 . The electronic device of claim 8 , wherein the context features include at least one of: at least one user characteristic, a launch method, and foreground application information.

14 . The electronic device of claim 11 , wherein;

to determine the background noise environment associated with the speech signal, the at least one processing device is configured to determine that the background noise environment is associated with an environment in which use of a voice assistant is likely; and

to set the second threshold based on the determined background noise environment, the at least one processing device is configured to decrease the second threshold.

15 . A non-transitory machine readable medium containing instructions that when executed cause at least one processor of an electronic device to:

obtain a speech signal;

predict a first likelihood of a wake word or phrase being spoken in the speech signal using a first machine learning model trained to receive the speech signal as input;

responsive to the first likelihood exceeding a first threshold, perform automatic speech recognition on the speech signal to determine a textual representation of the speech signal;

predict a second likelihood of the wake word or phrase being spoken in the speech signal using a second machine learning model trained to receive at least one of the textual representation, audio features associated with the speech signal, and context features associated with the electronic device;

responsive to the second likelihood exceeding a second threshold, generate instructions to perform an action requested in the speech signal;

determine an accuracy of the first machine learning model used to predict the first likelihood using a plurality of previous speech signals; and

update the second threshold based on the determined accuracy.

16 . The non-transitory machine readable medium of claim 15 , wherein the instructions that when executed cause the at least one processor to predict the second likelihood using the second machine learning model comprise instructions that when executed cause the at least one processor to:

perform natural language processing on the textual representation using a deep transformer model included in the second machine learning model and outputting one or more results;

provide, to a multi-layer perceptron model of the second machine learning model, the one or more results of the natural language processing, the audio features, and the context features; and

receive an output from the multi-layer perceptron model indicating whether or not to suppress a wake-up operation of a voice assistant.

17 . The non-transitory machine readable medium of claim 15 , wherein the instructions that when executed cause the at least one processor to predict the second likelihood using the second machine learning model comprise instructions that when executed cause the at least one processor to:

perform natural language processing on the textual representation using a deep transformer model included in the second machine learning model and outputting one or more results;

provide, to a random forest classifier model of the second machine learning model, the one or more results of the natural language processing, the audio features, and the context features; and

receive an output from the random forest classifier model indicating whether or not to suppress a wake-up operation of a voice assistant.

18 . The non-transitory machine readable medium of claim 15 , further containing instructions that when executed cause the at least one processor to:

determine a background noise environment associated with the speech signal; and

set the second threshold based on the determined background noise environment.

19 . The non-transitory machine readable medium of claim 15 , wherein the audio features include at least one of: a bag of words, a total word count, a total character count, a unique word count, a stop word count, an audio time, and a signal-to-noise ratio.

20 . The non-transitory machine readable medium of claim 15 , wherein the context features include at least one of: at least one user characteristic, a launch method, and foreground application information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2023
From: KANUNGO, TAPAS; SARASWAT, PREETI; WALSH, STEPHEN MICHAEL
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 063842/0668 →
Continuity (2)
Provisional Application 63414648 · Oct 10, 2022
Related Publication 20240119925A1 · Apr 11, 2024
References Cited (38)
US 9418656B2 · Foerster · 2016 [cited by examiner]
US 9697828B1 · Prasad · 2017 [cited by examiner]
US 10510340B1 · Fu · 2019 [cited by examiner]
US 11195522B1 · Makashir et al. · 2021 [cited by applicant]
US 11232788B2 · Yavagal · 2022 [cited by examiner]
US 11361763B1 · Maas · 2022 [cited by examiner]
US 11557293B2 · Carbune · 2023 [cited by examiner]
US 11577379B2 · Kim · 2023 [cited by applicant]
US 11721338B2 · Kwatra · 2023 [cited by examiner]
US 20180012593A1 · Prasad et al. · 2018 [cited by applicant]
US 20180233150A1 · Gruenstein · 2018 [cited by examiner]
US 20190251963A1 · Li et al. · 2019 [cited by applicant]
US 20190371325A1 · Nakada · 2019 [cited by examiner]
US 20200105256A1 · Fainberg · 2020 [cited by examiner]
US 20200168207A1 · Wang et al. · 2020 [cited by applicant]
US 20210104221A1 · Sharifi et al. · 2021 [cited by applicant]
US 20210249005A1 · Bromand · 2021 [cited by examiner]
US 20210256965A1 · Kim et al. · 2021 [cited by applicant]
US 20220020357A1 · Rastrow et al. · 2022 [cited by applicant]
US 20220115015A1 · Elkhatib · 2022 [cited by examiner]
US 20230245648A1 · Thyssen · 2023 [cited by examiner]
US 20230298578A1 · Delaney · 2023 [cited by examiner]
US 20240029736A1 · Xiao · 2024 [cited by examiner]
US 20240265921A1 · Dureau · 2024 [cited by examiner]
CN 112002317A · 2020 [cited by applicant]
CN 114360522B · 2022 [cited by applicant]
CN 113450771B · 2022 [cited by applicant]
CN 115064160A · 2022 [cited by applicant]
CN 115132195A · 2022 [cited by applicant]
KR 1020190090424A · 2019 [cited by applicant]
KR 1020200013152A · 2020 [cited by applicant]
KR 1020200025226A · 2020 [cited by applicant]
KR 1020210010270A · 2021 [cited by applicant]
Devlin, Jacob, et al. “Bert: Pre-training of deep bidirectional transformers for language understanding.” Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics… [cited by examiner]
Ho, Tin Kam. “Random decision forests.” Proceedings of 3rd international conference on document analysis and recognition. vol. 1. IEEE, 1995. (Year: 1995). [cited by examiner]
International Search Report and Written Opinion of the International Searching Authority dated Sep. 20, 2023 in connection with International Patent Application No. PCT/KR2023/008934, 12 pages. [cited by applicant]
Supplementary European Search Report dated Aug. 19, 2025, in connection with European Application No. 23877426.9, 9 pages. [cited by applicant]
Gruenstein, et al., “A Cascade Architecture for Keyword Spotting on Moble Devices,” arXiv:1712.03603v1 [cs.SD], Dec. 2017, 4 pages. [cited by applicant]