IP Library Granted Patent US 12,626,697
Granted Patent B2
US 12,626,697 · App. 18/352,601 · Granted May 12, 2026

System and method for keyword false alarm reduction

Inventors: Rakshith Sharma Srinivasa (Sunnyvale, CA); Yashas Malur Saidutta (Menlo Park, CA); Ching-Hua Lee (Mountain View, CA); Chou-Chang Yang (San Jose, CA); Yilin Shen (San Jose, CA); Hongxia Jin (San Jose, CA)
Assignee: Samsung Electronics Co., Ltd.
G10L15/22G10L15/02G10L15/063G10L15/18G10L25/78G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,697
App. No.
18/352,601
Granted
May 12, 2026
Kind
B2
Abstract

A method includes extracting, using a keyword detection model, audio features from audio data. The method also includes processing the audio features by a first layer of the keyword detection model configured to predict a first likelihood that the audio data includes speech. The method also includes processing the audio features by a second layer of the keyword detection model configured to predict a second likelihood that the audio data includes keyword-like speech. The method also includes processing the audio features by a third layer of the keyword detection model configured to predict a third likelihood, for each of a plurality of possible keywords, that the audio data includes the keyword. The method also includes identifying a keyword included in the audio data. The method also includes generating instructions to perform an action based at least in part on the identified keyword.

Claims (62)

1 . A method comprising:

obtaining audio data from an audio input device;

providing the audio data as input to a keyword detection model;

extracting, using the keyword detection model, audio features from the audio data;

processing the audio features by a first layer of the keyword detection model configured to predict a first likelihood that the audio data includes speech;

processing the audio features by a second layer of the keyword detection model configured to predict a second likelihood that the audio data includes keyword-like speech;

processing the audio features by a third layer of the keyword detection model configured to predict a third likelihood, for each of a plurality of possible keywords, that the audio data includes the keyword;

providing each of the first likelihood, the second likelihood, and the third likelihood to an inference combination layer of the keyword detection model;

identifying a keyword included in the audio data based on a combination of the first likelihood, the second likelihood, and the third likelihood by the inference combination layer; and

generating instructions to perform an action based at least in part on the identified keyword.

2 . The method of claim 1 , further comprising:

processing the audio features by the second layer of the keyword detection model in response to the first likelihood exceeding a first threshold; and

processing the audio features by the third layer of the keyword detection model in response to the second likelihood exceeding a second threshold,

wherein identifying the keyword includes identifying which one of the plurality of possible keywords is associated with a highest third likelihood.

3 . The method of claim 1 , wherein identifying the keyword includes determining, based on the first likelihood, the second likelihood, and each of the third likelihood, that the audio data includes the keyword.

4 . The method of claim 1 , wherein the audio features are extracted from the audio data using a backbone model of the keyword detection model.

5 . The method of claim 1 , wherein the keyword detection model is trained using a training dataset that includes a first set of audio data samples including non-speech audio, a second set of audio data samples including non-keyword speech, and a third set of audio data samples including a keyword, and wherein each audio data sample is annotated with a speech label indicating whether the audio data sample includes speech, a keyword-like label indicating whether the audio data sample includes keyword-like speech, and a keyword label identifying which keyword, if any, is in the audio data sample.

6 . The method of claim 5 , wherein the first layer of the keyword detection model is trained using the first set of audio data samples, the second set of audio data samples, and the third set of audio data samples to distinguish between speech and non-speech audio, wherein the audio data samples including non-keyword speech and the audio data samples including a keyword are pooled into a speech class.

7 . The method of claim 5 , wherein the second layer of the keyword detection model is trained using the second set of audio data samples and the third set of audio data samples to distinguish between non-keyword and keyword-like speech, wherein the audio data samples including keyword-like speech are pooled into a keyword-like class.

8 . The method of claim 5 , wherein the third layer of the keyword detection model is trained using the third set of audio data samples.

9 . An electronic device comprising:

at least one processing device configured to:

obtain audio data from an audio input device;

provide the audio data as input to a keyword detection model;

extract, using the keyword detection model, audio features from the audio data;

process the audio features by a first layer of the keyword detection model configured to predict a first likelihood that the audio data includes speech;

process the audio features by a second layer of the keyword detection model configured to predict a second likelihood that the audio data includes keyword-like speech;

process the audio features by a third layer of the keyword detection model configured to predict a third likelihood, for each of a plurality of possible keywords, that the audio data includes the keyword;

provide each of the first likelihood, the second likelihood, and the third likelihood to an inference combination layer of the keyword detection model;

identify a keyword included in the audio data based on a combination of the first likelihood, the second likelihood, and the third likelihood by the inference combination layer; and

generate instructions to perform an action based at least in part on the identified keyword.

10 . The electronic device of claim 9 , wherein the at least one processing device is further configured to:

process the audio features by the second layer of the keyword detection model in response to the first likelihood exceeding a first threshold; and

process the audio features by the third layer of the keyword detection model in response to the second likelihood exceeding a second threshold,

wherein, to identify the keyword, the at least one processing device is further configured to identify which one of the plurality of possible keywords is associated with a highest third likelihood.

11 . The electronic device of claim 9 , wherein, to identify the keyword, the at least one processing device is further configured to determine, based on the first likelihood, the second likelihood, and each of the third likelihood, that the audio data includes the keyword.

12 . The electronic device of claim 9 , wherein the audio features are extracted from the audio data using a backbone model of the keyword detection model.

13 . The electronic device of claim 9 , wherein the keyword detection model is trained using a training dataset that includes a first set of audio data samples including non-speech audio, a second set of audio data samples including non-keyword speech, and a third set of audio data samples including a keyword, and wherein each audio data sample is annotated with a speech label indicating whether the audio data sample includes speech, a keyword-like label indicating whether the audio data sample includes keyword-like speech, and a keyword label identifying which keyword, if any, is in the audio data sample.

14 . The electronic device of claim 13 , wherein the first layer of the keyword detection model is trained using the first set of audio data samples, the second set of audio data samples, and the third set of audio data samples to distinguish between speech and non-speech audio, wherein the audio data samples including non-keyword speech and the audio data samples including a keyword are pooled into a speech class.

15 . The electronic device of claim 13 , wherein the second layer of the keyword detection model is trained using the second set of audio data samples and the third set of audio data samples to distinguish between non-keyword and keyword-like speech, wherein the audio data samples including keyword-like speech are pooled into a keyword-like class.

16 . The electronic device of claim 13 , wherein the third layer of the keyword detection model is trained using the third set of audio data samples.

17 . A non-transitory machine readable medium containing instructions that when executed cause at least one processor of an electronic device to:

obtain audio data from an audio input device;

provide the audio data as input to a keyword detection model;

extract, using the keyword detection model, audio features from the audio data;

process the audio features by a first layer of the keyword detection model configured to predict a first likelihood that the audio data includes speech;

process the audio features by a second layer of the keyword detection model configured to predict a second likelihood that the audio data includes keyword-like speech;

process the audio features by a third layer of the keyword detection model configured to predict a third likelihood, for each of a plurality of possible keywords, that the audio data includes the keyword;

provide each of the first likelihood, the second likelihood, and the third likelihood to an inference combination layer of the keyword detection model;

identify a keyword included in the audio data based on a combination of the first likelihood, the second likelihood, and the third likelihood by the inference combination layer, and

generate instructions to perform an action based at least in part on the identified keyword.

18 . The non-transitory machine readable medium of claim 17 , further comprising instructions that when executed cause the at least one processor to:

process the audio features by the second layer of the keyword detection model in response to the first likelihood exceeding a first threshold; and

process the audio features by the third layer of the keyword detection model in response to the second likelihood exceeding a second threshold,

wherein, to identify the keyword, the instructions when executed further cause the at least one processor to identify which one of the plurality of possible keywords is associated with a highest third likelihood.

19 . The non-transitory machine readable medium of claim 17 , wherein to identify the keyword, the instructions when executed further cause the at least one processor to determine, based on the first likelihood, the second likelihood, and each of the third likelihood, that the audio data includes the keyword.

20 . The non-transitory machine readable medium of claim 17 , wherein:

the keyword detection model is trained using a training dataset that includes a first set of audio data samples including non-speech audio, a second set of audio data samples including non-keyword speech, and a third set of audio data samples including a keyword;

each audio data sample is annotated with a speech label indicating whether the audio data sample includes speech, a keyword-like label indicating whether the audio data sample includes keyword-like speech, and a keyword label identifying which keyword, if any, is in the audio data sample;

the first layer of the keyword detection model is trained using the first set of audio data samples, the second set of audio data samples, and the third set of audio data samples to distinguish between speech and non-speech audio, wherein the audio data samples including non-keyword speech and the audio data samples including a keyword are pooled into a speech class;

the second layer of the keyword detection model is trained using the second set of audio data samples and the third set of audio data samples to distinguish between non-keyword and keyword-like speech, wherein the audio data samples including keyword-like speech are pooled into a keyword-like class; and

the third layer of the keyword detection model is trained using the third set of audio data samples.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2023
From: SRINIVASA, RAKSHITH SHARMA; SAIDUTTA, YASHAS MALUR; LEE, CHING-HUA; YANG, CHOU-CHANG; SHEN, YILIN; JIN, HONGXIA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 064260/0815 →
Continuity (2)
Provisional Application 63419268 · Oct 25, 2022
Related Publication 20240185850A1 · Jun 6, 2024
References Cited (36)
US 10304440B1 · Panchapagesan · 2019 [cited by examiner]
US 10872599B1 · Wu et al. · 2020 [cited by applicant]
US 11222623B2 · Wang et al. · 2022 [cited by applicant]
US 11264036B2 · Kang et al. · 2022 [cited by applicant]
US 11556793B2 · Gruenstein · 2023 [cited by applicant]
US 11823669B2 · Ding · 2023 [cited by examiner]
US 20150340032A1 · Gruenstein · 2015 [cited by examiner]
US 20170148429A1 · Hayakawa · 2017 [cited by applicant]
US 20190318727A1 · Lopez Moreno · 2019 [cited by examiner]
US 20200336846A1 · Rohde · 2020 [cited by examiner]
US 20210055778A1 · Myer et al. · 2021 [cited by applicant]
US 20220343895A1 · Tomar et al. · 2022 [cited by applicant]
US 20230104431A1 · Smyth et al. · 2023 [cited by applicant]
WO 2021030918A1 · 2021 [cited by applicant]
WO 2023044836A1 · 2023 [cited by applicant]
Ming Sun, et al., “Max-Pooling Loss Training of Long Short-Term Memory Networks for Small-Footprint Keyword Spotting,” 2016 IEEE Spoken Language Technology Workshop (SLT), San Diego, California, Dec. 13-16, 2016, pp. 47… [cited by applicant]
Sercan Ö. Arik, et al., “Convolutional Recurrent Neural Networks for Small-Footprint Keyword Spotting,” Interspeech 2017, Acoustic Models for ASR 1, Stockholm, Sweden, Aug. 20-24, 2017, pp. 1606-1610. [cited by applicant]
Tara N. Sainath, et al., “Convolutional Neural Networks for Small-Footprint Keyword Spotting,” Interspeech 2015, Fast Efficient and Scalable Computing for Neural Nets, Dresden, Germany, Sep. 6-10, 2015, pp. 1478-1482. [cited by applicant]
Raphael Tang, et al., “Deep Residual Learning for Small-Footprint Keyword Spotting,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, Alberta, Canada, Apr. 15-20, 2018, pp… [cited by applicant]
Seungwoo Choi, et al., “Temporal Convolution for Real-time Keyword Spotting on Mobile Devices” Interspeech 2019, Speech and Audio Classification 2, Graz, Austria, Sep. 15-19, 2019, pp. 3372-3376. [cited by applicant]
Ximin Li, et al., “Small-Footprint Keyword Spotting with Multi-Scale Temporal Convolution,” Interspeech 2020, Speech Classification, Shanghai, China, Oct. 25-29, 2020, pp. 1987-1991. [cited by applicant]
Menglong Xu, et al. “Depthwise Separable Convolutional ResNet with Squeeze-and-Excitation Blocks for Small-footprint Keyword Spotting,” Interspeech 2020, Spoken Term Detection, Shanghai, China, Oct. 25-29, 2020, pp. 254… [cited by applicant]
Byeonggeun Kim, et al., “Broadcasted Residual Learning for Efficient Keyword Spotting,” arXiv:2106.04140, Jun. 8, 2021, 5 pages. [cited by applicant]
Mengjun Zeng, et al., “Effective Combination of DenseNet and BiLSTM for Keyword Spotting,” IEEE Access, vol. 7, Jan. 10, 2019, pp. 10767-10775. [cited by applicant]
Oleg Rybakov, et al., “Streaming Keyword Spotting on Mobile Devices,” Interspeech 2020, Applications of ASR, Shanghai, China, Oct. 25-29, 2020, pp. 2277-2281. [cited by applicant]
Axel Berg, et al., “Keyword Transformer: A Self-Attention Model for Keyword Spotting,” Interspeech 2021, Spoken Term Detection & Voice Search, Brno, Czechia, Aug. 30-Sep. 3, 2021, pp. 4249-4253. [cited by applicant]
Ye Bai, et al., “A Time Delay Neural Network with Shared Weight Self-Attention for Small-Footprint Keyword Spotting,” Interspeech 2019, Spoken Term Detection, Confidence Measure, and End-to-End Speech Recognition, Graz,… [cited by applicant]
Douglas Coimbra de Andrade, et al., “A neural attention model for speech command recognition,” arXiv:1808.08929, Aug. 27, 2018, 18 pages. [cited by applicant]
Simon Mittermaier, et al. “Small-Footprint Keyword Spotting on Raw Audio Data with Sinc-Convolutions,” ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain… [cited by applicant]
Bo Zhang, et al., “AutoKWS: Keyword Spotting with Differentiable Architecture Search,” ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, Ontario, Canada, Jun. 6… [cited by applicant]
Tong Mo, et al., “Neural Architecture Search for Keyword Spotting,” Interspeech 2020, Speech Classification, Shanghai, China, Oct. 25-29, 2020, pp. 1982-1986. [cited by applicant]
Bruno U. Pedroni, et al. “Small-footprint Spiking Neural Networks for Power-efficient Keyword Spotting,” 2018 IEEE Biomedical Circuits and Systems Conference (BioCAS), Cleveland, Ohio, Oct. 17-19, 2018, 4 pages. [cited by applicant]
Samuel Myer, et al., “Efficient Keyword Spotting Using Time Delay Neural Networks,” arXiv:1807.04353, Jul. 11, 2018, 5 pages. [cited by applicant]
Vineet Garg, et al., “Streaming Transformer for Hardware Efficient Voice Trigger Detection and False Trigger Mitigation,” Interspeech 2021, Spoken Term Detection & Voice Search, Brno, Czechia, Aug. 30-Sep. 3, 2021, pp. … [cited by applicant]
Yiming Wang, et al., “Wake Word Detection with Alignment-Free Lattice-Free MMI,” Interspeech 2020, Summarization, Semantic Analysis and Classification, Shanghai, China, Oct. 25-29, 2020, pp. 4258-4262. [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority dated Jan. 24, 2024, in connection with International Application No. PCT/IB2023/060599, 7 pages. [cited by applicant]