IP Library Granted Patent US 12,626,690
Granted Patent B2
US 12,626,690 · App. 18/449,237 · Granted May 12, 2026

Systems, methods, and devices for low-power audio signal detection

Inventors: Aidan Smyth (Irvine, CA); Ashutosh Pandey (Irvine, CA); Niall Lyons (Irvine, CA); Ted Wada (Irvine, CA); Robert Zopf (Rancho Santa Margarita, CA)
Assignee: Infineon Technologies Americas Corp.
G10L15/063G10L15/01G10L15/02G10L15/187G10L2015/025G10L2015/0635G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,690
App. No.
18/449,237
Granted
May 12, 2026
Kind
B2
Abstract

Systems, methods, and devices detect wake signals included in audio signals. Methods include receiving a dataset including raw audio data, the raw audio data comprising a plurality of audio samples and associated metadata, and generating, using one or more processing elements, an augmented dataset based on the raw audio data, the augmented dataset comprising a plurality of annotations identifying types of raw audio data. Methods further include generating, using the one or more processing elements, a feature dataset by extracting features from the augmented dataset based, at least in part, on the plurality of annotations, and generating, using the one or more processing elements, a wake signal detection model based, at least in part, on the feature dataset, the wake signal detection model being a machine learning model trained based on the feature dataset.

Claims (58)

1 . A method comprising:

receiving a dataset including raw audio data, the raw audio data comprising a plurality of audio samples and associated metadata;

generating, using one or more processing elements, an augmented dataset based on the raw audio data, the augmented dataset comprising a plurality of annotations identifying types of raw audio data, the plurality of annotations being generated based, at least in part, on the metadata associated with the plurality of audio samples;

generating, using the one or more processing elements, a feature dataset by extracting features from the augmented dataset based, at least in part, on the plurality of annotations and using the plurality of annotations to serialize the extracted features to generate an input for a wake signal detection model; and

generating, using the one or more processing elements, the wake signal detection model based, at least in part, on the feature dataset, the wake signal detection model being a machine learning model trained based on the feature dataset, the generating of the wake signal detection model further comprising reducing a number of dimensions of the machine learning model based, at least in part, on power consumption characteristics of a target audio signal processing device.

2 . The method of claim 1 further comprising:

generating, using the one or more processing elements, training data based, at least in part, on the feature dataset.

3 . The method of claim 2 , wherein the generating of the training data further comprises:

concatenating at least some of the extracted features included in the feature dataset; and

generating an output file based on the concatenation of the at least some of the extracted features.

4 . The method of claim 1 , wherein the generating of the augmented dataset further comprises:

classifying a plurality of phonemes included in the raw audio data;

generating a plurality of tokens based on the plurality of phonemes; and

generating the plurality of annotations based on the plurality of tokens.

5 . The method of claim 4 , wherein the classifying is performed by an automatic speech recognition model, and wherein the plurality of tokens identify whether or not speech is present in each of the plurality of phonemes.

6 . The method of claim 1 further comprising:

testing the wake signal detection model using test data.

7 . The method of claim 6 further comprising:

modifying one or more weights associated with the wake signal detection model based on a result of the testing.

8 . The method of claim 1 further comprising:

generating a low-power model based on the wake signal detection model, the low-power model having the reduced number of dimensions.

9 . The method of claim 8 , wherein the low-power model is configured to execute on a low-power device in real-time.

10 . A system comprising:

a communications interface configured to receive raw audio data, the raw audio data comprising a plurality of audio samples and associated metadata;

one or more processing elements configured to:

generate a dataset based on the raw audio data received from the communications interface;

generate an augmented dataset based on the raw audio data, the augmented dataset comprising a plurality of annotations identifying types of raw audio data, the plurality of annotations being generated based, at least in part, on the metadata associated with the plurality of audio samples;

generate a feature dataset by extracting features from the augmented dataset based, at least in part, on the plurality of annotations and using the plurality of annotations to serialize the extracted features to generate an input for a wake signal detection model; and

generate the wake signal detection model based, at least in part, on the feature dataset, the wake signal detection model being a machine learning model trained based on the feature dataset, the generating of the wake signal detection model further comprising reducing a number of dimensions of the machine learning model based, at least in part, on power consumption characteristics of a target audio signal processing device.

11 . The system of claim 10 , wherein the one or more processing elements are further configured to:

generate training data based, at least in part, on the feature dataset.

12 . The system of claim 11 , wherein the one or more processing elements are further configured to:

concatenate at least some of the extracted features included in the feature dataset; and

generate an output file based on the concatenation of the at least some of the extracted features.

13 . The system of claim 10 , wherein the one or more processing elements are further configured to:

classify a plurality of phonemes included in the raw audio data, wherein the classifying is performed by an automatic speech recognition model;

generate a plurality of tokens based on the plurality of phonemes, wherein the plurality of tokens identify whether or not speech is present in each of the plurality of phonemes; and

generate the plurality of annotations based on the plurality of tokens.

14 . The system of claim 10 , wherein the one or more processing elements are further configured to:

generate a low-power model based on the wake signal detection model, the low-power model having the reduced number of dimensions.

15 . The system of claim 14 , wherein the low-power model is configured to execute on a low-power device in real-time.

16 . A device comprising:

one or more processing elements configured to:

receive a dataset including raw audio data, the raw audio data comprising a plurality of audio samples and associated metadata;

generate an augmented dataset based on the raw audio data, the augmented dataset comprising a plurality of annotations identifying types of raw audio data, the plurality of annotations being generated based, at least in part, on the metadata associated with the plurality of audio samples;

generate a feature dataset by extracting features from the augmented dataset based, at least in part, on the plurality of annotations and using the plurality of annotations to serialize the extracted features to generate an input for a wake signal detection model; and

generate the wake signal detection model based, at least in part, on the feature dataset, the wake signal detection model being a machine learning model trained based on the feature dataset, the generating of the wake signal detection model further comprising reducing a number of dimensions of the machine learning model based, at least in part, on power consumption characteristics of a target audio signal processing device.

17 . The device of claim 16 , wherein the one or more processing elements are further configured to:

generate training data based, at least in part, on the feature dataset.

18 . The device of claim 17 , wherein the one or more processing elements are further configured to:

concatenate at least some of the extracted features included in the feature dataset; and

generate an output file based on the concatenation of the at least some of the extracted features.

19 . The device of claim 16 , wherein the one or more processing elements are further configured to:

classify a plurality of phonemes included in the raw audio data, wherein the classifying is performed by an automatic speech recognition model;

generate a plurality of tokens based on the plurality of phonemes, wherein the plurality of tokens identify whether or not speech is present in each of the plurality of phonemes; and

generate the plurality of annotations based on the plurality of tokens.

20 . The device of claim 16 , wherein the one or more processing elements are further configured to:

generate a low-power model based on the wake signal detection model, wherein the low-power model has the reduced number of dimensions, and wherein the low-power model is configured to execute on a low-power device in real-time.

Assignments (2)
MERGER AND CHANGE OF NAME Recorded Oct 21, 2025
From: CYPRESS SEMICONDUCTOR CORPORATION; INFINEON TECHNOLOGIES AMERICAS CORP.
To: INFINEON TECHNOLOGIES AMERICAS CORP.
Reel/Frame 073140/0554 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2023
From: SMYTH, AIDAN; PANDEY, ASHUTOSH; LYONS, NIALL; WADA, TED; ZOPF, ROBERT
To: CYPRESS SEMICONDUCTOR CORPORATION
Reel/Frame 065685/0952 →
Continuity (2)
Provisional Application 63398370 · Aug 16, 2022
Related Publication 20240062745A1 · Feb 22, 2024
References Cited (16)
US 11227122B1 · Gill · 2022 [cited by examiner]
US 11355102B1 · Mishchenko · 2022 [cited by examiner]
US 11551670B1 · Smith · 2023 [cited by examiner]
US 12112752B1 · Gupta · 2024 [cited by examiner]
US 20170270919A1 · Parthasarathi et al. · 2017 [cited by applicant]
US 20200279561A1 · Sheeder et al. · 2020 [cited by applicant]
US 20200349925A1 · Shahid · 2020 [cited by examiner]
US 20200365138A1 · Kim · 2020 [cited by examiner]
US 20210050003A1 · Zaheer · 2021 [cited by examiner]
US 20210174794A1 · Mont-Reynaud · 2021 [cited by examiner]
US 20210249035A1 · Bone et al. · 2021 [cited by applicant]
US 20210350798A1 · Zopf · 2021 [cited by examiner]
US 20210385319A1 · Holleman, III · 2021 [cited by applicant]
US 20220068272A1 · Kwatra et al. · 2022 [cited by applicant]
Gao, Yixin, et al. “Towards data-efficient modeling for wake word spotting.” ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020. (Year: 2020). [cited by examiner]
USPTO Search Report and Written Opinion from Application PCT/US23/30277 dated Nov. 9, 2023 14 pages. [cited by applicant]