IP Library Granted Patent US 12,512,093
Granted Patent B2
US 12,512,093 · App. 16/529,456 · Granted Dec 30, 2025

Sensor-processing systems including neuromorphic processing modules and methods thereof

Inventors: Kurt F. Busch (Laguna Hills, CA); Jeremiah H. Holleman, III (Davidson, NC); Pieter Vorenkamp (Laguna Beach, CA); Stephen W. Bailey (Irvine, CA); David Christopher Garrett (Tustin, CA)
Assignee: SYNTIANT
G10L15/16G06N3/08G10L15/02G10L15/22G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,512,093
App. No.
16/529,456
Granted
Dec 30, 2025
Kind
B2
Abstract

Disclosed is a sensor-processing system including, in some embodiments, a sensor, one or more sample pre-processing modules, one or more sample-processing modules, one or more neuromorphic integrated circuits (“ICs”), and a microcontroller. The one or more sample pre-processing modules are configured to process raw sensor data for use in the sensor-processing system. The one or more sample-processing modules are configured to process pre-processed sensor data including extracting features from the pre-processed sensor data. Each of the neuromorphic ICs includes at least one neural network configured to arrive at actionable decisions of the neural network from the features extracted from the pre-processed sensor data. The microcontroller includes a CPU along with memory including instructions for operating the sensor-processing system. In some embodiments, the sensor is a pulse-density modulation (“PDM”) microphone, and the sensor-processing system is configured for keyword spotting. Also disclosed are methods of such a keyword spotting sensor-processing system.

Claims (49)

1 . A sensor-processing system, comprising:

one or more sensors that may be operated in first mode to conserve power, and a second mode to achieve a desired signal-to-noise ratio for a subsequent audio sample;

one or more sample pre-processing modules configured to process raw sensor data for use in the sensor-processing system, wherein the one or more sample pre-processing modules include a PDM decimation module configured to decimate audio samples from a pulse-density modulation (“PDM”) microphone to a baseband audio sampling rate for use in the sensor-processing system;

one or more sample-processing modules comprising a time domain-processing module configured to process amplitudes of signals in the audio samples multiplexed by a mux module, and a frequency domain-processing module configured to process frequencies of the signals in the audio sample;

wherein the time domain-processing module is configured to process amplitudes of signals in the audio samples multiplexed by a mux module;

one or more sample-processing modules coupled with a feature store is configured to temporarily store features extracted from pre-processed sensor data and the one or more sample pre-processing modules, wherein the one or more sample-processing modules are configured to process pre-processed sensor data including extracting features from the pre-processed sensor data;

a sample holding tank configured to temporarily store audio samples formatted for processing, wherein the sample holding tank is configured to provide audio samples for subsequent analysis;

one or more neuromorphic Integrated Circuits (“ICs”), each neuromorphic IC comprising: a smaller, secondary neural network configured for initial keyword detection and assigned speaker identification and a larger, primary neural network configured to confirm keyword detection;

an initial firmware comprising synaptic weights that may be updated;

a microcontroller including at least one central-processing unit (“CPU”) along with memory including instructions for operating the sensor-processing system, wherein the one or more sensors include at least one of: an accelerometer, a temperature sensor, and a microphone;

in response to determining that a signal is present in a time domain of an audio sample, determining whether the signal represents speech in the frequency domain of the audio sample and determining whether the speech includes features which are characteristic of an assigned speaker based on further analysis of the audio samples in the sample holding tank;

and, in response to determining that the signal represents speech and the speech includes features which are characteristic of an assigned speaker, operating the primary neural network to confirm presence of a keyword after initial detection by the secondary neural network.

2 . The sensor-processing system of claim 1 , wherein the feature store is configured to at least temporarily store the features extracted from the pre-processed sensor data for the one or more neuromorphic ICs.

3 . The sensor-processing system of claim 1 , wherein the sensor-processing system includes a single neuromorphic IC including a single neural network configured as a classifier.

4 . The sensor-processing system of claim 1 , wherein the sensor-processing system includes at least a first neuromorphic IC including a relatively larger, primary neural network and a second neuromorphic IC including a relatively smaller, secondary neural network; and

wherein the primary neural network is configured to power on and operate on the features extracted from the pre-processed sensor data after the secondary neural network arrives at an actionable decision on the features extracted from the pre-processed sensor data, thereby lowering power consumption of a sensor-processing multi-chip.

5 . The sensor-processing system of claim 1 , wherein the sensor-processing system is configured as a keyword spotter; and

wherein the features are one or more signals in a time domain, a frequency domain, or both the time and frequency domains characteristic of keywords one or more neural networks are trained to recognize.

6 . A method of conditional neural network operation in a sensor-processing system upon detection of a credible signal, comprising:

operating a pulse-density modulation (“PDM”) microphone, a PDM decimation module, a time domain-processing module, and a frequency domain-processing module, operating the PDM microphone in a first mode to conserve power, and wherein the PDM decimation module decimates an audio sample from the PDM microphone to a baseband audio sampling rate;

wherein operating the time domain-processing module and the frequency domain-processing module includes identifying one or more signals of the audio sample in a time domain or a frequency domain if the one or more signals of the audio sample are present;

wherein the time domain-processing module is configured to process amplitudes of signals in the audio samples multiplexed by a mux module;

temporarily storing extracted features in a feature store for subsequent analysis;

determining whether the signal represents speech in the frequency domain of the audio sample;

operating a smaller, secondary neural network to determine if the one or more signals includes a portion of a keyword and to identify an assigned speaker;

configuring a sample holding tank to temporarily store audio samples formatted for processing;

operating the PDM microphone in a second mode to achieve a desired signal-to-noise ratio for a subsequent audio sample in response to determining that the one or more signals are present in the time domain;

determining whether the one or more signals represents speech in the frequency domain of the audio sample;

determining whether the speech includes features which are characteristic of the assigned speaker based on further analysis of the audio samples in the sample holding tank;

and, in response to determining that the signal represents speech and the speech includes features which are characteristic of the assigned speaker, powering on and operating a larger, primary neural network to confirm if the one or more signals includes the keyword or a portion thereof.

7 . The method of claim 6 , further comprising:

pulling the audio sample from the sample holding tank to either:

confirm the one or more signals includes a keyword; or

process the audio sample via an alternative method.

8 . A method of conditional neural network operation in a sensor-processing system upon detection of a credible keyword, comprising:

operating a pulse-density modulation (“PDM”) microphone, a PDM decimation module, a time domain-processing module, and a frequency domain-processing module, operating the PDM microphone in a first mode to conserve power, and wherein the PDM decimation module decimates an audio sample from the PDM microphone to a baseband audio sampling rate;

wherein operating the time domain-processing module and the frequency domain-processing module includes identifying one or more signals of the audio sample in a time domain or a frequency domain if the one or more signals of the audio sample are present;

wherein the time domain-processing module is configured to process amplitudes of signals in the audio samples multiplexed by a mux module;

temporarily storing extracted features in a feature store for subsequent analysis by a neural network;

wherein in response to determining that a signal is present in the time domain of the audio sample, determining whether the signal represents speech in the frequency domain of the audio sample;

operating a low-powered secondary neural network in response to determining that the one or more signals represent speech in the frequency domain, to determine if the one or more signals includes a keyword;

configuring a sample holding tank to temporarily store audio samples formatted for processing without real-time re-acquisition of the audio samples;

operating the PDM microphone in a second mode to achieve a desired signal-to-noise ratio for a subsequent audio sample in response to determining that a signal is present in the time domain;

operating the low-powered secondary neural network to perform further analysis of the audio samples in the sample holding tank, the further analysis for determining whether the speech includes features which are characteristic of an assigned speaker;

and powering on and operating a high-powered primary neural network if both the one or more signals include a keyword or a portion of the keyword and the speech including features which are characteristic of an assigned speaker, wherein the primary neural network confirms the one or more signals include a keyword or a portion of the keyword.

9 . The method of claim 8 , further comprising:

pulling the audio sample from the sample holding tank to either:

confirm the one or more signals includes a keyword; or

process the audio sample via an alternative method.

Assignments (3)
SECURITY INTEREST Recorded Dec 27, 2024
From: SYNTIANT CORP.; PILOT AI LABS, INC.; SYNTIANT TAIWAN LLC; SYNTIANT HOLDINGS LLC
To: OCEAN II PLO LLC
Reel/Frame 069687/0757 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 21, 2020
From: GARRETT, DAVID CHRISTOPHER
To: SYNTIANT
Reel/Frame 053263/0391 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2019
From: BUSCH, KURT F.; HOLLEMAN, JEREMIAH H., III; VORENKAMP, PIETER; BAILEY, STEPHEN W.
To: SYNTIANT
Reel/Frame 050918/0770 →
Continuity (2)
Provisional Application 62713423 · Aug 1, 2018
Related Publication 20200043477A1 · Feb 6, 2020
References Cited (21)
US 9478231B1 · Soman · 2016 [cited by examiner]
US 9953634B1 · Pearce et al. · 2018 [cited by applicant]
US 20140089232A1 · Buibas · 2014 [cited by examiner]
US 20140244273A1 · Laroche et al. · 2014 [cited by applicant]
US 20150066498A1 · Ma · 2015 [cited by examiner]
US 20150269954A1 · Ryan · 2015 [cited by examiner]
US 20160196838A1 · Rossum et al. · 2016 [cited by applicant]
US 20160351197A1 · Tan · 2016 [cited by examiner]
US 20170229117A1 · van der Made · 2017 [cited by examiner]
US 20170236051A1 · van der Made · 2017 [cited by examiner]
US 20180039768A1 · Roberts · 2018 [cited by examiner]
US 20180276537A1 · Wood · 2018 [cited by examiner]
US 20180315416A1 · Berthelsen · 2018 [cited by examiner]
US 20190013037A1 · Haiut · 2019 [cited by examiner]
US 20190042910A1 · Krishnamurthy · 2019 [cited by examiner]
US 20190333522A1 · Lesso · 2019 [cited by examiner]
US 20200035233A1 · Lee · 2020 [cited by examiner]
US 20210304734A1 · Kang · 2021 [cited by examiner]
Price, M., Glass, J. and Chandrakasan, A.P., 2017. A low-power speech recognizer and voice activity detector using deep neural networks. IEEE Journal of Solid-State Circuits, 53(1), pp. 66-75. (Year: 2017). [cited by examiner]
International Search Report and Written Opinion, PCT Application No. PCT/US19/44713, mailed Oct. 25, 2019. [cited by applicant]
Supplementary Partial European Search Report issued in corresponding EP Application No. 19844888.8, dated Aug. 23, 2022. [cited by applicant]