IP Library Patent Application 19191556
Patent Application
App. No. 19/191,556

Data Free Speech Recognition

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/191,556
Abstract

Described are techniques to recognize spoken wake word (WW) or command for human-machine interface using a speech recognition system that does not require any WW or command-matching speech data for training. The system uses the text or grapheme representation of the WW or commands for training before deployment. The technique includes receiving by the system a text representation of a target phrase of a target language. It includes training an acoustic model based on a speech database to distinguish speech signals according to acoustic units of the target language. The training of the acoustic model is independent of the target phrase, It includes constructing a recognition model based on the text representation of the target phrase and the acoustic model to recognize the target phrase in speech; and processing speech from a speaker based on the acoustic model and the recognition model to detect a presence of the target phrase.

Claims (58)

1 . A method of recognizing speech by a device, the method comprising:

receiving a text representation of a target phrase of a target language;

training an acoustic model based on a speech database to distinguish speech signals according to acoustic units of the target language, said training of the acoustic model being independent of the target phrase;

constructing a recognition model based on the text representation of the target phrase and the acoustic model to recognize the target phrase in speech; and

processing speech from a speaker based on the acoustic model and the recognition model to detect a presence of the target phrase.

2 . The method of claim 1 , wherein the target phrase comprises at least one of:

a wake word spoken by the speaker to address the device;

a command spoken by the speaker to control the device; or

a command spoken by the speaker following the wake word to control the device.

3 . The method of claim 1 , wherein the speech database comprises:

a plurality of speech files, wherein each of the speech files represents spoken speech data;

annotations indicating true acoustic units of phonetic content of the plurality of speech files; and

annotations indicating boundaries of each of the true acoustic units.

4 . The method of claim 3 , wherein the acoustic model is a machine learning model, and wherein training the acoustic model comprises:

analyzing spectral and temporal characteristics of the phonetic content of the plurality of speech files to generate observation vectors;

training the mAachine learning model based on the observation vectors and the true acoustic units of the phonetic content to distinguish the phonetic content according to the acoustic units of the target language;

generating a probability estimate for each of the acoustic units of the target language that said acoustic unit represents the phonetic content on a frame-by-frame basis; and

generating statistics characterizing the acoustic model based on the probability estimates of the acoustic units for a plurality of frames, the true acoustic units of the phonetic content, and the boundaries of the true acoustic units.

5 . The method of claim 4 , wherein the statistics comprise, for each of the true acoustic units:

a probability estimate that each of all remaining acoustic units or a silent class is present when said true acoustic unit is present.

6 . The method of claim 4 wherein constructing the recognition model comprises:

constructing the recognition model to recognize acoustically similar acoustic units based on the statistics.

7 . The method of claim 4 wherein constructing the recognition model comprises:

converting, by a tokenizer, the text representation of the target phrase into a sequence of acoustic units representing the target phrase; and

constructing the recognition model based on the sequence of acoustic units and the statistics characterizing the acoustic model to recognize a plurality of pronunciations of the target phrase.

8 . The method of claim 7 , wherein the tokenizer is trained to support a plurality of pronunciations or a plurality of accents of the target phrase.

9 . The method of claim 7 , wherein training of the tokenizer is independent of training of the acoustic model to enable constructing a plurality of recognition models to recognize a plurality of target phrases based on an identical acoustic model.

10 . The method of claim 7 , wherein the tokenizer is trained based on a phonetic dictionary of word-token transcriptions of the target language.

11 . The method of claim 1 , wherein processing speech from a speaker comprises:

detecting an onset of the speech;

analyzing spectral and temporal characteristics of the speech to generate observation vectors;

generating by the acoustic model a sequence of acoustic units based on the observation vectors; and

detecting by the recognition model the presence of the target phrase based on the sequence of acoustic unit.

12 . The method of claim 11 , wherein generating by the acoustic model the sequence of acoustic units comprises:

generating a probability estimate for each of the acoustic units of the target language that said acoustic unit represents the speech on a frame-by-frame basis, and wherein detecting by the recognition model the presence of the target phrase comprises:

determining the presence or an absence of the target phrase based on the probability estimates associated with the acoustic units for a plurality of frames of the speech.

13 . The method of claim 11 , wherein the recognition model comprises a sequence decoding model.

14 . The method of claim 13 , wherein the sequence decoding model comprises a Hidden Markov Model (HMM), and wherein the HMM comprises:

a plurality of states, wherein each of the states decodes an acoustic unit of the target phrase.

15 . The method of claim 14 , wherein at least one of the states decodes a plurality of acoustic units to support alternate pronunciations of the target phrase.

16 . The method of claim 14 , wherein at least one of the states decodes an acoustic unit based on a probability estimate associated with each of all remaining acoustic units of the target language when decoding said acoustic unit.

17 . The method of claim 14 , wherein detecting by the recognition model the presence of the target phrase comprises:

computing a score for determining the presence of the target phrase based on behavior of the plurality of states and statistical information for the plurality of states, wherein the statistical information includes one or more of:

statistics on expected behavior for transitioning between the states;

statistics on expected behavior for transitioning within a same state; or

statistics on acoustically similar acoustic units for each of the plurality of states decoding an acoustic unit of the target phrase.

18 . The method of claim 17 , wherein constructing the recognition model comprises:

generating the statistical information for the plurality of states.

19 . The method of claim 1 , wherein the acoustic units comprise:

phonemes of the target language; and

a silence class.

20 . An apparatus comprising:

an input terminal configured to receive an audio signal from one or more microphones; and

a processing system configured to:

receive a text representation of a target phrase of a target language;

train an acoustic model based on a speech database to distinguish speech signals according to acoustic units of the target language, said training of the acoustic model being independent of the target phrase;

construct a recognition model based on the text representation of the target phrase and the acoustic model to recognize the target phrase in speech; and

process the audio signal based on the acoustic model and the recognition model to detect a presence of the target phrase.

Assignments (1)
MERGER AND CHANGE OF NAME Recorded Oct 21, 2025
From: CYPRESS SEMICONDUCTOR CORPORATION; INFINEON TECHNOLOGIES AMERICAS CORP.
To: INFINEON TECHNOLOGIES AMERICAS CORP.
Reel/Frame 073140/0554 →