IP Library Granted Patent US 10,719,115
Granted Patent B2
US 10,719,115 · App. 14/606,588 · Granted Jul 21, 2020

Isolated word training and detection using generated phoneme concatenation models of audio inputs

Inventors: Robert W. Zopf (Rancho Santa Margarita, CA); Bengt J. Borgstrom (Santa Monica, CA)
Assignee: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
G06F1/3231G10L15/063G10L15/144G10L2015/025G10L2015/0635G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,719,115
App. No.
14/606,588
Granted
Jul 21, 2020
Kind
B2
Abstract

Methods, systems, and apparatuses are described for isolated word training and detection. Isolated word training devices and systems are provided in which a user may provide a wake-up phrase from 1 to 3 times to train the device or system. A concatenated phoneme model of the user-provided wake-up phrase may be generated based on the provided wake-up phrase and a pre-trained phoneme model database. A word model of the wake-up phrase may be subsequently generated from the concatenated phoneme model and the provided wake-up phrase. Once trained, the user-provided wake-up phrase may be used to unlock the device or system and/or to wake up the device or system from a standby mode of operation. The word model of the user-provided wake-up phrase may be further adapted based on additional provisioning of the wake-up phrase.

Claims (68)

1. An isolated word training system that comprises:

a hardware processor;

a hardware input device configured to receive at least one audio input representation;

a recognition component executing on the hardware processor, the recognition component configured to generate a phoneme concatenation model of the at least one audio input representation based on a phoneme transcription; and

an adaptation component executing on the hardware processor, the adaptation component configured to generate a first word model of the at least one audio input based on the phoneme concatenation model.

2. The isolated word training system of claim 1 , wherein the adaptation component is further configured to generate a second word model based on the first word model and at least one additional audio input representation.

3. The isolated word training system of claim 1 , further comprising:

a phoneme model database that includes the one or more stored phoneme models, each of the one or more stored phoneme models being a pre-trained Hidden Markov Model.

4. The isolated word training system of claim 1 , wherein the recognition component is configured to generate the phoneme concatenation model by:

decoding feature vectors of the at least one audio input using the one or more stored phoneme models to generate the phoneme transcription that comprises one or more phoneme identifiers;

selecting a subset of phoneme identifiers from the phoneme transcription; and

selecting corresponding phoneme models from a phoneme model database for each phoneme identifier in the subset as the phoneme concatenation model.

5. The isolated word training system of claim 4 , further comprising:

a feature extraction component configured to:

derive speech features of speech of the at least one audio input; and

generate the feature vectors based on the speech features.

6. The isolated word training system of claim 5 , further comprising:

a voice activity detection component configured to:

detect an onset of the speech;

determine a termination of the speech; and

provide the speech to the feature extraction component.

7. The isolated word training system of claim 1 , wherein the adaptation component is configured to generate the first word model by at least one of:

removing one or more phonemes from the phoneme concatenation model;

combining one or more phonemes of the phoneme concatenation model;

adapting one or more state transition probabilities of the phoneme concatenation model;

adapting an observation symbol probability distribution of the phoneme concatenation model; or

adapting the phoneme concatenation model as a whole.

8. An electronic device that comprises:

a first processing circuit configured to:

receive a user-specified wake-up phrase that is previously unassociated with the electronic device as a wake-up phrase,

generate a phoneme concatenation model of the user-specified wake-up phrase, and

generate a word model of the user-specified wake-up phrase based on the phoneme concatenation model; and

a second processing circuit configured to:

detect audio activity of the user, and

determine if the user-specified wake-up phrase is present within the audio activity based on the word model.

9. The electronic device of claim 8 , wherein the first processing component is further configured to operate in a training mode of the electronic device, and

wherein the second processing component is further configured to:

operate in a stand-by mode of the electronic device, and

provide, in the stand-by mode, an indication to the electronic device to operate in a normal operating mode subsequent to a determination that the user-specified wake-up phrase is present in the audio activity.

10. The electronic device of claim 9 , further comprising:

a memory component configured to:

buffer the audio activity and the user-specified wake-up phrase of the user; and

provide the audio activity or the user-specified wake-up phrase to at least one of a voice activity detector or an automatic speech recognition component.

11. The electronic device of claim 9 , wherein the first processing component, in the training mode, is further configured to:

decode feature vectors of the at least one audio input using the one or more stored phoneme models to generate a phoneme transcription that comprises one or more phoneme identifiers;

select a subset of phoneme identifiers from the phoneme transcription; and

select corresponding phoneme models from a phoneme model database for each phoneme identifier in the subset as the phoneme concatenation model.

12. The electronic device of claim 11 , wherein the first processing component, in the training mode, is further configured to:

derive speech features of the user-specified wake-up phrase; and

generate the feature vectors based on the speech features.

13. The electronic device of claim 9 , wherein the second processing component, in the stand-by mode, is further configured to:

derive speech features of the audio activity;

generate feature vectors based on the speech features; and

determine if the user-specified wake-up phrase is present by comparing the feature vectors to the word model.

14. The electronic device of claim 8 , wherein the first processing component is further configured to:

generate the first word model by at least one of:

removing one or more phonemes from the phoneme concatenation model;

combining one or more phonemes of the phoneme concatenation model;

adapting one or more state transition probabilities of the phoneme concatenation model;

adapting an observation symbol probability distribution of the phoneme concatenation model; or

adapting the phoneme concatenation model as a whole.

15. The electronic device of claim 8 , wherein the first processing component is further configured to update the word model based on the user-specified wake-up phrase in a subsequent audio input.

16. The electronic device of claim 8 , further comprising:

a phoneme model database that includes one or more stored phoneme models, each of the one or more stored phoneme models being a pre-trained Hidden Markov Model.

17. A non-transitory computer-readable storage medium having program instructions recorded thereon that, when executed by an electronic device, perform a method for utilizing a user-specified wake-up phrase in the electronic device, the method comprising:

training the electronic device using the user-specified wake-up phrase, received from a user, to generate, at the electronic device, a phoneme transcription of the user-specified wake-up phrase that includes one or more phoneme identifiers and a phoneme concatenation model of the user-specified wake-up phrase based on stored phoneme models that correspond to the one or more phoneme identifiers, the user-specified wake-up phrase being previously unassociated with the electronic device as a wake-up phrase; and

transitioning the electronic device from a stand-by mode to a normal operating mode subsequent to a detection, that is based on the word model, of the user-specified wake-up phrase.

18. The non-transitory computer-readable storage medium of claim 17 , wherein training the electronic device using the user-specified wake-up phrase includes receiving the user-specified wake-up phrase from the user no more than three times.

Assignments (6)
CORRECTIVE ASSIGNMENT TO CORRECT THE EXECUTION DATE OF THE MERGER AND APPLICATION NOS. 13/237,550 AND 16/103,107 FROM THE MERGER PREVIOUSLY RECORDED ON REEL 047231 FRAME 0369. ASSIGNOR(S) HEREBY CONFIRMS THE MERGER. Recorded Mar 8, 2019
From: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
To: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
Reel/Frame 048549/0113 →
MERGER Recorded Oct 4, 2018
From: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
To: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
Reel/Frame 047231/0369 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Feb 3, 2017
From: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
To: BROADCOM CORPORATION
Reel/Frame 041712/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2017
From: BROADCOM CORPORATION
To: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
Reel/Frame 041706/0001 →
PATENT SECURITY AGREEMENT Recorded Feb 11, 2016
From: BROADCOM CORPORATION
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 037806/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2015
From: ZOPF, ROBERT W.; BORGSTROM, BENGT J.
To: BROADCOM CORPORATION
Reel/Frame 034822/0821 →
Continuity (2)
Provisional Application 62098143 · Dec 30, 2014
Related Publication 20160189706A1 · Jun 30, 2016