IP Library Granted Patent US 9,953,634
Granted Patent B1
US 9,953,634 · App. 14/573,846 · Granted Apr 24, 2018

Passive training for automatic speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,953,634
App. No.
14/573,846
Granted
Apr 24, 2018
Kind
B1
Abstract

Provided are methods and systems for passive training for automatic speech recognition. An example method includes utilizing a first, speaker-independent model to detect a spoken keyword or a key phrase in spoken utterances. While utilizing the first model, a second model is passively trained to detect the spoken keyword or the key phrase in the spoken utterances using at least partially the spoken utterances. The second, speaker dependent model may utilize deep neural network (DNN) or convolutional neural network (CNN) techniques. In response to completion of the training, a switch is made from utilizing the first model to utilizing the second model to detect the spoken keyword or the key phrase in spoken utterances. While utilizing the second model, parameters associated therewith are updated using the spoken utterances in response to detecting the keyword or the key phrase in the spoken utterances. User authentication functionality may be provided.

Claims (33)

1. A method for passive training for automatic speech recognition, the method comprising:

maintaining both a first model and a second model that are configured to detect a spoken keyword or a key phrase in spoken utterances, wherein the spoken utterances comprise words and/or phrases spoken by a user;

utilizing only the first model, the first model being a speaker-independent model, to detect the spoken keyword or the key phrase in spoken utterances until the second model has been trained;

passively training the second model, using the spoken utterances comprising words and/or phrases spoken by the user, to detect the spoken keyword or the key phrase in the spoken utterances; and

in response to completion of the passive training, wherein completion requires performing the passive training with a plurality of the spoken utterances satisfying a predetermined set of criteria, switching from utilizing only the first, speaker-independent model to utilizing only the second, passively trained model to detect the spoken keyword or the key phrase in the spoken utterances,

wherein passively training the second model includes selecting at least one utterance from the spoken utterances comprising words and/or phrases spoken by the user according to at least one predetermined criterion, and wherein the at least one pre-determined criterion includes one or more of a determination that a signal-to-noise ratio level of the selected utterance is below a pre-determined level and a determination that a duration of the selected utterance is below a pre-determined time period.

2. The method of claim 1 , wherein the second model is a speaker-dependent model for the user.

3. The method of claim 1 , wherein the second model is operable to provide at least one additional functionality different from detecting the spoken keyword or the key phrase.

4. The method of claim 3 , wherein the at least one additional functionality includes an authentication of the user.

5. The method of claim 1 , wherein the second model includes one or more of the following: a deep neural network (DNN) and a convolutional neural network (CNN).

6. The method of claim 1 , wherein a threshold of detecting the keyword or the key phrase by the second model is narrower than a threshold of detecting the keyword or the key phrase by the first model, such that the second model provides substantially more sensitive keyword detection compared to the first model.

7. The method of claim 1 , wherein the passive training of the second model completes upon collecting a pre-determined number of the selected at least one utterance.

8. The method of claim 1 , further comprising updating parameters associated with the second model using the spoken utterances in response to detecting the keyword or the key phrase in the spoken utterances.

9. A system for passive training for automatic speech recognition, comprising:

at least one processor; and

a memory communicatively coupled to the at least one processor, the memory storing instructions which when executed by the at least one processor perform a method comprising:

maintaining both a first model and a second model that are configured to detect a spoken keyword or a key phrase in spoken utterances, wherein the spoken utterances comprise words and/or phrases spoken by a user;

utilizing only the first model, the first model being a speaker-independent model, to detect the spoken keyword or the key phrase in spoken utterances until the second model has been trained;

passively training the second model, using the spoken utterances comprising words and/or phrases spoken by the user, to detect the spoken keyword or the key phrase in the spoken utterances; and

in response to completion of the passive training, wherein completion requires performing the passive training with a plurality of the spoken utterances satisfying a predetermined set of criteria, switching from utilizing only the first, speaker-independent model to utilizing only the second, passively trained model to detect the spoken keyword or the key phrase in the spoken utterances,

wherein passively training the second model includes selecting at least one utterance from the spoken utterances comprising words and/or phrases spoken by the user according to at least one predetermined criterion, and wherein the at least one pre-determined criterion includes one or more of a determination that a signal-to-noise ratio level of the selected utterance is below a pre-determined level and a determination that a duration of the selected utterance is below a pre-determined time period.

10. The system of claim 9 , wherein the second model is a speaker-dependent model for the user.

11. The system of claim 9 , wherein the second model includes one or more of the following: a deep neural network (DNN) and a convolutional neural network (CNN).

12. The system of claim 9 , wherein the second model is operable to provide at least one additional functionality different from detecting the spoken keyword or the key phrase.

13. The system of claim 12 , wherein the at least one additional functionality includes an authentication of the user.

14. The system of claim 9 , wherein a threshold of detecting the keyword or the key phrase by the second model is narrower than a threshold of detecting the keyword or the key phrase by the first model.

15. The system of claim 9 , further comprising updating parameters associated with the second model using the spoken utterances in response to detecting the keyword or the key phrase in the spoken utterances.

16. A non-transitory processor-readable medium having embodied thereon a program being executable by at least one processor to perform a method for passive training for automatic speech recognition, the method comprising:

maintaining both a first model and a second model that are configured to detect a spoken keyword or a key phrase in spoken utterances, wherein the spoken utterances comprise words and/or phrases spoken by a user;

utilizing only the first model, the first model being a speaker-independent model, to detect the spoken keyword or the key phrase in spoken utterances until the second model has been trained;

passively training the second model, using the spoken utterances comprising words and/or phrases spoken by the user, to detect the spoken keyword or the key phrase in the spoken utterances; and

in response to completion of the passive training, wherein completion requires performing the passive training with a plurality of the spoken utterances satisfying a predetermined set of criteria, switching from utilizing only the first, speaker-independent model to utilizing only the second, passively trained model to detect the spoken keyword or the key phrase in the spoken utterances,

wherein passively training the second model includes selecting at least one utterance from the spoken utterances comprising words and/or phrases spoken by the user according to at least one predetermined criterion, and wherein the at least one pre-determined criterion includes one or more of a determination that a signal-to-noise ratio level of the selected utterance is below a pre-determined level and a determination that a duration of the selected utterance is below a pre-determined time period.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2024
From: KNOWLES ELECTRONICS, LLC
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 066216/0464 →
CHANGE OF NAME Recorded Feb 25, 2016
From: AUDIENCE, INC.
To: AUDIENCE LLC
Reel/Frame 037927/0424 →
MERGER Recorded Feb 25, 2016
From: AUDIENCE LLC
To: KNOWLES ELECTRONICS, LLC
Reel/Frame 037927/0435 →
EMPLOYMENT, CONFIDENTIAL INFORMATION AND INVENTION ASSIGNMENT AGREEMENT Recorded May 29, 2015
From: PEARCE, DAVID
To: AUDIENCE, INC.
Reel/Frame 035797/0496 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 27, 2015
From: CLARK, BRIAN
To: AUDIENCE, INC.
Reel/Frame 035726/0205 →