IP Library Granted Patent US 11,049,495
Granted Patent B2
US 11,049,495 · App. 16/086,294 · Granted Jun 29, 2021

Method and device for automatically learning relevance of words in a speech recognition system

Inventors: Vikrant Tomar (Montreal, CA); Vincent P. G. Renkens (Herent, BE); Hugo R. J. G. Van Hamme (Leuven, BE)
Assignee: Fluent.ai Inc.
G10L15/063G10L15/07G10L15/16G10L15/1815G10L15/22G10L15/30G10L2015/0635G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,049,495
App. No.
16/086,294
Granted
Jun 29, 2021
Kind
B2
Abstract

There is provided a system and method for processing and/or recognizing acoustic signals. The method comprises obtaining at least one pre-existing speech recognition model; adapting and/or training the at least one pre-existing speech recognition model incrementally when new, previously unseen, user-specific data is received, the data comprising input acoustic signals and/or user action demonstrations and/or semantic information about a meaning of the acoustic signals, wherein the at least one model is incrementally updated by associating new input acoustic signals with input semantic frames to enable recognition of changed input acoustic signals. The method further comprises adapting to a user's vocabulary over time by learning new words and/or removing words no longer being used by the user, generating a semantic frame from an input acoustic signal according to the at least one model, and mapping the semantic frame to a predetermined action.

Claims (65)

1. A method of processing acoustic signals, the method comprising:

obtaining at least one pre-existing speech recognition model;

adapting the at least one pre-existing speech recognition model to a user's vocabulary over time by one or more of:

learning new words; and

removing words no longer being used by the user,

wherein adapting the at least one pre-existing speech recognition model comprises:

receiving previously unseen user-specific data comprising one or more of:

acoustic signals;

user action demonstrations; and

semantic information about a meaning of an acoustic signal; and

incrementally updating the at least one pre-existing speech recognition model by associating acoustic signals with input semantic frames to enable recognition of new or changed input acoustic signals, wherein the incrementally updating comprises adapting one or more of a clustering layer and a latent variable layer, wherein the clustering layer uses one or more of:

a deep neural network (DNN);

a convolutional neural network (CNN);

a recurrent neural network (RNN) with long short term memory (LSTM) units; and

a recurrent neural network (RNN) with gated recurrent units (GRU);

generating a semantic frame from an input acoustic signal using one or more of the at least one pre-existing speech recognition models;

mapping the semantic frame to a predefined action; and

performing the predefined action when the mapping is successful.

2. The method of claim 1 , wherein said at least one pre-existing speech recognition model comprises a dictionary for the acoustic signals and a set of activations or weights for the dictionary.

3. The method of claim 1 , wherein the clustering layer is trained on one or more of:

a separate vocabulary speech recognition database;

the pre-existing speech recognition model; and

a database of known acoustic signals.

4. The method according to claim 1 , wherein in the latent variable layer or the clustering layer comprises one or more of a non-negative matrix factorization (NMF), histogram of acoustic co-occurrence (HAC), deep neural network (DNN), convolutional neural network (CNN), and recurrent neural network (RNN) that associates input features into the layer to the semantics are trained.

5. The method according to claim 1 , wherein the clustering layer and the latent variable layer are trained on one or more of:

a separate vocabulary speech recognition database in advance; and

the pre-existing speech recognition model or database of acoustic signals.

6. The method of claim 1 , wherein the latent variable layer is incrementally learned using automatic relevance determination (ARD) or incremental ARD.

7. The method according to claim 1 , wherein mapping the semantic frame to a predefined action comprises incrementally updating the latent variable layer.

8. The method according to claim 1 , wherein the semantic frames are generated from user actions performed on an alternate, non-vocal user interface.

9. The method according to claim 1 , wherein the semantic frames are generated from automatically analyzing text associated with input acoustic signals.

10. The method according to claim 1 , wherein semantic concepts of the semantic frames are relevant semantics that a user refers to when controlling or addressing a device or object by voice using a vocal user interface (VUI).

11. The method according to claim 1 , wherein semantic concepts of semantic frames are predefined and a vector is composed in which entries represent a presence or absence of an acoustic signal referring to one of the predefined semantic concepts.

12. The method of claim 1 , wherein, when a plurality of models is learned, the relevance of each model in the plurality of models is determined automatically.

13. The method of claim 12 , wherein, when a plurality of models is learned, each irrelevant model in the plurality of models is removed.

14. The method of claim 12 , wherein, when a plurality of models is learned, a required number of models is incrementally learned through incremental automatic relevance (IARD) detection.

15. The method of claim 12 , wherein, when a plurality of models is learned, the number of models is increased upon receiving a new acoustic signal and irrelevant models are removed using automatic relevance detection.

16. The method of claim 1 , wherein the incremental learning and automatic relevance determination are achieved through maximum a posteriori (MAP) estimation.

17. The method of claim 1 , wherein a forgetting factor is included in the clustering layer, latent variable layer and/or the automatic relevance determination.

18. The method of claim 1 , wherein incremental automatic relevance determination is achieved by imposing an inverse gamma prior on all relevance parameters.

19. The method of claim 1 , wherein incremental learning is achieved by imposing a gamma prior on the at least one model.

20. The method of claim 19 , wherein the incremental automatic relevance determination is achieved by imposing an exponential prior to the activation of the at least one model.

21. A non-transitory computer readable medium comprising computer executable instructions for performing the method of claim 1 .

22. A system for processing acoustic signals, the system comprising:

a processor for executing instructions; and

memory comprising computer executable instructions which when executed configure the system to:

obtain at least one pre-existing speech recognition model;

adapt the at least one pre-existing speech recognition model to a user's vocabulary over time by one or more of:

learning new words; and

removing words no longer being used by the user,

wherein adapting the at least one pre-existing speech recognition model comprises:

receiving previously unseen user-specific data comprising one or more of:

acoustic signals;

user action demonstrations; and

semantic information about a meaning of an acoustic signal; and

incrementally updating the at least one pre-existing speech recognition model by associating acoustic signals with input semantic frames to enable recognition of new or changed input acoustic signals, wherein the incrementally updating comprises adapting one or more of a clustering layer and a latent variable layer, wherein the clustering layer uses one or more of:

a deep neural network (DNN);

a convolutional neural network (CNN);

a recurrent neural network (RNN) with long short term memory (LSTM) units; and

a recurrent neural network (RNN) with gated recurrent units (GRU);

generate a semantic frame from an input acoustic signal using one or more of the at least one pre-existing speech recognition models;

map the semantic frame to a predefined action; and

perform the predefined action when the mapping is successful.

23. The system of claim 22 , wherein the system comprises a cloud-based device for performing cloud-based processing.

24. The system of claim 23 wherein the system further comprises an acoustic sensor for receiving acoustic signals.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2025
From: FLUENT.AI INC.
To: LALA, PROBAL
Reel/Frame 070651/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2021
From: TOMAR, VIKRANT; RENKENS, VINCENT P. G.; VAN HAMME, HUGO R. J. G.
To: FLUENT.AI INC.
Reel/Frame 056224/0708 →
Priority Claims (2)
GB 1604592 · Mar 18, 2016 · national
GB 1604594 · Mar 18, 2016 · national
Continuity (1)
Related Publication 20190108832A1 · Apr 11, 2019