IP Library › Granted Patent US 11,544,463
Granted Patent B2
US 11,544,463 · App. 16/407,766 · Granted Jan 3, 2023

Time asynchronous spoken intent detection

Inventors: Munir Georges (Kehl, DE); Wenda Chen (Singapore, SG); Tobias Bocklet (Munich, DE); Jonathan Huang (Pleasanton, CA)
Assignee: Intel Corporation
G06F40/289G06F17/18G06N3/0454G06N3/08G06N20/20G10L15/142G10L15/16G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,544,463
App. No.
16/407,766
Granted
Jan 3, 2023
Kind
B2
Abstract

An embodiment of a spoken intent detection device includes technology to detect a phrase in an electronic representation of an audio stream based on a pre-defined vocabulary, associate a time stamp with the detected phrase, and classify a spoken intent based on a sequence of detected phrases and the respective associated time stamps. Other embodiments are disclosed and claimed.

Claims (58)

1. An electronic system, comprising:

an electronic memory to store an electronic representation of an audio stream;

an electronic processor coupled to the memory; and

logic circuitry coupled to the electronic processor and the electronic memory, the logic circuitry to:

electronically detect a phrase in the stored electronic representation of the audio stream based on a pre-defined vocabulary,

electronically compute a time stamp for the detected phrase which is relative to an adjacent previously detected phrase,

electronically associate the time stamp with the detected phrase, and

electronically classify a spoken intent based on a sequence of detected phrases and the respective associated time stamps.

2. The system of claim 1 , wherein the logic circuitry is further to:

electronically monitor a continuous audio stream; and

electronically detect the phrase in the continuous audio stream.

3. The system of claim 1 , wherein the logic circuitry comprises:

an always-on phrase spotter circuit with a first neural network with an acoustic model and a hidden Markov model to detect the phrase in the audio stream; and

a selectively powered intent classification circuit coupled to the always-on phrase spotter circuit, wherein the selectively powered intent classification circuit is to power up in response to a signal from the always-on phrase spotter circuit to electronically classify the spoken intent.

4. The system of claim 3 , wherein the acoustic model is further configured to:

automatically add time stamp information to text data for the detected phrase.

5. The system of claim 3 , wherein the selectively powered intent classification circuit comprises:

a second neural network trained to return a probability for each of two or more intent classifications based on detected phrases and time stamps respectively associated with the detected phrases as input features to the second neural network.

6. The system of claim 5 , wherein the logic circuitry is further to:

classify the spoken intent in accordance with a highest probability of the two or more intent classifications.

7. The system of claim 5 , wherein the always-on phrase spotter circuit is further to:

asynchronously signal the intent classification circuit to power up and trigger the second neural network when a sequence of detected phrases is ready for classification.

8. A method of detecting spoken intent, comprising:

electronically detecting a phrase in an electronic representation of an audio stream based on a pre-defined vocabulary;

electronically associating a time stamp with the detected phrase which is relative to an adjacent previously detected phrase; and

electronically classifying a spoken intent based on a sequence of detected phrases and the respective associated time stamps.

9. The method of claim 8 , further comprising:

electronically monitoring a continuous audio stream;

electronically detecting the phrase in an electronic representation of the continuous audio stream; and

electronically computing the time stamp for the detected phrase which is relative to the adjacent previously detected phrase.

10. The method of claim 8 , further comprising:

detecting the phrase in the audio stream with a first neural network which includes both an acoustic model and a hidden Markov model.

11. The method of claim 10 , further comprising:

automatically adding time stamp information to text data for the detected phrase by the acoustic model.

12. The method of claim 8 , further comprising:

returning a probability for each of two or more intent classifications from a second neural network based on detected phrases and time stamps respectively associated with the detected phrases as input features to the second neural network.

13. The method of claim 12 , further comprising:

classifying the spoken intent in accordance with a highest probability of the two or more intent classifications.

14. The method of claim 12 , further comprising:

asynchronously triggering the second neural network when a sequence of detected phrases is ready for classification.

15. At least one non-transitory machine readable medium comprising a plurality of instructions that, in response to being executed on a computing device, cause the computing device to:

electronically detect a phrase in an electronic representation of an audio stream based on a pre-defined vocabulary;

electronically compute a time stamp for the detected phrase which is relative to an adjacent previously detected phrase;

electronically associate the time stamp with the detected phrase; and

electronically classify a spoken intent based on a sequence of detected phrases and the respective associated time stamps.

16. The machine readable medium of claim 15 , comprising a plurality of further instructions that, in response to being executed on the computing device, cause the computing device to:

electronically monitor a continuous audio stream; and

electronically detect the phrase in an electronic representation of the continuous audio stream.

17. The machine readable medium of claim 15 , comprising a plurality of further instructions that, in response to being executed on the computing device, cause the computing device to:

detect the phrase in the audio stream with a first neural network which includes both an acoustic model and a hidden Markov model.

18. The machine readable medium of claim 17 , comprising a plurality of further instructions that, in response to being executed on the computing device, cause the computing device to:

automatically add time stamp information to text data for the detected phrase by the acoustic model.

19. The machine readable medium of claim 15 , comprising a plurality of further instructions that, in response to being executed on the computing device, cause the computing device to:

return a probability for each of two or more intent classifications from a second neural network based on detected phrases and time stamps respectively associated with the detected phrases as input features to the second neural network.

20. The machine readable medium of claim 19 , comprising a plurality of further instructions that, in response to being executed on the computing device, cause the computing device to:

classify the spoken intent in accordance with a highest probability of the two or more intent classifications.

21. The machine readable medium of claim 19 , comprising a plurality of further instructions that, in response to being executed on the computing device, cause the computing device to:

asynchronously trigger the second neural network when a sequence of detected phrases is ready for classification.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 10, 2019
From: GEORGES, MUNIR; CHEN, WENDA; BOCKLET, TOBIAS; HUANG, JONATHAN
To: INTEL CORPORATION
Reel/Frame 049152/0426 →
Continuity (1)
Related Publication 20190266240A1 · Aug 29, 2019