IP Library › Granted Patent US 10,019,992
Granted Patent B2
US 10,019,992 · App. 14/754,457 · Granted Jul 10, 2018

Speech-controlled actions based on keywords and context thereof

Inventors: Jill Fain Lehman (Pittsburgh, PA); Samer Al Moubayed (Pittsburgh, PA)
Assignee: Disney Enterprises, Inc.
G10L15/22G10L15/19G10L2015/088G10L2015/223G10L2015/227G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,019,992
App. No.
14/754,457
Granted
Jul 10, 2018
Kind
B2
Abstract

A device includes a plurality of components, a memory having a keyword recognition module and a context recognition module, a microphone configured to receive an input speech spoken by a user, an analog-to-digital converter configured to convert the input speech from an analog form to a digital form and generate a digitized speech, and a processor. The processor is configured to detect, using the keyword recognition module, a keyword in the digitized speech, initiate, in response to detecting the keyword by the keyword recognition module, an action to be taken one of the plurality of components, wherein the keyword is associated with the action, determine, using the context recognition module, a context for the keyword, and execute the action if the context determined by the context recognition module indicates that the keyword is a command.

Claims (50)

1. A device comprising:

a plurality of components;

a memory including a keyword recognition module and a context recognition module;

a microphone configured to receive an input speech spoken by a user;

an analog-to-digital converter configured to convert the input speech from an analog form to a digital form and generate a digitized speech;

a processor configured to:

detect, using the keyword recognition module, a keyword in the digitized speech based on features extracted from the digitized speech;

initiate, in response to detecting the keyword by the keyword recognition module, an action to be taken by one of the plurality of components, wherein the keyword is defined as a command to take the action;

determine, using the context recognition module and after detecting the keyword, a context for the keyword based on additional features extracted from a portion of the digitized speech before and/or after the keyword, wherein the context is used to determine whether or not the keyword should be considered to be the command; and

execute the action if the context determined by the context recognition module indicates that the keyword should be considered to be the command.

2. The device of claim 1 , wherein the context recognition module utilizes a voice activity detector to determine the context.

3. The device of claim 1 , wherein the keyword recognition module is continuously listening for the keyword and the context recognition module is configured to begin listening when the input speech is received from the microphone.

4. The device of claim 1 , wherein the processor is configured to:

prior to determining the context of the keyword, receive one or more second inputs from the user; and

analyze the context of the keyword based on the one or more second inputs.

5. The device of claim 4 , wherein the one or more second inputs is a non-verbal input including a physical gesture.

6. The device of claim 4 , wherein the one or more second inputs are received from one of a motion sensor and a video camera.

7. The device of claim 1 , wherein the context recognition module determines that the keyword is in the digitized speech based on an indication received from the keyword recognition module.

8. The device of claim 1 , wherein the context of the keyword includes a location of the user.

9. The device of claim 1 , wherein the processor is further configured to:

display a result of executing the action on a display.

10. The device of claim 1 , wherein the processor is further configured to:

terminate the action if the context determined by the context recognition module indicates that the keyword should not be considered to be the command.

11. A method for speech recognition by a device having a microphone, a processor, and a memory including a keyword recognition module and a context recognition module, the method comprising:

detecting, using the keyword recognition module, a keyword in a digitized speech based on features extracted from the digitized speech;

initiating, in response to detecting the keyword by the keyword recognition module, an action to be taken by one of the plurality of components, wherein the keyword is defined as a command to take the action;

determining, using the context recognition module and after detecting the keyword, a context for the keyword based on additional features extracted from a portion of the digitized speech before and/or after the keyword, wherein the context is used to determine whether or not the keyword should be considered to be the command; and

executing the action if the context determined by the context recognition module indicates that the keyword should be considered to be the command.

12. The method of claim 11 , wherein the context recognition module utilizes a voice activity detector to determine the context.

13. The method of claim 11 , wherein the keyword recognition module is continuously listening for the keyword and the context recognition module is configured to begin listening when the input speech is received from the microphone.

14. The method of claim 11 , further comprising:

prior to determining the context of the keyword, receiving one or more second inputs from the user; and

analyzing the context of the keyword based on the one or more second inputs.

15. The method of claim 14 , wherein the one or more second inputs include a non-verbal input including a physical gesture.

16. The method of claim 14 , wherein the one or more second inputs are received from one of a motion sensor and a video camera.

17. The method of claim 11 , wherein the context recognition module determines that the keyword is in the digitized speech based on an indication received from the keyword recognition module.

18. The method of claim 11 , wherein the context of the keyword includes a location of a user.

19. The method of claim 11 further comprising:

terminating the action if the context determined by the context recognition module indicates that the keyword should not be considered to be the command.

20. A device comprising:

a plurality of components;

a memory including a keyword recognition module and a context recognition module;

a microphone configured to receive an input speech spoken by a user;

an analog-to-digital converter configured to convert the input speech from an analog form to a digital form and generate a digitized speech;

a processor configured to:

detect, using the keyword recognition module, a keyword in the digitized speech based on features extracted from the digitized speech;

initiate, in response to detecting the keyword by the keyword recognition module, an action to be taken by one of the plurality of components, wherein the keyword is defined as a command to take the action;

determine, using the context recognition module and after detecting the keyword, a context for the keyword based on additional features extracted from the digitized speech before and after the keyword, wherein the context is used to determine whether or not the keyword should be considered to be the command;

execute the action if the context determined by the context recognition module indicates that the keyword should be considered to be the command; and

terminate the action if the context determined by the context recognition module indicates that the keyword should not be considered to be the command.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2015
From: LEHMAN, JILL FAIN; AL MOUBAYED, SAMER
To: DISNEY ENTERPRISES, INC.
Reel/Frame 035932/0572 →
Continuity (1)
Related Publication 20160379633A1 · Dec 29, 2016
Cited By (1)
US 12,254,717