IP Library › Granted Patent US 11,423,879
Granted Patent B2
US 11,423,879 · App. 15/653,486 · Granted Aug 23, 2022

Verbal cues for high-speed control of a voice-enabled device

Inventor: William Valentine Zajac, III (La Crescenta, CA)
Assignee: Disney Enterprises, Inc.
G10L15/02A63F13/215A63F13/424G06F3/167G10L2015/025G10L2015/027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,423,879
App. No.
15/653,486
Filed
Jul 18, 2017
Granted
Aug 23, 2022
Kind
B2
Art Unit
2658
USPC
704/254
Abstract

A technique for controlling a voice-enabled device using voice commands includes receiving an audio signal that is generated in response to a verbal utterance, generating a verbal utterance indicator for the verbal utterance based on the audio signal, selecting a first command for a voice-controlled application residing within the voice-enabled device based on the verbal utterance indicator, and transmitting the first command to the voice-controlled application as an input.

Claims (53)

1. A method for controlling a voice-enabled device using voice commands, the method comprising:

receiving an audio signal that is generated in response to a verbal utterance;

generating a verbal utterance indicator for the verbal utterance based on the audio signal;

selecting a first application included in a plurality of applications currently executing on the voice-enabled device that is to receive one or more commands;

selecting a first command to transmit to the first application based on the verbal utterance indicator comprising:

upon determining that the verbal utterance indicator comprises a phonetic fragment, performing the steps of:

performing a first mapping from the verbal utterance indicator to a first complete word via a first mapping table associated with the first application; and

performing a second mapping from the first complete word to the first command via a second mapping table associated with the first application; and

upon determining that the verbal utterance indicator comprises the first complete word, performing the second mapping from the first complete word to the first command via the second mapping table without performing the first mapping from the verbal utterance indicator to the first complete word via the first mapping table; and

transmitting the first command to the first application as an input.

2. The method of claim 1 , wherein generating a verbal utterance indicator comprises transmitting the audio signal to a speech recognition application for processing.

3. The method of claim 1 , wherein the verbal utterance indicator comprises a textual representation of the verbal utterance.

4. The method of claim 1 , wherein selecting the first application included in the plurality of applications comprises at least one of selecting an application included in the plurality of applications that is causing a visual output to be displayed or selecting an application included in the plurality of applications that is causing an audio output to be generated.

5. The method of claim 1 , wherein the verbal utterance consists of only a single phonetic fragment.

6. The method of claim 1 , wherein the phonetic fragment consists of a single consonant and a single vowel.

7. The method of claim 1 , wherein the first application is further selected due to at least one of a current location of a user, a current time, a current day, the first application being associated with a home automation system, or the first application currently controlling a home automation device.

8. The method of claim 1 , wherein the first mapping table comprises a plurality of word mappings, each word mapping included in the plurality of word mappings comprising a mapping from a particular phonetic fragment to a particular complete word, the plurality of word mappings including a word mapping for the first complete word.

9. The method of claim 1 , wherein the second mapping table comprises a plurality of command mappings, each command mapping included in the plurality of command mappings comprising a mapping from a particular complete word to a particular command, the plurality of command mappings including a command mapping for the first command.

10. The method of claim 1 , wherein selecting the first application included in the plurality of applications comprises selecting the first application due to the first application currently outputting video or audio.

11. A non-transitory computer-readable storage medium including instructions that, when executed by a processor, cause the processor to perform the steps of:

receiving an audio signal that is generated in response to a verbal utterance;

generating a verbal utterance indicator for the verbal utterance based on the audio signal;

selecting a first application included in a plurality of applications currently executing on the voice-enabled device that is to receive one or more commands;

selecting a first command to transmit to the first application based on the verbal utterance indicator comprising:

upon determining that the verbal utterance indicator comprises a phonetic fragment, performing the steps of:

performing a first mapping from the verbal utterance indicator to a first complete word via a first mapping table associated with the first application; and

performing a second mapping from the first complete word to the first command via a second mapping table associated with the first application; and

upon determining that the verbal utterance indicator comprises the first complete word, performing the second mapping from the first complete word to the first command via the second mapping table without performing the first mapping from the verbal utterance indicator to the first complete word via the first mapping table; and

transmitting the first command to the first application as an input.

12. The non-transitory computer-readable storage medium of claim 11 , wherein the verbal utterance consists of only a single consonant and a single vowel.

13. The non-transitory computer-readable storage medium of claim 11 , further comprising:

after receiving the audio signal, receiving an additional signal that is generated in response to an additional verbal utterance, wherein the addition verbal utterance is different than the verbal utterance; and

transmitting an additional command to the first application based on the additional verbal utterance, wherein the additional command acts to cancel the first command transmitted to the first application.

14. The non-transitory computer-readable storage medium of claim 11 , further comprising:

after receiving the audio signal, receiving an additional signal that is generated by the microphone in response to an additional verbal utterance, wherein the additional verbal utterance is different than the verbal utterance; and

with the processing unit, based on the additional verbal utterance, transmitting the first command to the first application.

15. The non-transitory computer-readable storage medium of claim 11 , wherein generating a verbal utterance indicator comprises transmitting the audio signal to a speech recognition application for processing.

16. The non-transitory computer-readable storage medium of claim 11 , wherein the verbal utterance indicator comprises a textual representation of the verbal utterance.

17. A system, comprising:

a microphone;

a memory storing a speech recognition application and a verbal utterance translator; and

one or more processors that are coupled to the memory and the microphone, and when executing the speech recognition application or the verbal utterance translator, are configured to:

receive an audio signal that is generated in response to a verbal utterance;

generate a verbal utterance indicator for the verbal utterance based on the audio signal;

select a first application included in a plurality of applications currently executing on the voice-enabled device that is to receive one or more commands;

select a first command to transmit to the first application based on the verbal utterance indicator comprising:

upon determining that the verbal utterance indicator comprises a phonetic fragment, performing the steps of:

performing a first mapping from the verbal utterance indicator to a first complete word via a first mapping table associated with the first application; and

performing a second mapping from the first complete word to the first command via a second mapping table associated with the first application; and

upon determining that the verbal utterance indicator comprises the first complete word, performing the second mapping from the first complete word to the first command via the second mapping table without performing the first mapping from the verbal utterance indicator to the first complete word via the first mapping table; and

transmit the first command to the first application as an input.

18. The system of claim 17 , wherein the speech recognition application is running on the processor.

19. The system of claim 17 , wherein the speech recognition application is running on a remote processor that is communicatively connected to the computing device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2017
From: ZAJAC, WILLIAM VALENTINE, III
To: DISNEY ENTERPRISES, INC.
Reel/Frame 043037/0741 →
Continuity (1)
Related Publication 20190027131A1 · Jan 24, 2019