IP Library Granted Patent US 9,305,548
Granted Patent B2
US 9,305,548 · App. 14/083,061 · Granted Apr 5, 2016

System and method for an integrated, multi-modal, multi-device natural language voice services environment

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,305,548
App. No.
14/083,061
Granted
Apr 5, 2016
Kind
B2
Abstract

A system and method for an integrated, multi-modal, multi-device natural language voice services environment may be provided. In particular, the environment may include a plurality of voice-enabled devices each having intent determination capabilities for processing multi-modal natural language inputs in addition to knowledge of the intent determination capabilities of other devices in the environment. Further, the environment may be arranged in a centralized manner, a distributed peer-to-peer manner, or various combinations thereof. As such, the various devices may cooperate to determine intent of multi-modal natural language inputs, and commands, queries, or other requests may be routed to one or more of the devices best suited to take action in response thereto.

Claims (114)

1. A method of processing natural language utterances, the method being implemented by a first device that comprises one or more physical processors executing one or more computer program instructions which, when executed, perform the method, the method comprising:

receiving, by the first device, a natural language utterance spoken by a user;

performing, by the first device, speech recognition to determine one or more words of the natural language utterance;

determining, by the first device, based on the one or more words, a first prediction of an intent of the user and a domain relating to the natural language utterance;

determining, by the first device, processing capabilities associated with a plurality of devices, wherein the plurality of devices comprises a second device;

selecting, by the first device, based on the processing capabilities associated with the second device and the domain, the second device to determine the second prediction;

transmitting, by the first device, the natural language utterance to the second device based on the selection;

receiving, from the second device by the first device, a second prediction of the intent of the user; and

determining, by the first device, the intent of the user based on the first prediction and the second prediction.

2. The method of claim 1 , wherein the processing capabilities associated with the plurality of devices comprise processing power, storage resources, information, or services available at individual ones of the plurality of devices.

3. The method of claim 1 , wherein the plurality of devices are associated with different domains, the second device is associated with the domain, and the different domains comprise the domain.

4. The method of claim 1 , further comprising:

determining, by the first device, whether a prediction of the intent of the user is to be obtained from at least one other device,

wherein transmitting the natural language utterance comprises transmitting the natural language utterance to the second device based on a determination that a prediction of the intent of the user is to be obtained from at least one other device.

5. The method of claim 4 , further comprising:

determining, by the first device, a confidence level associated with the first prediction; and

determining, by the first device, whether the confident level satisfies a threshold level of confidence relating to intent prediction accuracy,

wherein the determination that a prediction of the intent of the user is to be obtained from at least one other device is based on a determination that the confidence level does not satisfy the threshold level of confidence.

6. A method of processing natural language utterances, the method being implemented by a first device that comprises one or more physical processors executing one or more computer program instructions which, when executed, perform the method, the method comprising:

receiving, by the first device, a natural language utterance spoken by a user;

performing, by the first device, speech recognition to determine one or more words of the natural language utterance;

determining, by the first device, based on the one or more words, a first prediction of an intent of the user;

transmitting, by the first device, the natural language utterance to a second device;

receiving, from the second device by the first device, a second prediction of the intent of the user;

transmitting, by the first device, the natural language utterance to a third device;

receiving, from the third device by the first device, a third prediction of the intent of the user;

receiving, from the second device by the first device, information regarding a second confidence level associated with the second prediction;

receiving, from the third device by the first device, information regarding a third confidence level associated with the third prediction;

comparing, by the first device, the second and third confidence levels with one another; and

determining, by the first device, the intent of the user based on the first prediction, the second prediction, the third prediction, and the comparison.

7. The method of claim 6 , further comprising:

determining, by the first device, processing capabilities associated with the second device and processing capabilities associated with the third device,

wherein determining the intent comprises determining the intent based on the comparison and the processing capabilities associated with the second and third devices.

8. A method of processing natural language utterances, the method being implemented by a first device that comprises one or more physical processors executing one or more computer program instructions which, when executed, perform the method, the method comprising:

receiving, by the first device, a natural language utterance spoken by a user;

performing, by the first device, speech recognition to determine one or more words of the natural language utterance;

determining, by the first device, based on the one or more words, a first prediction of an intent of the user;

determining, by the first device, a first confidence level associated with the first prediction;

transmitting, by the first device, the natural language utterance to a second device;

receiving, from the second device by the first device, a second prediction of the intent of the user;

receiving, from the second device by the first device, information regarding a second confidence level associated with the second prediction;

determining, by the first device, a highest confidence level among the first confidence level, the second confidence level, and one or more other confidence levels associated with one or more other predictions of the intent of the user, wherein the one or more other predictions are received from one or more other devices; and

determining, by the first device, the intent of the user based on the first prediction and the second prediction, wherein determining the intent comprises selecting, from the first prediction, the second prediction, and the one or more other predictions, a prediction of the intent of the user that is associated with the highest confidence level.

9. A method of processing natural language utterances, the method being implemented by a first device that comprises one or more physical processors executing one or more computer program instructions which, when executed, perform the method, the method comprising:

receiving, by the first device, a natural language utterance spoken by a user, wherein the first device comprises an input device at which the natural language utterance is initially received from the user;

performing, by the first device, speech recognition to determine one or more words of the natural language utterance;

determining, by the first device, based on the one or more words, a first prediction of an intent of the user;

determining, by the first device, whether a third device is available to manage determination of the intent of the user;

transmitting, by the first device, the natural language utterance to a second device, wherein transmitting the natural language utterance comprises transmitting the natural language utterance to the second device based on a determination that the third device is not available to manage determination of the intent of the user;

receiving, from the second device by the first device, a second prediction of the intent of the user; and

determining, by the first device, the intent of the user based on the first prediction and the second prediction.

10. The method of claim 9 , further comprising:

transmitting, by the first device, a request to a plurality of devices for information regarding processing capabilities associated with the plurality of devices based on a determination that the third device is not available to manage determination of the intent, wherein the plurality of devices comprise the second device;

receiving, by the first device, information regarding the processing capabilities; and

selecting, by the first device based on the processing capabilities, the second device to determine the second prediction,

wherein transmitting the natural language utterance comprises transmitting the natural language utterance to the second device based on the selection.

11. A system for processing natural language utterances, the system comprising:

a first device having one or more physical processors programmed to execute one or more computer program instructions which, when executed, cause the one or more physical processors to:

receive a natural language utterance spoken by a user;

perform speech recognition to determine one or more words of the natural language utterance;

determine, based on the one or more words, a first prediction of an intent of the user and a domain relating to the natural language utterance;

determine processing capabilities associated with a plurality of devices, wherein the plurality of devices comprises the second device;

select, based on the processing capabilities associated with the second device and the domain, the second device to determine the second prediction;

transmit the natural language utterance to a second device based on the selection;

receive, from the second device, a second prediction of the intent of the user; and

determine the intent of the user based on the first prediction and the second prediction.

12. The system of claim 11 , wherein the processing capabilities associated with the plurality of devices comprise processing power, storage resources, information, or services available at individual ones of the plurality of devices.

13. The system of claim 11 , wherein the plurality of devices are associated with different domains, the second device is associated with the domain, and the different domains comprise the domain.

14. The system of claim 11 , wherein the one or more physical processors are further caused to:

determine whether a prediction of the intent of the user is to be obtained from at least one other device,

wherein transmitting the natural language utterance comprises transmitting the natural language utterance to the second device based on a determination that a prediction of the intent of the user is to be obtained from at least one other device.

15. The system of claim 14 , wherein the one or more physical processors are further caused to:

determine a confidence level associated with the first prediction; and

determine whether the confident level satisfies a threshold level of confidence relating to intent prediction accuracy,

wherein the determination that a prediction of the intent of the user is to be obtained from at least one other device is based on a determination that the confidence level does not satisfy the threshold level of confidence.

16. A system for processing natural language utterances, the system comprising:

a first device having one or more physical processors programmed to execute one or more computer program instructions which, when executed, cause the one or more physical processors to:

receive a natural language utterance spoken by a user;

perform speech recognition to determine one or more words of the natural language utterance;

determine, based on the one or more words, a first prediction of an intent of the user;

transmit the natural language utterance to a second device;

receive, from the second device, a second prediction of the intent of the user;

transmit the natural language utterance to a third device;

receive, from the third device, a third prediction of the intent of the user;

receive, from the second device, information regarding a second confidence level associated with the second prediction;

receive, from the third device, information regarding a third confidence level associated with the third prediction;

compare, by the first device, the second and third confidence levels with one another; and

determine the intent of the user based on the first prediction, the second prediction, the third prediction, and the comparison.

17. The system of claim 16 , wherein the one or more physical processors are further caused to:

determine processing capabilities associated with the second device and processing capabilities associated with the third device,

wherein determining the intent comprises determining the intent based on the comparison and the processing capabilities associated with the second and third devices.

18. A system for processing natural language utterances, the system comprising:

a first device having one or more physical processors programmed to execute one or more computer program instructions which, when executed, cause the one or more physical processors to:

receive a natural language utterance spoken by a user;

perform speech recognition to determine one or more words of the natural language utterance;

determine, based on the one or more words, a first prediction of an intent of the user;

determine a first confidence level associated with the first prediction;

transmit the natural language utterance to a second device;

receive, from the second device, a second prediction of the intent of the user;

receive, from the second device, information regarding a second confidence level associated with the second prediction;

determine a highest confidence level among the first confidence level, the second confidence level, and one or more other confidence levels associated with one or more other predictions of the intent of the user, wherein the one or more other predictions are received from one or more other devices; and

determine the intent of the user based on the first prediction and the second prediction, wherein determining the intent comprises selecting, from the first prediction, the second prediction, and the one or more other predictions, a prediction of the intent of the user that is associated with the highest confidence level.

19. A system for processing natural language utterances, the system comprising:

a first device having one or more physical processors programmed to execute one or more computer program instructions which, when executed, cause the one or more physical processors to:

receive a natural language utterance spoken by a user, wherein the first device comprises an input device at which the natural language utterance is initially received from the user;

perform speech recognition to determine one or more words of the natural language utterance;

determine, based on the one or more words, a first prediction of an intent of the user;

determine whether a third device is available to manage determination of the intent of the user;

transmit a request to a plurality of devices for information regarding processing capabilities associated with the plurality of devices based on a determination that the third device is not available to manage determination of the intent, wherein the plurality of devices comprise the second device;

receive information regarding the processing capabilities;

select, based on the processing capabilities, the second device to determine the second prediction;

transmit the natural language utterance to a second device, wherein transmitting the natural language utterance comprises transmitting the natural language utterance to the second device based on the selection;

receive, from the second device, a second prediction of the intent of the user; and

determine the intent of the user based on the first prediction and the second prediction.

Assignments (10)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR'S NAME AND ASSIGNEE'S NAME PREVIOUSLY RECORDED ON REEL 051582 FRAME 0067. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNOR'S INTEREST. Recorded Oct 30, 2020
From: VOICEBOX TECHNOLOGIES CORPORATION
To: VB ASSETS, LLC
Reel/Frame 054260/0811 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2020
From: ORACLE INTERNATIONAL CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 051582/0067 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2019
From: VB ASSETS, LLC
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 049506/0531 →
RELEASE OF SECURITY INTEREST Recorded Jun 13, 2019
From: DELPHI ASSET MANAGEMENT CORPORATION
To: VB ASSETS, LLC
Reel/Frame 049459/0596 →
SECURITY INTEREST Recorded Apr 12, 2019
From: VB ASSETS, LLC
To: DELPHI ASSET MANAGEMENT CORPORATION
Reel/Frame 048872/0831 →
NUNC PRO TUNC ASSIGNMENT Recorded Jul 25, 2018
From: VOICEBOX TECHNOLOGIES CORPORATION
To: VB ASSETS, LLC
Reel/Frame 046456/0128 →
RELEASE OF SECURITY INTEREST Recorded Apr 5, 2018
From: ORIX GROWTH CAPITAL, LLC
To: VOICEBOX TECHNOLOGIES CORPORATION
Reel/Frame 045581/0630 →
SECURITY INTEREST Recorded Dec 22, 2017
From: VOICEBOX TECHNOLOGIES CORPORATION
To: ORIX GROWTH CAPITAL, LLC
Reel/Frame 044949/0948 →
MERGER Recorded May 1, 2014
From: VOICEBOX TECHNOLOGIES, INC.
To: VOICEBOX TECHNOLOGIES CORPORATION
Reel/Frame 032796/0671 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2014
From: KENNEWICK, ROBERT A.; WEIDER, CHRIS
To: VOICEBOX TECHNOLOGIES, INC.
Reel/Frame 032792/0285 →