IP Library Granted Patent US 10,089,984
Granted Patent B2
US 10,089,984 · App. 15/632,713 · Granted Oct 2, 2018

System and method for an integrated, multi-modal, multi-device natural language voice services environment

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,089,984
App. No.
15/632,713
Granted
Oct 2, 2018
Kind
B2
Abstract

A system and method for an integrated, multi-modal, multi-device natural language voice services environment may be provided. In particular, the environment may include a plurality of voice-enabled devices each having intent determination capabilities for processing multi-modal natural language inputs in addition to knowledge of the intent determination capabilities of other devices in the environment. Further, the environment may be arranged in a centralized manner, a distributed peer-to-peer manner, or various combinations thereof. As such, the various devices may cooperate to determine intent of multi-modal natural language inputs, and commands, queries, or other requests may be routed to one or more of the devices best suited to take action in response thereto.

Claims (37)

1. A method of providing an integrated multi-modal, natural language voice services environment comprising one or more of an input device that receives a multi-modal natural language input comprising at least a natural language utterance and a non-voice input related to the natural language utterance, a first device, or one or more secondary devices, the method being implemented in the first device having one or more physical processors programmed with computer program instructions that, when executed by the one or more physical processors, program the first device to perform the method, the method comprising:

obtaining, by the first device from the input device, the multi-modal natural language input;

transcribing, by the first device, the natural language utterance;

determining, by the first device, a preliminary intent prediction of the multi-modal natural language input based on the transcribed utterance and the non-voice input; and

invoking, by the first device, at least one action at one or more of the input device, the first device, or the one or more secondary devices based on the preliminary intent prediction.

2. The method of claim 1 , wherein invoking the at least one action at one or more of the input device, the first device, or the one or more secondary devices comprises transmitting a request related to the multi-modal natural language input based on the preliminary intent prediction.

3. The method of claim 1 , wherein the one or more secondary devices include at least a second device, the method further comprising:

transmitting, by the first device, the multi-modal natural language input to the second device;

receiving, by the first device from the second device, a second intent prediction of the multi-modal natural language input; and

determining, by the first device, an intent of the multi-modal natural language input based on the preliminary intent prediction and the second intent prediction, wherein the at least one action is invoked based on the determined intent.

4. The method of claim 3 , the method further comprising: determining, by the first device, processing capabilities associated with the one or more secondary devices; and selecting, by the first device, based on the processing capabilities associated with the one or more secondary devices, the second device to make the second intent prediction of the multi-modal natural language input.

5. The method of claim 4 , the method further comprising: maintaining, by the first device, a constellation model that describes natural language resources, dynamic states, and intent determination capabilities associated with the input device and the one or more secondary devices, wherein the processing capabilities associated with the one or more secondary devices are determined based on the constellation model.

6. The method of claim 5 , wherein the intent determination capabilities for a given one of the input device, the first device, or the one or more secondary devices are based on at least one of processing power, storage resources, natural language processing capabilities, or local knowledge.

7. The method of claim 3 , the method further comprising: determining, by the first device, a domain relating to the multi-modal natural language input; and selecting, by the first device, based on the domain, the second device to make the second intent prediction of the multi-modal natural language input.

8. The method of claim 7 , wherein the one or more secondary devices are associated with different domains, the second device is associated with the domain, and the different domains comprise the domain.

9. The method of claim 1 , the method further comprising:

communicating, by the first device, the multi-modal natural language input to each of the one or more secondary devices, wherein each of the one or more secondary devices determine an intent of the multi-modal natural language input received at the input device using local intent determination capabilities;

receiving, by the first device, the intent determined by each of the secondary devices; and

arbitrating, by the first device, among the intent determinations of the secondary devices to determine the intent of the multi-modal natural input, wherein the at least one action is invoked based on the arbitrated intent determinations.

10. The method of claim 1 , wherein the input device initially received the multi-modal natural language input.

11. A system for processing a multi-modal natural language input, the system comprising:

an input device that receives a multi-modal natural language input comprising at least a natural language utterance and a non-voice input related to the natural language utterance; one or more secondary devices; and

a first device having one or more physical processors programmed with computer program instructions that, when executed by the one or more physical processors, program the first device to:

obtain, from the input device, the multi-modal natural language input; transcribe the natural language utterance;

determine a preliminary intent prediction of the multi-modal natural language input based on the transcribed utterance and the non-voice input; and

invoke at least one action at one or more of the input device, the first device, or the one or more secondary devices based on the preliminary intent prediction.

12. The system of claim 11 , wherein to invoke the at least one action at one or more of the input device, the first device, or the one or more secondary devices, the first device is further programmed to: transmit a request related to the multi-modal natural language input based on the preliminary intent prediction.

13. The system of claim 11 , wherein the one or more secondary devices include at least a second device, and wherein the first device is further programmed to: transmit the multi-modal natural language input to the second device; receive, from the second device, a second intent prediction of the multi-modal natural language input; and determine an intent of the multi-modal natural language input based on the preliminary intent prediction and the second intent prediction, wherein the at least one action is invoked based on the determined intent.

14. The system of claim 13 , wherein the first device is further programmed to: determine processing capabilities associated with the one or more secondary devices; and select based on the processing capabilities associated with the one or more secondary devices, the second device to make the second intent prediction of the multi-modal natural language input.

15. The system of claim 14 , wherein the first device is further programmed to: maintain a constellation model that describes natural language resources, dynamic states, and intent determination capabilities associated with the input device and the one or more secondary devices, wherein the processing capabilities associated with the one or more secondary devices are determined based on the constellation model.

16. The system of claim 15 , wherein the intent determination capabilities for a given one of the input device, the first device, or the one or more secondary devices are based on at least one of processing power, storage resources, natural language processing capabilities, or local knowledge.

17. The system of claim 13 , wherein the first device is further programmed to: determine a domain relating to the multi-modal natural language input; and select, based on the domain, the second device to make the second intent prediction of the multi-modal natural language input.

18. The system of claim 17 , wherein the one or more secondary devices are associated with different domains, the second device is associated with the domain, and the different domains comprise the domain.

19. The system of claim 11 , wherein the first device is further programmed to:

communicate the multi-modal natural language input to each of the one or more secondary devices, wherein each of the one or more secondary devices determine an intent of the multi-modal natural language input received at the input device using local intent determination capabilities;

receive the intent determined by each of the secondary devices; and arbitrate among the intent determinations of the secondary devices to determine the intent of the multi-modal natural input, wherein the at least one action is invoked based on the arbitrated intent determinations.

20. The system of claim 11 , wherein the input device initially received the multi-modal natural language input.

Assignments (10)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR'S NAME AND ASSIGNEE'S NAME PREVIOUSLY RECORDED ON REEL 051582 FRAME 0067. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNOR'S INTEREST. Recorded Oct 30, 2020
From: VOICEBOX TECHNOLOGIES CORPORATION
To: VB ASSETS, LLC
Reel/Frame 054260/0811 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2020
From: ORACLE INTERNATIONAL CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 051582/0067 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2019
From: VB ASSETS, LLC
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 049506/0531 →
RELEASE OF SECURITY INTEREST Recorded Jun 13, 2019
From: DELPHI ASSET MANAGEMENT CORPORATION
To: VB ASSETS, LLC
Reel/Frame 049459/0596 →
SECURITY INTEREST Recorded Apr 12, 2019
From: VB ASSETS, LLC
To: DELPHI ASSET MANAGEMENT CORPORATION
Reel/Frame 048872/0831 →
NUNC PRO TUNC ASSIGNMENT Recorded Jul 25, 2018
From: VOICEBOX TECHNOLOGIES CORPORATION
To: VB ASSETS, LLC
Reel/Frame 046456/0128 →
RELEASE OF SECURITY INTEREST Recorded Apr 5, 2018
From: ORIX GROWTH CAPITAL, LLC
To: VOICEBOX TECHNOLOGIES CORPORATION
Reel/Frame 045581/0630 →
SECURITY INTEREST Recorded Dec 22, 2017
From: VOICEBOX TECHNOLOGIES CORPORATION
To: ORIX GROWTH CAPITAL, LLC
Reel/Frame 044949/0948 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2017
From: KENNEWICK, ROBERT A.; WEIDER, CHRIS
To: VOICEBOX TECHNOLOGIES, INC.
Reel/Frame 042813/0156 →
MERGER Recorded Jun 26, 2017
From: VOICEBOX TECHNOLOGIES, INC.
To: VOICEBOX TECHNOLOGIES CORPORATION
Reel/Frame 042992/0079 →