IP Library Granted Patent US 9,105,266
Granted Patent B2
US 9,105,266 · App. 14/278,645 · Granted Aug 11, 2015

System and method for processing multi-modal device interactions in a natural language voice services environment

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,105,266
App. No.
14/278,645
Granted
Aug 11, 2015
Kind
B2
Abstract

A system and method for processing multi-modal device interactions in a natural language voice services environment may be provided. In particular, one or more multi-modal device interactions may be received in a natural language voice services environment that includes one or more electronic devices. The multi-modal device interactions may include a non-voice interaction with at least one of the electronic devices or an application associated therewith, and may further include a natural language utterance relating to the non-voice interaction. Context relating to the non-voice interaction and the natural language utterance may be extracted and combined to determine an intent of the multi-modal device interaction, and a request may then be routed to one or more of the electronic devices based on the determined intent of the multi-modal device interaction.

Claims (58)

1. A method for facilitating natural language processing of user inputs via multiple input modes where each user input alone may be insufficient to completely and/or accurately determine a user request intended by a user, the method being implemented by a computer system that includes one or more physical processors executing computer program instructions which, when executed, perform the method, the method comprising:

receiving, at the computer system, a first user input of a user from a first input device via a first input mode, wherein the first user input is generated responsive to the user interacting with the first input device in a manner corresponding to the first input mode to provide the first user input;

receiving, at the computer system, a second user input of the user from a second input device via a second input mode, wherein the second user input is generated responsive to the user interacting with the second input device in a manner corresponding to the second input mode to provide the second user input, wherein the first user input and the second user input are related to one another, and wherein one of the first user input or the second user input comprises a voice input received from at least one of the first input device or the second input device via a voice input mode, and the other one of the first user input or the second user input comprises a non-voice input received from at least one of the first input device or the second input device via a non-voice input mode;

determining, by the computer system, based on the second user input, context information for interpreting the first user input, wherein the context information identifies a first item of a first item type;

determining, by the computer system, further context information based on the first user input, wherein the further context information identifies a second item of a second item type that is related to the first item of the first item type;

generating, by the computer system, a query based on the context information and the further context information to obtain one or more intermediary results, wherein the generated query comprises a query related to the second item of the second item type;

determining, by the computer system, a user request based on the one or more intermediary results;

providing, by the computer system, a response to the user request; and

providing, by the computer system, based on at least one of the context information for interpreting the first user input or the further context information, an advertisement for presentation to the user.

2. The method of claim 1 , further comprising:

providing, by the computer system, the context information for interpreting the first user input to an advertiser system; and

obtaining, by the computer system, the advertisement from the advertiser system responsive to providing the context information to the advertiser system,

wherein providing the advertisement comprises providing the advertisement obtained from the advertiser system.

3. The method of claim 1 , wherein the context information for interpreting the first user input is used as input for selecting the advertisement, and wherein providing the advertisement comprises providing the selected advertisement.

4. The method of claim 1 , wherein information about the user request is used as input for selecting the advertisement, and wherein providing the advertisement comprises providing the selected advertisement.

5. The method of claim 1 , wherein the first user input comprises the voice input received via the voice input mode, and the second user input comprises the non-voice input received via the non-voice input mode,

wherein determining the context information comprises determining, based on the non-voice input, the context information for interpreting the voice input,

wherein determining the user request comprises determining the user request based on the voice input and the context information for interpreting the voice input, and

wherein providing the advertisement comprises providing the advertisement based on the context information for interpreting the voice input.

6. The method of claim 5 , wherein the context information for interpreting the voice input is used as input for selecting the advertisement, and wherein providing the advertisement comprises providing the selected advertisement.

7. The method of claim 5 , further comprising:

processing, by the computer system, the voice input to recognize one or more words of the voice input;

interpreting, by the computer system, the one or more recognized words based on the context information determined from the non-voice input for interpreting the voice input,

wherein determining the user request comprises determining the user request based on the interpretation of the one or more recognized words.

8. The method of claim 7 , wherein at least one of the one or more recognized words is associated with at least two meanings,

wherein interpreting the one or more recognized words comprises selecting, based on the context information determined from the non-voice input for interpreting the voice input, one of the at least two meanings associated with the at least one recognized word to determine the user request.

9. The method of claim 1 , wherein the first user input comprises the non-voice input received via the non-voice input mode, and the second user input comprises the voice input received via the voice input mode,

wherein determining the context information comprises determining, based on the voice input, the context information for interpreting the non-voice input,

wherein determining the user request comprises determining the user request based on the non-voice input and the context information for interpreting the non-voice input, and

wherein providing the advertisement comprises providing the advertisement based on the context information for interpreting the non-voice input.

10. The method of claim 9 , wherein the context information for interpreting the non-voice input is used as input for selecting the advertisement, and wherein providing the advertisement comprises providing the selected advertisement.

11. The method of claim 1 , wherein the first item of the first item type comprises one of a command or a music-related product, and the second item of the second item type comprises the other one of the command or the music-related product.

12. The method of claim 1 , further comprising:

determining, by the computer system, prior context information associated with one or more prior voice inputs, wherein the one or more prior voice inputs are received by the computer system before the voice input is received, and

wherein determining the user request comprises determining the user request further based on the prior context information.

13. The method of claim 1 , wherein the context information for interpreting the first user input comprises information identifying at least one of a product, a service, a place, a location, an entity, or a content item.

14. The method of claim 1 , wherein the receipt of first user input is prior to, contemporaneously with, or subsequent to the receipt of the second user input.

15. A system for facilitating natural language processing of user inputs via multiple input modes where each user input alone may be insufficient to completely and/or accurately determine a user request intended by a user, the system comprising:

one or more physical processors programmed with computer program instructions which, when executed, cause the one or more physical processors to:

receive a first user input of a user from a first input device via a first input mode, wherein the first user input is generated responsive to the user interacting with the first input device in a manner corresponding to the first input mode to provide the first user input;

receive a second user input of the user from a second input device via a second input mode, wherein the second user input is generated responsive to the user interacting with the second input device in a manner corresponding to the second input mode to provide the second user input, wherein the first user input and the second user input are related to one another, and wherein one of the first user input or the second user input comprises a voice input received from at least one of the first input device or the second input device via a voice input mode, and the other one of the first user input or the second user input comprises a non-voice input received from at least one of the first input device or the second input device via a non-voice input mode;

determine, based on the second user input, context information for interpreting the first user input, wherein the context information identifies a first item of a first item type;

determine further context information based on the first user input, wherein the further context information identifies a second item of a second item type that is related to the first item of the first item type;

generate a query based on the context information and the further context information to obtain one or more intermediary results, wherein the generated query comprises a query related to the second item of the second type;

determine a user request based on the one or more intermediary results;

provide a response to the user request; and

provide, based on at least one of the context information for interpreting the first user input or the further context information, an advertisement for presentation to the user.

16. The system of claim 15 , wherein the first item of the first item type comprises one of a command or a music-related product, and the second item of the second item type comprises the other one of the command or the music-related product.

17. The system of claim 15 , further comprising:

provide the context information for interpreting the first user input to an advertiser system; and

obtain the advertisement from the advertiser system responsive to providing the context information to the advertiser system,

wherein providing the advertisement comprises providing the advertisement obtained from the advertiser system.

18. The system of claim 15 , wherein the context information for interpreting the first user input is used as input for selecting the advertisement, and wherein providing the advertisement comprises providing the selected advertisement.

19. The system of claim 15 , wherein information about the user request is used as input for selecting the advertisement, and wherein providing the advertisement comprises providing the selected advertisement.

20. The system of claim 15 , wherein the first user input comprises the voice input received via the voice input mode, and the second user input comprises the non-voice input received via the non-voice input mode,

wherein determining the context information comprises determining, based on the non-voice input, the context information for interpreting the voice input,

wherein determining the user request comprises determining the user request based on the voice input and the context information for interpreting the voice input, and

wherein providing the advertisement comprises providing the advertisement based on the context information for interpreting the voice input.

Assignments (10)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR'S NAME AND ASSIGNEE'S NAME PREVIOUSLY RECORDED ON REEL 051582 FRAME 0067. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNOR'S INTEREST. Recorded Oct 30, 2020
From: VOICEBOX TECHNOLOGIES CORPORATION
To: VB ASSETS, LLC
Reel/Frame 054260/0811 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2020
From: ORACLE INTERNATIONAL CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 051582/0067 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2019
From: VB ASSETS, LLC
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 049506/0531 →
RELEASE OF SECURITY INTEREST Recorded Jun 13, 2019
From: DELPHI ASSET MANAGEMENT CORPORATION
To: VB ASSETS, LLC
Reel/Frame 049459/0596 →
SECURITY INTEREST Recorded Apr 12, 2019
From: VB ASSETS, LLC
To: DELPHI ASSET MANAGEMENT CORPORATION
Reel/Frame 048872/0831 →
NUNC PRO TUNC ASSIGNMENT Recorded Jul 25, 2018
From: VOICEBOX TECHNOLOGIES CORPORATION
To: VB ASSETS, LLC
Reel/Frame 046456/0128 →
RELEASE OF SECURITY INTEREST Recorded Apr 5, 2018
From: ORIX GROWTH CAPITAL, LLC
To: VOICEBOX TECHNOLOGIES CORPORATION
Reel/Frame 045581/0630 →
SECURITY INTEREST Recorded Dec 22, 2017
From: VOICEBOX TECHNOLOGIES CORPORATION
To: ORIX GROWTH CAPITAL, LLC
Reel/Frame 044949/0948 →
MERGER Recorded Jul 9, 2014
From: VOICEBOX TECHNOLOGIES, INC.
To: VOICEBOX TECHNOLOGIES CORPORATION
Reel/Frame 033269/0503 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2014
From: BALDWIN, LARRY; WEIDER, CHRIS
To: VOICEBOX TECHNOLOGIES, INC.
Reel/Frame 033082/0247 →