IP Library Granted Patent US 10,521,186
Granted Patent B2
US 10,521,186 · App. 13/847,974 · Granted Dec 31, 2019

Systems and methods for prompting multi-token input speech

Inventors: Ciprian Agapi (Lake Worth, FL); Soonthorn Ativanichayaphong (Boca Raton, FL); Leslie R. Wilson (Boca Raton, FL)
Assignee: Nuance Communications, Inc.
G06F3/167G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,521,186
App. No.
13/847,974
Granted
Dec 31, 2019
Kind
B2
Abstract

A method for prompting user input for a multimodal interface including the steps of providing a multimodal interface to a user, where the interface includes a visual interface having a plurality of input regions, each having at least one input field; selecting an input region and processing a multi-token speech input provided by the user, where the processed speech input includes at least one value for at least one input field of the selected input region; and storing at least one value in at least one input field.

Claims (48)

1. A method for prompting and processing user input for a multimodal interface, the method comprising acts of:

electronically rendering a visual interface comprising a plurality of input fields, wherein the plurality of input fields are arranged in a plurality of input regions, the plurality of input regions comprising a first input region, and wherein the first input region comprises at least first and second input fields;

electronically rendering one or more first indicia that distinguish the first input region from at least one third input field outside the first input region, wherein the first, second, and third input fields are concurrently displayed; and

processing, using at least a first grammar and a second grammar different from the first grammar, a speech input provided by the user in connection with the first input region, comprising:

determining, using the first grammar, whether the speech input provided by the user includes a request to switch to a second input region different from the first input region;

in response to determining that the speech input provided by the user includes a request to switch to a second input region different from the first input region, electronically rendering one or more second indicia that distinguish the second input region from the first input region;

in response to determining that the speech input provided by the user does not include a request to switch to a second input region different from the first input region, determining, using the second grammar, whether the speech input provided by the user comprises a multi-token speech input for the first input region; and

in response to determining that the speech input provided by the user comprises a multi-token speech input for the first input region:

processing a plurality of values identified from the multi-token speech input to determine whether the plurality of values comprise sufficient input for the first and second input fields; and

in response to determining that the plurality of values identified from the multi-token speech input comprise sufficient input for the first input field but incomplete input for the second input field, electronically rendering one or more third indicia to prompt the user to provide input for the second input field, the one or more third indicia distinguishing the second input field from the first input field, wherein the one or more third indicia are configured to indicate to the user that input is expected for the second input field, but not the first input field.

2. The method of claim 1 , wherein the act of using the first grammar to process the speech input comprises an act of:

identifying the second input region of the plurality of input regions to which the user wishes to switch.

3. The method of claim 2 , wherein the multi-token speech input is a first speech input, and wherein the method further comprises an act of:

using at least one third grammar associated with the second input region to process a second speech input provided by the user subsequent to the first speech input.

4. The method of claim 2 , wherein the method further comprises an act of:

electronically rendering one or more fourth indicia that distinguish the second input region from at least one other input region of the plurality of input regions.

5. A system for prompting and processing user input for a multimodal interface, the system comprising at least one processor programmed to:

electronically render a visual interface comprising a plurality of input fields, wherein the plurality of input fields are arranged in a plurality of input regions, the plurality of input regions comprising a first input region, and wherein the first input region comprises at least first and second input fields;

electronically render one or more first indicia that distinguish the first input region from at least one third input field outside the first input region, wherein the first, second, and third input fields are concurrently displayed; and

process, using at least a first grammar and a second grammar different from the first grammar, a speech input provided by the user in connection with the first input region, wherein the at least one processor is programmed to:

determine, using the first grammar, whether the speech input provided by the user includes a request to switch to a second input region different from the first input region;

in response to determining that the speech input provided by the user includes a request to switch to a second input region different from the first input region, electronically render one or more second indicia that distinguish the second input region from the first input region;

in response to determining that the speech input provided by the user does not include a request to switch to a second input region different from the first input region, determine, using the second grammar, whether the speech input provided by the user comprises a multi-token speech for the first input region; and

in response to determining that the speech input provided by the user comprises a multi-token speech input for the first input region:

process a plurality of values identified from the multi-token speech input to determine whether the plurality of values comprise sufficient input for the first and second input fields; and

in response to determining that the plurality of values identified from the multi-token speech input comprise sufficient input for the first input field but incomplete input for the second input field, electronically render one or more third indicia to prompt the user to provide input for the second input field, the one or more third indicia distinguishing the second input field from the first input field, wherein the one or more third indicia are configured to indicate to the user that input is expected for the second input field, but not the first input field.

6. The system of claim 5 , wherein, in using the first grammar to process the speech input, the at least one processor is programmed to:

identify the second input region of the plurality of input regions to which the user wishes to switch.

7. The system of claim 6 , wherein the multi-token speech input is a first speech input, and wherein the at least one processor is further programmed to:

use at least one third grammar associated with the second input region to process a second speech input provided by the user subsequent to the first speech input.

8. The system of claim 6 , wherein the at least one processor is further programmed to:

electronically render one or more fourth indicia that distinguish the second input region from at least one other input region of the plurality of input regions.

9. At least one non-transitory computer-readable medium having encoded thereon instructions that, when executed by at least one processor, perform a method for prompting and processing user input for a multimodal interface, the method comprising acts of:

electronically rendering a visual interface comprising a plurality of input fields, wherein the plurality of input fields are arranged in a plurality of input regions, the plurality of input regions comprising a first input region, and wherein the first input region comprises at least first and second input fields;

electronically rendering one or more first indicia that distinguish the first input region from at least one third input field outside the first input region, wherein the first, second, and third input fields are concurrently displayed; and

processing, using at least a first grammar and a second grammar different from the first grammar, a speech input provided by the user in connection with the first input region, comprising:

determining, using the first grammar, whether the speech input provided by the user includes a request to switch to a second input region different from the first input region;

in response to determining that the speech input provided by the user includes a request to switch to a second input region different from the first input region, electronically rendering one or more second indicia that distinguish the second input region from the first input region;

in response to determining that the speech input provided by the user does not include a request to switch to a second input region different from the first input region, determining, using the second grammar, whether the speech input provided by the user comprises a multi-token speech input for the first input region; and

in response to determining that the speech input provided by the user comprises a multi-token speech input for the first input region:

processing a plurality of values identified from the multi-token speech input to determine whether the plurality of values comprise sufficient input for the first and second input fields; and

in response to determining that the plurality of values identified from the multi-token speech input comprise sufficient input for the first input field but incomplete input for the second input field, electronically rendering one or more third indicia to prompt the user to provide input for the second input field, the one or more third indicia distinguishing the second input field from the first input field, wherein the one or more third indicia are configured to indicate to the user that input is expected for the second input field, but not the first input field.

10. The at least one non-transitory computer-readable medium of claim 9 , wherein the act of using the first grammar to process the speech input comprises an act of:

identifying the second input region of the plurality of input regions to which the user wishes to switch.

11. The at least one non-transitory computer-readable medium of claim 10 , wherein the multi-token speech input is a first speech input, and wherein the method further comprises an act of:

using at least one third grammar associated with the second input region to process a second speech input provided by the user subsequent to the first speech input.

12. The at least one non-transitory computer-readable medium of claim 10 , wherein the method further comprises an act of:

electronically rendering one or more fourth indicia that distinguish the second input region from at least one other input region of the plurality of input regions.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065531/0665 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2013
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 030257/0668 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2013
From: AGAPI, CIPRIAN; ATIVANICHAYAPHONG, SOONTHORN; WILSON, LESLIE R.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 030258/0096 →