IP Library Patent Application 11679292
Patent Application
App. No. 11/679,292

Enabling Natural Language Understanding In An X+V Page Of A Multimodal Application

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
11/679,292
Abstract

Enabling natural language understanding using an X+V page of a multimodal application implemented with a statistical language model (‘SLM’) grammar of the multimodal application in an automatic speech recognition (‘ASR’) engine, with the multimodal application operating in a multimodal browser on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to the ASR engine through a VoiceXML interpreter, including: receiving, in the ASR engine from the multimodal application, a voice utterance; generating, by the ASR engine according to the SLM grammar, at least one recognition result for the voice utterance; determining, by an action classifier for the VoiceXML interpreter, an action identifier in dependence upon the recognition result, the action identifier specifying an action to be performed by the multimodal application; and interpreting, by the VoiceXML interpreter, the multimodal application in dependence upon the action identifier.

Claims (50)

1 . A method of enabling natural language understanding using an X+V page of a multimodal application, the method implemented with a statistical language model (‘SLM’) grammar of the multimodal application in an automatic speech recognition (‘ASR’) engine, with the multimodal application operating in a multimodal browser on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to the ASR engine through a VoiceXML interpreter, the method comprising:

receiving, in the ASR engine from the multimodal application, a voice utterance;

generating, by the ASR engine according to the SLM grammar, at least one recognition result for the voice utterance;

determining, by an action classifier for the VoiceXML interpreter, an action identifier in dependence upon the recognition result, the action identifier specifying an action to be performed by the multimodal application; and

interpreting, by the VoiceXML interpreter, the multimodal application in dependence upon the action identifier.

2 . The method of claim 1 wherein the action identifier further comprises a plurality of action class identifiers, each action class identifier specifying an action class to which the specified action belongs.

3 . The method of claim 1 wherein:

the multimodal application specifies a particular action for the action identifier; and

interpreting, by the VoiceXML interpreter, the multimodal application in dependence upon the action identifier further comprises performing the particular action specified in the multimodal application for the action identifier.

4 . The method of claim 1 wherein:

the recognition result and the action identifier are represented as ECMAScript data structures; and

determining, by an action classifier for the VoiceXML interpreter, an action identifier in dependence upon the recognition result further comprises linking the ECMAScript data structure representing the action identifier to the ECMAScript data structure representing the recognition result.

5 . The method of claim 1 wherein determining, by an action classifier for the VoiceXML interpreter, an action identifier in dependence upon the recognition result further comprises determining identifier attributes for the action identifier, the identifier attributes specifying characteristics for the action identifier.

6 . The method of claim 5 wherein:

the multimodal application specifies a particular action for the action identifier; and

interpreting, by the VoiceXML interpreter, the multimodal application in dependence upon the action identifier further comprises performing, in dependence upon the identifier attributes, the particular action specified in the multimodal application for the action identifier.

7 . Apparatus for enabling natural language understanding using an X+V page of a multimodal application, the apparatus implemented with a statistical language model (‘SLM’) grammar of the multimodal application in an automatic speech recognition (‘ASR’) engine, with the multimodal application operating in a multimodal browser on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to the ASR engine through a VoiceXML interpreter, the apparatus comprising a computer processor and a computer memory operatively coupled to the computer processor, the computer memory having disposed within it computer program instructions capable of:

receiving, in the ASR engine from the multimodal application, a voice utterance;

generating, by the ASR engine according to the SLM grammar, at least one recognition result for the voice utterance;

determining, by an action classifier for the VoiceXML interpreter, an action identifier in dependence upon the recognition result, the action identifier specifying an action to be performed by the multimodal application; and

interpreting, by the VoiceXML interpreter, the multimodal application in dependence upon the action identifier.

8 . The apparatus of claim 7 wherein the action identifier further comprises a plurality of action class identifiers, each action class identifier specifying an action class to which the specified action belongs.

9 . The apparatus of claim 7 wherein:

the multimodal application specifies a particular action for the action identifier; and

interpreting, by the VoiceXML interpreter, the multimodal application in dependence upon the action identifier further comprises performing the particular action specified in the multimodal application for the action identifier.

10 . The apparatus of claim 7 wherein:

the recognition result and the action identifier are represented as ECMAScript data structures; and

determining, by an action classifier for the VoiceXML interpreter, an action identifier in dependence upon the recognition result further comprises linking the ECMAScript data structure representing the action identifier to the ECMAScript data structure representing the recognition result.

11 . The apparatus of claim 7 wherein determining, by an action classifier for the VoiceXML interpreter, an action identifier in dependence upon the recognition result further comprises determining identifier attributes for the action identifier, the identifier attributes specifying characteristics for the action identifier.

12 . The apparatus of claim 11 wherein:

the multimodal application specifies a particular action for the action identifier; and

interpreting, by the VoiceXML interpreter, the multimodal application in dependence upon the action identifier further comprises performing, in dependence upon the identifier attributes, the particular action specified in the multimodal application for the action identifier.

13 . A computer program product for enabling natural language understanding using an X+V page of a multimodal application, the computer program product implemented with a statistical language model (‘SLM’) grammar of the multimodal application in an automatic speech recognition (‘ASR’) engine, with the multimodal application operating in a multimodal browser on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to the ASR engine through a VoiceXML interpreter, the computer program product disposed upon a computer-readable, signal-bearing medium, the computer program product comprising computer program instructions capable of:

receiving, in the ASR engine from the multimodal application, a voice utterance;

generating, by the ASR engine according to the SLM grammar, at least one recognition result for the voice utterance;

determining, by an action classifier for the VoiceXML interpreter, an action identifier in dependence upon the recognition result, the action identifier specifying an action to be performed by the multimodal application; and

interpreting, by the VoiceXML interpreter, the multimodal application in dependence upon the action identifier.

14 . The computer program product of claim 13 wherein the computer-readable, signal-bearing medium comprises a recordable medium.

15 . The computer program product of claim 13 wherein the computer-readable, signal-bearing medium comprises a transmission medium.

16 . The computer program product of claim 13 wherein the action identifier further comprises a plurality of action class identifiers, each action class identifier specifying an action class to which the specified action belongs.

17 . The computer program product of claim 13 wherein:

the multimodal application specifies a particular action for the action identifier; and

interpreting, by the VoiceXML interpreter, the multimodal application in dependence upon the action identifier further comprises performing the particular action specified in the multimodal application for the action identifier.

18 . The computer program product of claim 13 wherein:

the recognition result and the action identifier are represented as ECMAScript data structures; and

determining, by an action classifier for the VoiceXML interpreter, an action identifier in dependence upon the recognition result further comprises linking the ECMAScript data structure representing the action identifier to the ECMAScript data structure representing the recognition result.

19 . The computer program product of claim 13 wherein determining, by an action classifier for the VoiceXML interpreter, an action identifier in dependence upon the recognition result further comprises determining identifier attributes for the action identifier, the identifier attributes specifying characteristics for the action identifier.

20 . The computer program product of claim 19 wherein:

the multimodal application specifies a particular action for the action identifier; and

interpreting, by the VoiceXML interpreter, the multimodal application in dependence upon the action identifier further comprises performing, in dependence upon the identifier attributes, the particular action specified in the multimodal application for the action identifier.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ATTORNEY DOCKET NUMBER PREVIOUSLY RECORDED ON REEL 019075 FRAME 0378. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 5, 2007
From: ATIVANICHAYAPHONG, SOONTHORN; CROSS, CHARLES W., JR.; MCCOBB, GERALD M.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 019119/0402 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2007
From: ATIVANICHAYAPHONG, SOONTHORN; CROSS, CHARLES W., JR.; MCCOBB, GERALD M.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 019075/0378 →