IP Library Granted Patent US 8,380,513
Granted Patent B2
US 8,380,513 · App. 12/468,166 · Granted Feb 19, 2013

Improving speech capabilities of a multimodal application

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,380,513
App. No.
12/468,166
Granted
Feb 19, 2013
Kind
B2
Abstract

Improving speech capabilities of a multimodal application including receiving, by the multimodal browser, a media file having a metadata container; retrieving, by the multimodal browser, from the metadata container a speech artifact related to content stored in the media file for inclusion in the speech engine available to the multimodal browser; determining whether the speech artifact includes a grammar rule or a pronunciation rule; if the speech artifact includes a grammar rule, modifying, by the multimodal browser, the grammar of the speech engine to include the grammar rule; and if the speech artifact includes a pronunciation rule, modifying, by the multimodal browser, the lexicon of the speech engine to include the pronunciation rule.

Claims (25)

1. A method of improving speech capabilities of a multimodal application, the method implemented with a multimodal browser and a speech engine operating on a multimodal device supporting multiple modes of user interaction with the multimodal application, the modes of user interaction including a voice mode and one or more non-voice modes, wherein the voice mode includes accepting speech input from a user, digitizing the speech, and providing digitized speech to a speech engine available to the multimodal browser for recognition, and wherein the non-voice mode includes accepting input from a user through physical user interaction with a user input device for the multimodal device; wherein the multimodal browser comprises a module of automated computing machinery for executing the multimodal application and the multimodal browser supports execution of a media file player, a module of automated computing machinery for playing media files;

the method comprising:

receiving, by the multimodal browser, a media file having a metadata container;

retrieving, by the multimodal browser, from the metadata container a speech artifact related to content stored in the media file for inclusion in the speech engine available to the multimodal browser, wherein said retrieving, by the multimodal browser, from the metadata container the speech artifact for inclusion in a speech engine available to the multimodal browser comprises retrieving an XML document from the metadata container;

determining whether the speech artifact includes a grammar rule or a pronunciation rule;

if the speech artifact includes a grammar rule, modifying, by the multimodal browser, the grammar of the speech engine to include the grammar rule, wherein said modifying, by the multimodal browser, the grammar of the speech engine to include the grammar rule includes extracting from the XML document retrieved from the metadata container a grammar rule and including the grammar rule in an XML grammar document in the speech engine; and

if the speech artifact includes a pronunciation rule, modifying, by the multimodal browser, the lexicon of the speech engine to include the pronunciation rule, wherein said modifying, by the multimodal browser, the lexicon of the speech engine to include the pronunciation rule includes extracting from the XML document retrieved from the metadata container a pronunciation rule and including the pronunciation rule in an XML lexicon document in the speech engine.

2. The method of claim 1 wherein retrieving, by the multimodal browser, from the metadata container a speech artifact for inclusion in a speech engine available to the multimodal browser further comprises scanning the metadata container for a tag identifying the speech artifact.

3. The method of claim 2 wherein scanning the metadata container for a tag identifying the speech artifacts further comprises scanning an ID3 container of an MPEG media file for a frame identifying speech artifacts.

4. An apparatus for improving speech capabilities of a multimodal application, the apparatus including a multimodal browser and a multimodal application operating on a multimodal device supporting multiple modes of user interaction with the multimodal application, the modes of user interaction including a voice mode and one or more non-voice modes, the apparatus comprising a computer processor and a computer memory operatively coupled to the computer processor, the computer memory having disposed within it computer program instructions for:

receiving, by the multimodal browser, a media file having a metadata container;

retrieving, by the multimodal browser, from the metadata container a speech artifact related to content stored in the media file for inclusion in the speech engine available to the multimodal browser, wherein the computer program instructions for retrieving, by the multimodal browser, from the metadata container the speech artifact for inclusion in the speech engine available to the multimodal browser comprises computer program instructions for retrieving an XML document from the metadata container;

determining whether the speech artifact includes a grammar rule or a pronunciation rule;

if the speech artifact includes a grammar rule, modifying, by the multimodal browser, the grammar of the speech engine to include the grammar rule, wherein the computer program instructions for modifying, by the multimodal browser, the grammar of the speech engine to include the grammar rule includes computer program instructions for extracting from the XML document retrieved from the metadata container a grammar rule and including the grammar rule in an XML grammar document in the speech engine; and

if the speech artifact includes a pronunciation rule, modifying, by the multimodal browser, the lexicon of the speech engine to include the pronunciation rule, wherein the computer program instructions for modifying, by the multimodal browser, the lexicon of the speech engine to include the pronunciation rule include computer program instructions for extracting from the XML document retrieved from the metadata container a pronunciation rule and including the pronunciation rule in an XML lexicon document in the speech engine.

5. The apparatus of claim 4 wherein computer program instructions for retrieving, by the multimodal browser, from the metadata container a speech artifact for inclusion in a speech engine available to the multimodal browser further comprise computer program instructions for scanning the metadata container for a tag identifying the speech artifact.

6. The apparatus of claim 5 wherein computer program instructions for scanning the metadata container for a tag identifying the speech artifacts further comprise computer program instructions for scanning an ID3 container of an MPEG media file for a frame identifying speech artifacts.

7. A computer program product for improving speech capabilities of a multimodal application, the computer program product including a multimodal browser for operating on a multimodal device supporting multiple modes of user interaction with the multimodal application, the modes of user interaction including a voice mode and one or more non-voice modes, the computer program product disposed upon a computer-readable, recording medium, the computer program product comprising computer program instructions capable for:

receiving, by the multimodal browser, a media file having a metadata container;

retrieving, by the multimodal browser, from the metadata container a speech artifact related to content stored in the media file for inclusion in the speech engine available to the multimodal browser, wherein the computer program instructions for retrieving, by the multimodal browser, from the metadata container the speech artifact for inclusion in the speech engine available to the multimodal browser comprises computer program instructions for retrieving an XML document from the metadata container;

determining whether the speech artifact includes a grammar rule or a pronunciation rule;

if the speech artifact includes a grammar rule, modifying, by the multimodal browser, the grammar of the speech engine to include the grammar rule, wherein the computer program instructions for modifying, by the multimodal browser, the grammar of the speech engine to include the grammar rule includes computer program instructions for extracting from the XML document retrieved from the metadata container a grammar rule and including the grammar rule in an XML grammar document in the speech engine; and

if the speech artifact includes a pronunciation rule, modifying, by the multimodal browser, the lexicon of the speech engine to include the pronunciation rule, wherein the computer program instructions for modifying, by the multimodal browser, the lexicon of the speech engine to include the pronunciation rule include computer program instructions for extracting from the XML document retrieved from the metadata container a pronunciation rule and including the pronunciation rule in an XML lexicon document in the speech engine.

8. The computer program product of claim 7 wherein computer program instructions for retrieving, by the multimodal browser, from the metadata container a speech artifact for inclusion in a speech engine available to the multimodal browser further comprise computer program instructions for scanning the metadata container for a tag identifying the speech artifact.

9. The computer program product of claim 8 wherein computer program instructions for scanning the metadata container for a tag identifying the speech artifacts further comprise computer program instructions for scanning an ID3 container of an MPEG media file for a frame identifying speech artifacts.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2013
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 030381/0123 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2009
From: AGAPI, CIPRIAN; BODIN, WILLIAM K.; CROSS, CHARLES W., JR.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 022702/0810 →