IP Library Patent Application 11679225
Patent Application
App. No. 11/679,225

Presenting Supplemental Content For Digital Media Using A Multimodal Application

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
11/679,225
Abstract

Presenting supplemental content for digital media using a multimodal application, implemented with a grammar of the multimodal application in an automatic speech recognition (‘ASR’) engine, with the multimodal application operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to the ASR engine, includes: rendering, by the multimodal application, a portion of the digital media; receiving, by the multimodal application, a voice utterance from a user; determining, by the multimodal application using the ASR engine, a recognition result in dependence upon the voice utterance and the grammar; identifying, by the multimodal application, supplemental content for the rendered portion of the digital media in dependence upon the recognition result; and rendering, by the multimodal application, the supplemental content.

Claims (35)

1 . A method of presenting supplemental content for digital media using a multimodal application, the method implemented with a grammar of the multimodal application in an automatic speech recognition (‘ASR’) engine, with the multimodal application operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to the ASR engine, the method comprising:

rendering, by the multimodal application, a portion of the digital media;

receiving, by the multimodal application, a voice utterance from a user;

determining, by the multimodal application using the ASR engine, a recognition result in dependence upon the voice utterance and the grammar;

identifying, by the multimodal application, supplemental content for the rendered portion of the digital media in dependence upon the recognition result; and

rendering, by the multimodal application, the supplemental content.

2 . The method of claim 1 wherein rendering, by the multimodal application, the supplemental content further comprises supplementing the rendered portion of digital media with the supplemental content.

3 . The method of claim 1 wherein the supplemental content further comprises annotated content for the digital media.

4 . The method of claim 1 wherein the supplemental content further comprises another portion of the digital media.

5 . The method of claim 1 wherein identifying, by the multimodal application, supplemental content for the digital media in dependence upon the recognition result further comprises querying a content repository for supplemental content associated with at least a portion of the recognition result.

6 . The method of claim 1 wherein identifying, by the multimodal application, supplemental content for the digital media in dependence upon the recognition result further comprises searching the digital media for supplemental content associated with at least a portion of the recognition result.

7 . The method of claim 1 wherein the grammar further comprises grammar rules, the grammar rules specifying recognition results according to the supplemental content.

8 . The method of claim 1 wherein the digital media is digital video.

9 . Apparatus for presenting supplemental content for digital media using a multimodal application, the method implemented with a grammar of the multimodal application in an automatic speech recognition (‘ASR’) engine, with the multimodal application operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to the ASR engine, the apparatus comprising a computer processor and a computer memory operatively coupled to the computer processor, the computer memory having disposed within it computer program instructions capable of:

rendering, by the multimodal application, a portion of the digital media;

receiving, by the multimodal application, a voice utterance from a user;

determining, by the multimodal application using the ASR engine, a recognition result in dependence upon the voice utterance and the grammar;

identifying, by the multimodal application, supplemental content for the rendered portion of the digital media in dependence upon the recognition result; and

rendering, by the multimodal application, the supplemental content.

10 . The apparatus of claim 9 wherein rendering, by the multimodal application, the supplemental content further comprises supplementing the rendered portion of digital media with the supplemental content.

11 . The apparatus of claim 9 wherein identifying, by the multimodal application, supplemental content for the digital media in dependence upon the recognition result further comprises querying a content repository for supplemental content associated with at least a portion of the recognition result.

12 . The apparatus of claim 9 wherein identifying, by the multimodal application, supplemental content for the digital media in dependence upon the recognition result further comprises searching the digital media for supplemental content associated with at least a portion of the recognition result.

13 . The apparatus of claim 9 wherein the grammar further comprises grammar rules, the grammar rules specifying recognition results according to the supplemental content.

14 . The apparatus of claim 9 wherein the digital media is digital video.

15 . A computer program product for presenting supplemental content for digital media using a multimodal application, the method implemented with a grammar of the multimodal application in an automatic speech recognition (‘ASR’) engine, with the multimodal application operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application operatively coupled to the ASR engine, the computer program product disposed upon a recordable medium, the computer program product comprising computer program instructions capable of:

rendering, by the multimodal application, a portion of the digital media;

receiving, by the multimodal application, a voice utterance from a user;

determining, by the multimodal application using the ASR engine, a recognition result in dependence upon the voice utterance and the grammar;

identifying, by the multimodal application, supplemental content for the rendered portion of the digital media in dependence upon the recognition result; and

rendering, by the multimodal application, the supplemental content.

16 . The computer program product of claim 15 wherein rendering, by the multimodal application, the supplemental content further comprises supplementing the rendered portion of digital media with the supplemental content.

17 . The computer program product of claim 15 wherein identifying, by the multimodal application, supplemental content for the digital media in dependence upon the recognition result further comprises querying a content repository for supplemental content associated with at least a portion of the recognition result.

18 . The computer program product of claim 15 wherein identifying, by the multimodal application, supplemental content for the digital media in dependence upon the recognition result further comprises searching the digital media for supplemental content associated with at least a portion of the recognition result.

19 . The computer program product of claim 15 wherein the grammar further comprises grammar rules, the grammar rules specifying recognition results according to the supplemental content.

20 . The computer program product of claim 15 wherein the digital media is digital video.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →
CORRECTIVE ASSIGNMENT TO CORRECT THE DOCKET NUMBER PREVIOUSLY RECORDED ON REEL 019091 FRAME 0428. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 3, 2007
From: CROSS, CHARLES W., JR.; GOODMAN, BRIAN D.; JANIA, FRANK L.; SHAW, DARREN M.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 019101/0953 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2007
From: CROSS, CHARLES W., JR.; GOODMAN, BRIAN D.; JANIA, FRANK L.; SHAW, DARREN M.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 019091/0428 →