IP Library Granted Patent US 8,843,376
Granted Patent B2
US 8,843,376 · App. 11/685,350 · Granted Sep 23, 2014

Speech-enabled web content searching using a multimodal browser

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,843,376
App. No.
11/685,350
Granted
Sep 23, 2014
Kind
B2
Abstract

Speech-enabled web content searching using a multimodal browser implemented with one or more grammars in an automatic speech recognition (‘ASR’) engine, with the multimodal browser operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal browser operatively coupled to the ASR engine, includes: rendering, by the multimodal browser, web content; searching, by the multimodal browser, the web content for a search phrase, including yielding a matched search result, the search phrase specified by a first voice utterance received from a user and a search grammar; and performing, by the multimodal browser, an action in dependence upon the matched search result, the action specified by a second voice utterance received from the user and an action grammar.

Claims (64)

1. A method of speech-enabled searching of web content using a multimodal browser, the method implemented with one or more grammars in an automatic speech recognition (‘ASR’) engine, with the multimodal browser operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal browser operatively coupled to the ASR engine, the method comprising:

rendering, by the multimodal browser, web content;

searching, by the multimodal browser, the rendered web content for a search phrase, including matching the search phrase to at least one portion of the rendered web content, yielding a matched search result, the search phrase specified by a first voice utterance received from a user and a search grammar; and

in response to a second voice utterance received from the user:

using an action grammar comprising one or more entries to recognize the second voice utterance as corresponding to a first entry of the one or more entries, the action grammar specifying,

for the first entry of the one or more entries, an associated first action to be taken in dependence upon the matched search result, and

for a second entry of the one or more entries, an associated second action to be taken in dependence upon the same matched search result, the second action being different from the first action, and

performing, by the multimodal browser, the first action in dependence upon the matched search result associated with the first entry.

2. The method of claim 1 wherein searching, by the multimodal browser, the web content for a search phrase, including yielding a matched search result further comprises:

creating the search grammar in dependence upon the web content;

receiving the first voice utterance from a user; and

determining, using the ASR engine, the search phrase in dependence upon the first voice utterance and the search grammar.

3. The method of claim 2 wherein matching the search phrase to at least one portion of the web content, yielding a matched search result further comprises identifying a node of a Document Object Model (‘DOM’) representing the web content that contains the search phrase.

4. The method of claim 1 wherein performing, by the multimodal browser, an action in dependence upon the matched search result further comprises:

creating the action grammar in dependence upon the matched search result;

receiving the second voice utterance from the user;

determining, using the ASR engine, an action identifier in dependence upon the second voice utterance and the action grammar; and

performing the specified action in dependence upon the action identifier.

5. The method of claim 1 further comprising augmenting, by the multimodal browser, the matched search result with additional web content.

6. The method of claim 5 wherein augmenting, by the multimodal browser, the matched search result with additional web content further comprises inserting the additional web content into a node of a Document Object Model (‘DOM’) representing the web content that contains the matched search result.

7. The method of claim 1 wherein the web content is not speech-enabled.

8. Apparatus for speech-enabled searching of web content using a multimodal browser operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal browser operatively coupled to an automatic speech recognition (‘ASR’) engine, the apparatus comprising:

a computer processor; and

a computer memory operatively coupled to the computer processor, the computer memory having stored thereon computer program instructions that, when executed by the computer processor, perform a method comprising acts of:

rendering, by the multimodal browser, web content;

searching, by the multimodal browser, the rendered web content for a search phrase, including matching the search phrase to at least one portion of the rendered web content, yielding a matched search result, the search phrase specified by a first voice utterance received from a user and a search grammar; and

in response to a second voice utterance received from the user:

using an action grammar comprising one or more entries to recognize the second voice utterance as corresponding to a first entry of the one or more entries, the action grammar specifying,

for the first entry of the one or more entries, an associated first action to be taken in dependence upon the matched search result, and

for a second entry of the one or more entries, an associated second action to be taken in dependence upon the same matched search result, the second action being different from the first action, and

performing, by the multimodal browser, the first action in dependence upon the matched search result associated with the first entry.

9. The apparatus of claim 8 wherein searching, by the multimodal browser, the web content for a search phrase, including yielding a matched search result further comprises:

creating the search grammar in dependence upon the web content;

receiving the first voice utterance from a user; and

determining, using the ASR engine, the search phrase in dependence upon the first voice utterance and the search grammar.

10. The apparatus of claim 9 wherein matching the search phrase to at least one portion of the web content, yielding a matched search result further comprises identifying a node of a Document Object Model (‘DOM’) representing the web content that contains the search phrase.

11. The apparatus of claim 8 wherein performing, by the multimodal browser, an action in dependence upon the matched search result further comprises:

creating the action grammar in dependence upon the matched search result;

receiving the second voice utterance from the user;

determining, using the ASR engine, an action identifier in dependence upon the second voice utterance and the action grammar; and

performing the specified action in dependence upon the action identifier.

12. The apparatus of claim 8 further comprising computer program instructions capable of augmenting, by the multimodal browser, the matched search result with additional web content.

13. The apparatus of claim 12 wherein augmenting, by the multimodal browser, the matched search result with additional web content further comprises inserting the additional web content into a node of a Document Object Model (‘DOM’) representing the web content that contains the matched search result.

14. A computer-readable recordable medium encoded with instructions that, when executed, perform a method for speech-enabled searching of web content using a multimodal browser operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal browser operatively coupled to an automatic speech recognition (‘ASR’) engine, the method comprising acts of:

rendering, by the multimodal browser, web content;

searching, by the multimodal browser, the rendered web content for a search phrase, including matching the search phrase to at least one portion of the rendered web content, yielding a matched search result, the search phrase specified by a first voice utterance received from a user and a search grammar; and

in response to a second voice utterance received from the user:

using an action grammar comprising one or more entries to recognize the second voice utterance as corresponding to a first entry of the one or more entries, the action grammar specifying,

for the first entry of the one or more entries, an associated first action to be taken in dependence upon the matched search result, and

for a second entry of the one or more entries, an associated second action to be taken in dependence upon the same matched search result, the second action being different from the first action, and

performing, by the multimodal browser, the first action in dependence upon the matched search result associated with the first entry.

15. The computer-readable recordable medium of claim 14 wherein searching, by the multimodal browser, the web content for a search phrase, including yielding a matched search result further comprises:

creating the search grammar in dependence upon the web content;

receiving the first voice utterance from a user; and

determining, using the ASR engine, the search phrase in dependence upon the first voice utterance and the search grammar.

16. The computer-readable recordable medium of claim 15 wherein matching the search phrase to at least one portion of the web content, yielding a matched search result further comprises identifying a node of a Document Object Model (‘DOM’) representing the web content that contains the search phrase.

17. The computer-readable recordable medium of claim 14 wherein performing, by the multimodal browser, an action in dependence upon the matched search result further comprises:

creating the action grammar in dependence upon the matched search result;

receiving the second voice utterance from the user;

determining, using the ASR engine, an action identifier in dependence upon the second voice utterance and the action grammar; and

performing the specified action in dependence upon the action identifier.

18. The computer-readable recordable medium of claim 14 further comprising computer program instructions capable of augmenting, by the multimodal browser, the matched search result with additional web content.

19. The computer-readable recordable medium of claim 18 wherein augmenting, by the multimodal browser, the matched search result with additional web content further comprises inserting the additional web content into a node of a Document Object Model (‘DOM’) representing the web content that contains the matched search result.

20. The computer-readable recordable medium of claim 14 wherein the web content is not speech-enabled.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2007
From: CROSS, CHARLES W.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 019879/0090 →