IP Library Granted Patent US 9,781,262
Granted Patent B2
US 9,781,262 · App. 13/565,216 · Granted Oct 3, 2017

Methods and apparatus for voice-enabling a web application

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,781,262
App. No.
13/565,216
Granted
Oct 3, 2017
Kind
B2
Abstract

Methods and apparatus for voice-enabling a web application, wherein the web application includes one or more web pages rendered by a web browser on a computer. At least one information source external to the web application is queried to determine whether information describing a set of one or more supported voice interactions for the web application is available, and in response to determining that the information is available, the information is retrieved from the at least one information source. Voice input for the web application is then enabled based on the retrieved information.

Claims (48)

1. A method of determining a collective set of supported voice interactions for a plurality of frames of a web page displayed in a window of a web browser, wherein each of the plurality of frames includes content for a different web application, wherein the content for each of the plurality of frames is displayed simultaneously in the window of the web browser, wherein the plurality of frames includes a first frame and a second frame, wherein the first frame displays content for a first web application rendered by the web browser and the second frame displays content for a second web application rendered by the web browser, wherein the first web application is different from the second web application, the method comprising:

identifying a first data structure that includes information identifying a plurality of contexts of the first web application and supported voice interactions for the first web application in each of the plurality of contexts of the first web application;

determining a first current context of the first web application, wherein determining the first current context comprises analyzing whether a particular marker is present in the content displayed in the first frame;

determining based, at least in part, on the first current context of the first web application and the information included in the first data structure, a first set of supported voice interactions available for the first frame;

identifying a second data structure that includes information identifying a plurality of contexts of the second web application and supported voice interactions for the second web application in each of the plurality of contexts of the second web application;

determining based, at least in part, on a second current context of the second web application and the information included in the second data structure, a second set of supported voice interactions available for the second frame;

determining the collective set of supported voice interactions based on the first set of supported voice interactions and the second set of voice interactions; and

instructing an external speech engine to recognize voice input corresponding to the collective set of voice interactions.

2. The method of claim 1 , further comprising:

associating a first agent with the first frame and a second agent with the second frame, wherein the first agent is configured to monitor first events associated with the first frame and the second agent is configured to monitor second events associated with the second frame; and

wherein the first current context of the first web application is determined based, at least in part, on the first events and the second current context of the second web application is determined based, at least in part on the second events.

3. The method of claim 1 , wherein the collective set of supported voice interactions comprises a union of the supported voice interactions in the first set and the second set.

4. The method of claim 1 , wherein the collective set of supported voice interactions includes less than all of the supported voice interactions in the first set and the second set.

5. The method of claim 1 , wherein determining the first set of supported voice interactions comprises determining an identity of the first web application.

6. The method of claim 5 , wherein determining the identity of the first web application is performed based, at least in part, on an identifier for a web page of the first web application.

7. The method of claim 1 , wherein determining the first set of supported voice interactions available for the first frame comprises monitoring for browser events in the first frame.

8. A non-transitory computer-readable storage medium encoded with a plurality of instructions that, when executed by a computer, perform a method of determining a collective set of supported voice interactions for a plurality of frames of a web page displayed in a window of a web browser, wherein each of the plurality of frames includes content for a different web application, wherein the content for each of the plurality of frames is displayed simultaneously in the window of the web browser, wherein the plurality of frames includes a first frame and a second frame, wherein the first frame displays content for a first web application rendered by the web browser and the second frame displays content for a second web application rendered by the web browser, wherein the first web application is different from the second web application, the method comprising:

identifying a first data structure that includes information identifying a plurality of contexts of the first web application and supported voice interactions for the first web application in each of the plurality of contexts of the first web application;

determining a first current context of the first web application, wherein determining the first current context comprises analyzing whether a particular marker is present in the content displayed in the first frame;

determining based, at least in part, on the first current context of the first web application and the information included in the first data structure, a first set of supported voice interactions available for the first frame;

identifying a second data structure that includes information identifying a plurality of contexts of the second web application and supported voice interactions for the second web application in each of the plurality of contexts of the second web application;

determining based, at least in part, on a second current context of the second web application and the information included in the second data structure, a second set of supported voice interactions available for the second frame;

determining the collective set of supported voice interactions based on the first set of supported voice interactions and the second set of voice interactions; and

instructing an external speech engine to recognize voice input corresponding to the collective set of voice interactions.

9. The computer-readable storage medium of claim 8 , wherein the method further comprises:

associating a first agent with the first frame and a second agent with the second frame, wherein the first agent is configured to monitor first events associated with the first frame and the second agent is configured to monitor second events associated with the second frame; and

wherein the first current context of the first web application is determined based, at least in part, on the first events and the second current context of the second web application is determined based, at least in part on the second events.

10. The computer-readable storage medium of claim 8 , wherein the collective set of supported voice interactions comprises a union of the supported voice interactions in the first set and the second set.

11. The computer-readable storage medium of claim 8 , wherein the collective set of supported voice interactions includes less than all of the supported voice interactions in the first set and the second set.

12. The computer-readable storage medium of claim 8 , wherein determining the first set of supported voice interactions comprises determining an identity of the first web application.

13. The computer-readable storage medium of claim 12 , wherein determining the identity of the first web application is performed based, at least in part, on an identifier for a web page of the first web application.

14. The computer-readable storage medium of claim 8 , wherein determining the first set of supported voice interactions available for the first frame comprises monitoring for browser events in the first frame.

15. A computer for determining a collective set of supported voice interactions for a plurality of frames of a web page displayed in a window of a web browser, wherein each of the plurality of frames includes content for a different web application, wherein the content for each of the plurality of frames is displayed simultaneously in the window of the web browser, wherein the plurality of frames includes a first frame and a second frame, wherein the first frame displays content for a first web application rendered by the web browser and the second frame displays content for a second web application rendered by the web browser, wherein the first web application is different from the second web application, the computer comprising:

at least one processor programmed to:

identify a first data structure that includes information identifying a plurality of contexts of the first web application and supported voice interactions for the first web application in each of the plurality of contexts of the first web application;

determine a first current context of the first web application, wherein determining the first current context comprises analyzing whether a particular marker is present in the content displayed in the first frame;

determine based, at least in part, on the first current context of the first web application and the information included in the first data structure, a first set of supported voice interactions available for the first frame;

identify a second data structure that includes information identifying a plurality of contexts of the second web application and supported voice interactions for the second web application in each of the plurality of contexts of the second web application;

determine based, at least in part, on a second current context of the second web application and the information included in the second data structure, a second set of supported voice interactions available for the second frame;

determine the collective set of supported voice interactions based on the first set of supported voice interactions and the second set of voice interactions; and

instruct an external speech engine to recognize voice input corresponding to the collective set of voice interactions.

16. The computer of claim 15 , wherein the at least one processor is further programmed to:

associate a first agent with the first frame and a second agent with the second frame, wherein the first agent is configured to monitor first events associated with the first frame and the second agent is configured to monitor second events associated with the second frame; and

wherein the first current context of the first web application is determined based, at least in part, on the first events and the second current context of the second web application is determined based, at least in part on the second events.

17. The computer of claim 15 , wherein the collective set of supported voice interactions comprises a union of the supported voice interactions in the first set and the second set.

18. The computer of claim 15 , wherein the collective set of supported voice interactions includes less than all of the supported voice interactions in the first set and the second set.

19. The computer of claim 15 , wherein determining the first set of supported voice interactions comprises determining an identity of the first web application.

20. The computer of claim 19 , wherein determining the identity of the first web application is performed based, at least in part, on an identifier for a web page of the first web application.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2012
From: REICH, DAVID E.; HARDY, CHRISTOPHER
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 028781/0845 →