IP Library Granted Patent US 8,452,594
Granted Patent B2
US 8,452,594 · App. 12/091,079 · Granted May 28, 2013

Method and system for processing dictated information

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,452,594
App. No.
12/091,079
Granted
May 28, 2013
Kind
B2
Abstract

A method and a system for processing dictated information into a dynamic form are disclosed. The method comprises presenting an image ( 3 ) belonging to an image category to a user, dictating a first section of speech associated with the image category, retrieving an electronic document having a previously defined document structure ( 4 ) associated with the first section of speech, thus associating the document structure ( 4 ) with the image ( 3 ), wherein the document structure comprises at least one text field, presenting at least a part of the electronic document having the document structure ( 4 ) on a presenting unit ( 5 ), dictating a second section of speech and processing the second section of speech in a speech recognition engine ( 6 ) into dictated text and associating the dictated text with the text field.

Claims (72)

1. A method for processing dictated information into a dynamic form, the method comprising:

retrieving an electronic document having a document structure, wherein the document structure comprises at least one text field;

presenting at least a part of the electronic document to a user;

receiving speech input of first dictated information;

processing the speech input in a speech recognition engine into dictated text;

associating the dictated text with a first text field; and

based at least in part on the speech input, adjusting a display of the electronic document to control how much of the electronic document is displayed to the user.

2. The method according to claim 1 , wherein the electronic document having the document structure comprises is associated with a set of data that is specific to a topic and contains a population of words that are likely to be found in the at least one text field, and

wherein processing the speech input in a speech recognition engine into dictated text comprises using a statistical model of likelihood of use of the population of words.

3. The method according to claim 2 , wherein at least a portion of the set of data associated with the electronic document is associated with a specific text field in the document structure.

4. The method of claim 3 , wherein the set of data comprises a first subset of data and a second subset of data, the first subset of data being associated with at least one first text field of the at least one text field of the electronic document and the second subset of data being associated with at least one second text field of the at least one text field, and

wherein the method further comprises switching between the first subset of data and the second subset of data depending on a current text field selected by the user for dictation of text into the electronic document.

5. The method according to claim 1 , wherein the document structure comprises a plurality of text fields, and the method further comprises:

defining a voice macro associated with a specific text field of the text fields, such that the specific text field is chosen for receipt of the dictated text of the speech input when the voice macro is dictated by the user.

6. The method according to claim 5 , further comprising:

filling the plurality of text fields based on an order in which voice macros corresponding to each of the plurality of text fields are dictated by the user.

7. The method according to claim 1 , wherein adjusting the display of the electronic document comprises dynamically expanding or reducing a number of text fields, of the at least one text field of the electronic document, presented to the user.

8. The method of claim 1 , further comprising:

displaying an image to the user;

receiving, via a speech interface, a first speech input from the user regarding the image; and

based at least in part on the first speech input, selecting, from at least one available electronic document, the electronic document to be retrieved and displayed.

9. The method according to claim 8 , further comprising:

linking the image to the electronic document having the document structure and the dictated text, and

storing the image and the electronic document in a data store.

10. The method according to claim 9 , further comprising:

identifying the text field with a marker;

converting the marked text field into a code string;

storing the code string together with the image in the data store.

11. The method according to claim 10 , wherein identifying the text field with the marker comprises automatically performing the identifying of the text field with the marker.

12. The method according to claim 10 , wherein the converting the marked text field to a code string comprises exporting the marked text field as text and converting the markers into nodes, created by a general markup language, in a document having the document structure.

13. The method of claim 8 , wherein:

the image belongs to an image category,

receiving the first speech input comprises receiving an indication of the image category, and

selecting the electronic document based at least in part on the first speech input comprises selecting the electronic document associated with the image category.

14. A system for processing dictated information into a dynamic form, the system comprising:

at least one processor programmed to perform acts of:

retrieving an electronic document having a document structure, wherein the document structure comprises at least one text field;

presenting at least a part of the electronic document having the document structure;

receiving speech input of first dictated information;

using a speech recognition engine, processing the speech input into dictated text;

associating the dictated text with a first text field; and

based at least in part on the speech input, adjusting a display of the electronic document to control how much of the electronic document is displayed to the user.

15. The system of claim 14 , wherein the at least one processor is further programmed to perform acts of:

displaying an image to the user;

receiving, via a speech interface, a first speech input from the user regarding the image; and

based at least in part on the first speech input, selecting, from at least one available electronic document, the electronic document to be retrieved and displayed.

16. The system of claim 15 , wherein:

the image belongs to an image category,

receiving the first speech input comprises receiving an indication of the image category, and

selecting the electronic document based at least in part on the first speech input comprises selecting the electronic document associated with the image category.

17. The system of claim 15 , wherein the at least one processor is further programmed to perform acts of:

linking the image to the electronic document having the document structure and the dictated text; and

storing the image and the electronic document in a data store.

18. The system of claim 14 , wherein adjusting the display of the electronic document comprises dynamically expanding or reducing a number of text fields, of the at least one text field of the electronic document, presented to the user.

19. At least one non-transitory computer-readable storage medium having encoded thereon a computer program that, when executed by a computer, causes the computer to perform a method for processing dictated information into a dynamic form, the method comprising:

retrieving an electronic document having a document structure, wherein the document structure comprises at least one text field;

presenting at least a part of the electronic document to a user;

processing speech input in a speech recognition engine into dictated text;

associating the dictated text with a first text field; and

based at least in part on the speech input, adjusting a display of the electronic document to control how much of the electronic document is displayed to the user.

20. The at least one non-transitory computer-readable storage medium of claim 19 , wherein the method further comprises:

displaying an image to the user;

receiving, via a speech interface, a first speech input from the user regarding the image; and

based at least in part on the first speech input, selecting, from at least one available electronic document, the electronic document to be retrieved and displayed.

21. The at least one non-transitory computer-readable storage medium of claim 20 , wherein:

the image belongs to an image category,

receiving the first speech input comprises receiving an indication of the image category, and

selecting the electronic document based at least in part on the first speech input comprises selecting the electronic document associated with the image category.

22. The at least one non-transitory computer-readable storage medium of claim 20 , wherein the method further comprises:

linking the image to the electronic document having the document structure and the dictated text; and

storing the image and the electronic document in a data store.

23. The at least one non-transitory computer-readable storage medium of claim 19 , wherein adjusting the display of the electronic document comprises dynamically expanding or reducing a number of text fields, of the at least one text field of the electronic document, presented to the user.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065533/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 24, 2009
From: KONINKLIJKE PHILIPS ELECTRONICS N.V.
To: NUANCE COMMUNICATIONS AUSTRIA GMBH
Reel/Frame 022299/0350 →