IP Library › Granted Patent US 7,283,958
Granted Patent B2
US 7,283,958 · App. 10/807,532 · Granted Oct 16, 2007

Systems and method for resolving ambiguity

Assignee: Fuji Xexox Co., Ltd.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,283,958
App. No.
10/807,532
Granted
Oct 16, 2007
Kind
B2
Abstract

Techniques are provided for resolving ambiguity in natural language speech. Speech is recognized using automatic speech recognition. A theory of discourse analysis is determined and at least one set of candidate discourse functions is determined based on the theory of discourse analysis. Prosodic features in the speech and a correlation between the prosodic features and the discourse functions is determined. The sets of candidate discourse functions are ranked based on the prosodic features in the speech information and a correlation to the prosodic features expected for the determined discourse functions. Ambiguity is resolved between sets of candidate discourse functions based on the rank information.

Claims (33)

1. A method of resolving ambiguity comprising the steps of:

determining recognized speech information;

determining discourse functions in the recognized speech information;

determining a predictive model of discourse functions based on prosodic features;

determining at least one set of candidate discourse functions for the recognized speech information;

determining a rank of the at least one set of discourse functions based on the predictive model of discourse functions; and

resolving the ambiguity between the set of at least one discourse functions based on the determined rank.

2. The method of claim 1 , wherein the discourse functions are determined based on a theory of discourse analysis.

3. The method of claim 2 , in which the theory of discourse analysis is at least one of: the Linguistic Discourse Model; the Unified Linguistic Discourse Model; Rhetorical Structures Theory; Discourse Structure Theory; and Structured Discourse Representation Theory.

4. The method of claim 1 , wherein the recognized speech information is directed to at least one of: a dictation mode; and a command mode.

5. The method of claim 1 , in which the prosodic features occur in at least one of: a location preceding; within; and following the associated discourse function.

6. The method of claim 1 , in which the prosodic features are encoded within a prosodic feature vector.

7. The method of claim 6 , in which the prosodic feature vector is a multimodal feature vector.

8. The method of claim 1 , in which the discourse function is an intra-sentential discourse function.

9. The method of claim 1 , in which the discourse function is an inter-sentential discourse function.

10. A system for synthesizing speech using discourse function level prosodic features comprising:

an input/output circuit for retrieving recognized speech and prosodic features;

a processor that determines at least one set of candidate discourse functions in the recognized speech information; determines a predictive model of discourse functions; determines a rank of the at least one set of candidate discourse functions based on the predictive model of discourse functions and the prosodic features of the recognized speech and disambiguates between the at least one set of candidate discourse functions based on a measure of prosodic correlation between the prosodic features for the recognized speech and the expected prosodic features associated with each discourse function in the predictive model of discourse functions.

11. The system of claim 10 , wherein the discourse functions are determined based on a theory of discourse analysis.

12. The system of claim 11 , in which the theory of discourse analysis is at least one of: the Linguistic Discourse Model; the Unified Linguistic Discourse Model Rhetorical Structures Theory; Discourse Structure Theory; and Structured Discourse Representation Theory.

13. The system of claim 10 , wherein the recognized speech information is directed to at least one of: a dictation mode; and a command mode.

14. The system of claim 10 , in which the prosodic features occur in at least one of: a location preceding; within; and following the associated discourse function.

15. The system of claim 11 , in which the prosodic features are encoded within a prosodic feature vector.

16. The system of claim 15 , in which the prosodic feature vector is a multimodal feature vector.

17. The system of claim 11 , in which the discourse function is an intra-sentential discourse function.

18. The system of claim 11 , in which the discourse function is an inter-sentential discourse function.

19. Computer readable storage medium comprising: computer readable program code embodied on the computer readable storage medium, the computer readable program code usable to program a computer to resolve ambiguity comprising the steps of:

determining recognized speech information;

determining discourse functions in the recognized speech information;

determining a predictive model of discourse functions based on prosodic features;

determining at least one set of candidate discourse functions for the recognized speech information;

determining a rank of the at least one set of discourse functions based on the predictive model of discourse functions; and

resolving the ambiguity between the set of at least one discourse functions based on the determined rank.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2004
From: AZARA, MISTY; POLANYI, LIVIA; THIONE, GIOVANNI L.; VAN DEN BERG, MARTIN H.
To: FUJI XEROX
Reel/Frame 015132/0658 →
Continuity (2)
Continuation In Part 1078144300 · Feb 18, 2004
Related Publication 20050182619A1 · Aug 18, 2005