IP Library Granted Patent US 8,374,881
Granted Patent B2
US 8,374,881 · App. 12/324,388 · Granted Feb 12, 2013

System and method for enriching spoken language translation with dialog acts

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,374,881
App. No.
12/324,388
Granted
Feb 12, 2013
Kind
B2
Abstract

Disclosed herein are systems, computer-implemented methods, and tangible computer-readable media for enriching spoken language translation with dialog acts. The method includes receiving a source speech signal, tagging dialog acts associated with the received source speech signal using a classification model, dialog acts being domain independent descriptions of an intended action a speaker carries out by uttering the source speech signal, producing an enriched hypothesis of the source speech signal incorporating the dialog act tags, and outputting a natural language response of the enriched hypothesis in a target language. Tags can be grouped into sets such as statement, acknowledgement, abandoned, agreement, question, appreciation, and other. The step of producing an enriched translation of the source speech signal uses a dialog act specific translation model containing a phrase translation table.

Claims (41)

1. A method comprising:

tagging dialog acts associated with a user utterance in a source natural spoken language using a classification model to yield dialog act tags, the dialog act tags being domain independent descriptions of an intended action a speaker carries out by uttering the user utterance;

producing, via a processor, an enriched hypothesis of the user utterance incorporating the dialog act tags; and

outputting a version of the enriched hypothesis translated into a target natural spoken language to yield a translated speech output signal with a word order based on the dialog act tags.

2. The method of claim 1 , wherein the dialog acts are tagged using tags that are grouped into sets.

3. The method of claim 2 , wherein each of the sets is associated with a dialog act category selected from a group consisting of a statement category, an acknowledgement category, an abandoned category, an agreement category, a question category, and an appreciation category.

4. The method of claim 1 , wherein outputting the version of the enriched hypothesis uses a dialog act specific translation model containing a phrase translation table.

5. The method of claim 4 , the method further comprising:

appending to each phrase translation table associated with a particular dialog act specific translation model those entries from a complete model that are not present in the phrase table of the dialog act specific translation model, to yield appended entries; and

weighting the appended entries.

6. The method of claim 4 , wherein the dialog act specific translation model is a bag-of-words translation model.

7. The method of claim 1 , wherein the user utterance is part of a dialog turn having multiple sentences, the method further comprising:

segmenting the user utterance, to yield segments;

tagging second dialog acts in each segment using a maximum entropy model, to yield second tagged dialog acts; and

producing an enriched hypothesis of each segment incorporating the second tagged dialog acts.

8. The method of claim 1 , the method further comprising annotating tagged dialog acts.

9. The method of claim 1 , wherein the classification model is a maximum entropy model.

10. A system comprising:

a processor; and

a memory storing instructions for controlling the processor to perform a method comprising:

tagging dialog acts associated with a user utterance in a source natural spoken language using a classification model to yield dialog act tags, the dialog act tags being domain independent descriptions of an intended action a speaker carries out by uttering the user utterance;

producing an enriched hypothesis of the user utterance incorporating the dialog act tags; and

outputting a version of the enriched hypothesis translated into a target natural spoken language to yield a translated speech output signal with a word order based on the dialog act tags.

11. The system of claim 10 , wherein the dialog acts are tagged using tags that are grouped into sets.

12. The system of claim 11 , wherein each of the sets is associated with a dialog act category selected from a group consisting of a statement category, an acknowledgement category, an abandoned category, an agreement category, a question category, and an appreciation category.

13. The system of claim 10 , wherein outputting the version of the enriched hypothesis is based on a dialog act specific translation model containing a phrase translation table.

14. The system of claim 13 , the instructions further comprising:

appending to each phrase translation table associated with a particular dialog act specific translation model those entries from a complete model that are not present in the phrase table of the dialog act specific translation model, to yield appended entries; and

weighting the appended entries.

15. The system of claim 13 , wherein the dialog act specific translation model is a bag-of-words translation model.

16. The system of claim 10 , wherein the user utterance is part of a dialog turn having multiple sentences, the instructions further comprising:

segmenting the user utterance, to yield segments;

tagging second dialog acts in each segment using a maximum entropy model, to yield second tagged dialog acts; and

producing an enriched hypothesis of each segment incorporating the second tagged dialog acts.

17. A non-transitory computer-readable medium storing a computer program having instructions for controlling a computing device to perform a method comprising:

tagging dialog acts associated with a user utterance in a source natural spoken language using a classification model to yield dialog act tags, the dialog act tags being domain independent descriptions of an intended action a speaker carries out by uttering the user utterance;

producing, via a processor, an enriched hypothesis of the user utterance incorporating the dialog act tags; and

outputting a version of the enriched hypothesis translated into a target natural spoken language to yield a translated speech output signal with a word order based on the dialog act tags.

18. The non-transitory computer-readable medium of claim 17 , wherein the dialog acts are tagged using tags that are grouped into sets.

19. The non-transitory computer-readable medium of claim 18 , wherein each of the sets is associated with a dialog act category selected from a group consisting of a statement category, an acknowledgement category, an abandoned category, an agreement category, a question category, and an appreciation category.

20. The non-transitory computer-readable medium of claim 17 , wherein outputting the version of the enriched hypothesis uses a dialog act specific translation model containing a phrase translation table.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065566/0013 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 26, 2008
From: BANGALORE, SRINIVAS; RANGARAJAN SRIDHAR, VIVEK KUMAR
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 021897/0074 →