IP Library Granted Patent US 8,849,648
Granted Patent B1
US 8,849,648 · App. 10/329,138 · Granted Sep 30, 2014

System and method of extracting clauses for spoken language understanding

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,849,648
App. No.
10/329,138
Granted
Sep 30, 2014
Kind
B1
Abstract

A clausifier and method of extracting clauses for spoken language understanding are disclosed. The method relates to generating a set of clauses from speech utterance text and comprises inserting at least one boundary tag in speech utterance text related to sentence boundaries, inserting at least one edit tag indicating a portion of the speech utterance text to remove, and inserting at least one conjunction tag within the speech utterance text. The result is a set of clauses that may be identified within the speech utterance text according to the inserted at least one boundary tag, at least one edit tag and at least one conjunction tag. The disclosed clausifier comprises a sentence boundary classifier, an edit detector classifier, and a conjunction detector classifier. The clausifier may comprise a single classifier or a plurality of classifiers to perform the steps of identifying sentence boundaries, editing text, and identifying conjunctions within the text.

Claims (26)

1. A system comprising:

a processor; and

a computer-readable medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

inserting, via a discriminative classification approach independent of an n-gram approach, boundary tags, word by word, into speech utterance text, the boundary tags identifying boundaries selected from a group consisting of phrase boundaries, sentence boundaries, and paragraph boundaries, wherein the discriminative classification approach utilizes syntactic features before and after each word being tagged, to yield boundary marked speech utterance text;

thereafter, inserting, via the discriminative classification approach, edit tags in the boundary marked speech utterance text to identify at least one word of a set of duplicated words, to yield edited text and unedited text, wherein each edit tag comprises edit span information indicating how many words to the left of each edit tag to edit;

thereafter, identifying coordinating conjunctions within the unedited text based on a conjunction tag, to yield identified coordinating conjunctions, wherein the conjunction tag comprises conjunction span information indicating how many words to the left of the conjunction tag a corresponding conjunction includes; and

identifying clauses in the speech utterance text based on the boundary marked speech utterance text, the edited text, and the identified coordinating conjunctions.

2. The system of claim 1 , the computer-readable medium having additional instructions stored which result in the operations further comprising inserting editing tags into the speech utterance text.

3. The system of claim 2 , the computer-readable medium having additional instructions stored which result in the operations further comprising inserting conjunction tags into the speech utterance text.

4. A method comprising:

inserting, via a discriminative classification approach independent of an n-gram approach, boundary tags, word by word, into speech utterance text, the boundary tags identifying boundaries selected from a group consisting of phrase boundaries, sentence boundaries, and paragraph boundaries, wherein the discriminative classification approach utilizes syntactic features before and after each word being tagged, to yield boundary marked speech utterance text;

thereafter, inserting, via a processor and via the discriminative classification approach, edit tags in the boundary marked speech utterance text to identify at least one word of a set of duplicated words, to yield edited text and unedited text, wherein each edit tag comprises edit span information indicating how many words to the left of each edit tag to edit;

thereafter, identifying coordinating conjunctions within the unedited text based on a conjunction tag, to yield identified coordinating conjunctions, wherein the conjunction tag comprises conjunction span information indicating how many words to the left of the conjunction tag a corresponding conjunction includes; and

identifying clauses in the speech utterance text based on the boundary marked speech utterance text, the edited text, and the identified coordinating conjunctions.

5. The method of claim 4 , wherein the edit tags indicate a portion of the boundary marked speech utterance text to remove.

6. The method of claim 5 , further comprising inserting the conjunction tag within the edited text.

7. The method of claim 4 , wherein a different classifier performs each step of the method.

8. The method of claim 4 , wherein a single classifier performs all the steps of the method.

9. The method of claim 4 , wherein a plurality of classifiers perform all the steps of the method.

10. A computer-readable storage device having instructions stored which, when executed by a processor, cause the processor to perform operations comprising:

inserting, via a discriminative classification approach independent of an n-gram approach, boundary tags, word by word, into speech utterance text, the boundary tags identifying boundaries selected from a group consisting of phrase boundaries, sentence boundaries, and paragraph boundaries, wherein the discriminative classification approach utilizes syntactic features before and after each word being tagged, to yield boundary marked speech utterance text;

thereafter, inserting, via the discriminative classification approach, edit tags in the boundary marked speech utterance text to identify at least one word of a set of duplicated words, to yield edited text and unedited text, wherein each edit tag comprises edit span information indicating how many words to the left of each edit tag to edit;

thereafter, identifying coordinating conjunctions within the unedited text based on a conjunction tag, to yield identified coordinating conjunctions, wherein the conjunction tag comprises conjunction span information indicating how many words to the left of the conjunction tag a corresponding conjunction includes; and

identifying clauses in the speech utterance text based on the boundary marked speech utterance text, the edited text, and the identified coordinating conjunctions.

11. The computer-readable storage device of claim 10 , wherein the edit tags indicate a portion of the boundary marked speech utterance text to remove.

12. The computer-readable storage device of claim 11 , the computer-readable storage device having additional instructions stored which result in the operations further comprising inserting a conjunction tag within the speech utterance text.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 17, 2015
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 035436/0728 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 17, 2015
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 035436/0832 →