IP Library Granted Patent US 7,395,207
Granted Patent B1
US 7,395,207 · App. 11/466,815 · Granted Jul 1, 2008

Document expansion in speech retrieval

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,395,207
App. No.
11/466,815
Granted
Jul 1, 2008
Kind
B1
Abstract

Methods of document expansion for a speech retrieval document by a recognizer. A database of vectors of automatic transcriptions of documents is accessed and the vectors are truncated by removing all terms that are not recognizable by the recognizer to create truncated vectors. Terms in the vectors are then weighted to associate the truncated vectors with the untruncated vectors. Terms not recognized by the recognizer are then added back to the weighted, truncated vectors. The retrieval effectiveness may then be measured.

Claims (94)

1. A method of processing documents associated with speech retrieval, the method comprising:

removing terms in vectors that are not recognized by a recognizer, the vectors being associated with automatic transcriptions of documents;

generating weighted vectors by modifying weights of terms in the vectors; and

adding to the weighted vectors terms which were not recognized by the recognizer.

2. The method of claim 1 , further comprising:

accessing a database of vectors of automatic transcriptions of documents.

3. The method of claim 1 , wherein all terms not recognized by the recognizer are removed, and wherein any terms not recognized by the recognizer are added to the weighted vectors.

4. The method of claim 1 , further comprising:

comparing a document retrieved using original vectors with a document retrieval using weighted vectors with the added terms.

5. The method of claim 1 , further comprising:

measuring a loss in retrieval effectiveness due to the addition of terms not recognized into the weighted vectors.

6. The method of claim 5 , further comprising:

the step of determining final retrieval effectiveness of the speech retrieval document using automatic transcriptions.

7. The method of claim 1 , wherein generating the weighted vectors is based on the following function:

D

new

=

α

D

old

+

l

=

i

10

D

i

10

.

8. The computer-readable medium of claim 1 , wherein the instructions further comprise:

measuring a loss in retrieval effectiveness due to the addition of terms not recognized into the weighted vectors.

9. The computer-readable medium of claim 8 , wherein the instructions further comprise determining final retrieval effectiveness of the speech retrieval document using automatic transcriptions.

10. A computing device comprising:

a module configured to remove terms in vectors that are not recognized by a recognizer, the vectors being associated with automatic transcriptions of documents;

a module configured to generate weighted vectors by modifying weights of terms in the vectors; and

a module configured to add to the weighted vectors terms which are not recognized by the recognizer.

11. The computing device of claim 10 , further comprising:

a module configured to access a database of vectors of automatic transcriptions of documents.

12. The computing device of claim 10 , wherein all terms not recognized by the recognizer are removed, and wherein any terms not recognized by the recognizer are added to the weighted vectors.

13. The computing device of claim 10 , further comprising:

a module configured to compare a document retrieval using original vectors with a document retrieval using weighted vectors with the added terms.

14. The computing device of claim 10 , further comprising:

a module configured to measure a loss in retrieval effectiveness due to the addition of terms not recognized into the weighted vectors.

15. The computing device of claim 14 , further comprising:

a module configured to determine final retrieval effectiveness of the speech retrieval document using automatic transcriptions.

16. The computing device of claim 10 , wherein the module configured to generate the weighted vectors generates the weighted vectors based on the following function:

D

new

=

α

D

old

+

l

=

i

10

D

i

10

.

17. A computer-readable medium storing instructions for controlling a computing device, the instructions comprising:

removing terms in vectors that are not recognized by a recognizer, the vectors being associated with automatic transcriptions of documents;

generating weighted vectors by modifying weights of terms in the vectors; and

adding to the weighted vectors terms which were not recognized by the recognizer.

18. The computer-readable medium of claim 17 , wherein the instructions further comprise:

accessing a database of vectors of automatics transcriptions of documents.

19. The computer-readable medium of claim 15 , wherein all terms not recognized by the recognizer are removed, and wherein any terms not recognized by the recognizer are added to the weighted vectors.

20. The computer-readable medium of claim 17 , wherein the instructions further comprise:

comparing a document retrieved using original vectors with a document retrieval using weighted vectors with the added terms.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: PEREIRA, FERNANDO CARLOS; SINGHAL, AMITABH KUMAR
To: AT&T CORP.
Reel/Frame 038381/0302 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038529/0164 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038529/0240 →