IP Library Granted Patent US 7,809,568
Granted Patent B2
US 7,809,568 · App. 11/269,872 · Granted Oct 5, 2010

Indexing and searching speech with text meta-data

Assignee: Microsoft Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,809,568
App. No.
11/269,872
Granted
Oct 5, 2010
Kind
B2
Abstract

An index for searching spoken documents having speech data and text meta-data is created by obtaining probabilities of occurrence of words and positional information of the words of the speech data and combining it with at least positional information of the words in the text meta-data. A single index can be created because the speech data and the text meta-data are treated the same and considered only different categories.

Claims (37)

1. A method of indexing a spoken document comprising speech data and text meta-data, the method comprising:

using a processor to generate information pertaining to recognized speech from the speech data, the recognized speech comprising a sequence of textual words, the information comprising probabilities utilizing both a sum of length based probabilities and a word position probability to determine the words in the first sequence of words in the recognized speech and a position of each of the words in the first sequence of words;

using the processor to generate information pertaining to a second sequence of words in the text meta-data the text meta-data comprising a sequence of textual words, the information including at least positional information of a position of each of the words in the second sequence of words in the text meta-data with the same format as the positional information of the position of each of the words in the first sequence of words in the recognized speech;

using the processor to build an index based on processing text and the information pertaining to recognized speech including both the sum of length based probabilities and the word position probability and the information pertaining to the text meta-data wherein the index comprises position specific posterior lattices and wherein the position specific probability equals one for the text meta-data;

using the processor to output the index.

2. The method of claim 1 and further comprising categorizing at least one of speech data and text meta-data.

3. The method of claim 2 wherein categorizing includes categorizing different types of speech data.

4. The method of claim 2 wherein categorizing includes categorizing different types of text meta-data.

5. The method of claim 2 wherein building the index includes building the index with category information.

6. The method of claim 1 wherein generating information pertaining to recognized speech from the speech data comprises generating a lattice.

7. The method of claim 4 wherein generating information pertaining to the text meta-data comprises generating a lattice.

8. The method of claim 1 wherein generating information pertaining to recognized speech from the speech data includes identifying at least two alternative speech unit sequences based on the same portion of speech data; and wherein building an index based on the information pertaining to recognized speech includes for each speech unit in the at least two alternative speech unit sequences, placing information in an entry in the index that indicates a position of the speech unit in at least one of the two alternative speech unit sequences.

9. A non-transitory computer-readable storage medium having computer-executable instructions for performing steps comprising:

receiving a search query;

searching an index for an entry associated with a word in the search query, the index comprising:

information pertaining to a document identifier for a spoken document having speech data and text meta-data;

a category type identifier identifying at least one of different types of speech data, and speech data relative to text meta-data; and

positional information for the word based at least in part on the text meta-data comprising a plurality of words wherein the index comprises position specific posterior lattices and wherein the position specific probability equals one for each of the plurality of words in the text meta-data, the positional information indicating a position of the word in the plurality of words and a probability of the word appearing at the position based upon a summation of the probabilities of a word length along with a word position probability;

using the probabilities to rank spoken documents relative to each other; and

returning search results based on the ranked spoken documents.

10. The computer-readable storage medium of claim 9 wherein using the probabilities to rank the spoken documents comprises calculating a collection of composite n-gram scores for each spoken document.

11. The computer-readable storage medium of claim 10 wherein each composite n-gram score is formed by summing individual n-gram scores over all possible formations of an n-gram.

12. The computer-readable storage medium of claim 11 wherein the collection of composite n-gram scores is calculated based on different category types.

13. The computer-readable storage medium of claim 12 wherein a score for a category type is calculated by summing together each of the composite n-gram scores of each respective category type.

14. The computer-readable storage medium of claim 9 wherein using the probabilities to rank the spoken documents comprises calculating a document score as a combination of the category type scores.

15. The computer-readable storage medium of claim 14 wherein the category type scores are weighted.

16. A method of retrieving spoken documents based on a search query, the method comprising:

receiving the search query; and

using a processor for:

searching an index based on:

probabilities of positions for words in a sequence of words generated from speech data in the spoken documents, the probabilities of positions for words in the sequence of words referenced to different categories of speech data in the spoken document; and

positional information of a position of each of a plurality of words in a sequence of words in text meta-data associated with the speech data wherein the index comprises position specific posterior lattices and wherein the position specific probability equals one for each of the words in the text meta-data;

scoring each spoken document based on a set of probabilities for a word from the index for each category; and

returning search results based on the ranked spoken documents wherein the search results are pruned to remove the lower ranked documents.

17. The method of claim 16 wherein scoring each spoken document comprises calculating a document score as a weighted combination of scores for each different category of speech data.

18. The method of claim 16 wherein the index further includes probabilities of positions for words generated from text meta-data in the spoken documents, the probabilities of positions for words referenced to different categories of text meta-data in the spoken document.

19. The method of claim 18 wherein scoring each spoken document comprises calculating a document score as a weighted combination of scores for each different category of speech data and each different category of text meta-data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034543/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2005
From: SANCHEZ, JORGE F.SILVA; ACERO, ALEJANDRO; CHELBA, CIPRIAN I.
To: MICROSOFT CORPORATION
Reel/Frame 016927/0612 →
Continuity (1)
Related Publication 20070106509A1 · May 10, 2007