IP Library Patent Application 11702373
Patent Application
App. No. 11/702,373

Segmentation posterior based boundary point determination

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
11/702,373
Abstract

Boundary points for speech in an audio signal are determined based on posterior probabilities for the boundary points given a set of possible segmentations of the audio signal. The boundary point posterior probability is determined based on a set of level posterior probabilities that each provide the probability of a sequence of feature vectors given one of the segmentations in the set of possible segmentations.

Claims (32)

1 . A method comprising:

performing multiple segmentations of a sequence of feature vectors that represent an audio signal, each segmentation providing at least one segment boundary that is different from segment boundaries in other segmentations;

determining a separate segmentation probability for each segmentation, each segmentation probability providing the probability of the sequence of feature vectors given the segmentation;

determining a boundary point posterior probability for a possible boundary point of speech by summing over the segmentation probabilities for a set of segmentations that segment the sequence of feature vectors such that the possible boundary point represents a boundary point for speech in the segmentation; and

using the boundary point posterior probability to select a possible boundary point of speech as a boundary point of speech.

2 . The method of claim 1 wherein determining a segmentation probability comprises determining a separate segment posterior probability for each segment in the segmentation, each segment posterior probability providing the probability of a sub-sequence of the sequence of feature vectors given the segment that the sub-sequence of feature vectors is segmented into in the segmentation.

3 . The method of claim 2 wherein determining a segment posterior probability comprises determining a separate probability for each feature vector in the sub-sequence, wherein determining a probability for a feature vector comprises applying the feature vector to a normal distribution for a segment.

4 . The method of claim 3 wherein the normal distribution comprises a mean that is computed from the feature vectors in the segment.

5 . The method of claim 1 wherein the set of segmentations comprises levels of segmentation where each level comprises a different number of segments and wherein the number of segments is between a minimum level and a maximum level such that the set of levels contains fewer than all of the levels in which a segmentation was performed.

6 . The method of claim 5 further comprising determining the maximum level by identifying a level of segmentation at which a penalized homogeneity score is minimized.

7 . The method of claim 6 wherein the penalized homogeneity score comprises a distortion measure and a weighted penalty that is based on the number of segments in the level.

8 . The method of claim 7 wherein the weighted penalty is weighted by a weight that is trained on training data to minimize the mean square error of the estimated maximum level.

9 . The method of claim 8 wherein determining a boundary point posterior probability for a possible boundary point of speech comprises determining multiple boundary point posterior probabilities for multiple possible boundary points of speech.

10 . The method of claim 9 wherein determining multiple boundary point posterior probabilities for multiple possible boundary points of speech comprises determining a boundary point posterior probability for at least one starting boundary point that represents the start of speech and determining a boundary point posterior probability for at least one ending boundary point that represents the end of speech.

11 . The method of claim 10 wherein determining a boundary point posterior probability for a starting boundary point and a boundary point posterior probability for an ending boundary point comprises using a different maximum level for the starting boundary point than for the ending boundary point.

12 . A computer-readable medium having computer-executable instructions for performing steps comprising:

selecting a boundary point for speech in an audio signal from a plurality of possible boundary points found in a plurality of possible segmentations of the audio signal;

forming a summation of probabilities that is limited to probabilities associated with segmentations in the plurality of possible segmentations in which the selected possible boundary point is positioned as a boundary point;

using the summation of probabilities to determine a probability that the selected possible boundary point is a boundary point for speech; and

using the probability that the selected possible boundary point is a boundary point for speech to set a boundary point for speech in the audio signal.

13 . The computer-readable medium of claim 12 wherein each possible segmentation in the plurality of possible segmentations has a different number of segments.

14 . The computer-readable medium of claim 12 further comprising determining a maximum number of segments that can be in a segmentation associated with a probability used to form the summation of probabilities based on penalized homogeneity scores for the plurality of segmentations.

15 . The computer-readable medium of claim 12 wherein a probability associated with a segmentation is determined based in part on a probability of a feature vector given a segment in the segmentation.

16 . The computer-readable medium of claim 15 wherein the probability of a feature vector given a segment is based in part on a mean feature vector for the segment.

17 . A method comprising:

determining a maximum number of segments that can be found in segmentations used to determine the probability of a possible boundary point of speech in an audio signal based in part on distortion measures for a plurality of segmentations;

determining probabilities for a plurality of segmentations that have fewer than the maximum number of segments;

using the probabilities for the plurality of segmentations to determine a probability for each of at least two boundary points; and

selecting one of the at least two boundary points as a boundary point of speech in the audio signal based on the probabilities for the at least two boundary points.

18 . The method of claim 17 wherein the plurality of segmentations comprise at least one segmentation with more than the maximum number of segments.

19 . The method of claim 17 wherein determining a probability for a segmentation comprises determining a probability of a feature vector given a segment in the segmentation.

20 . The method of claim 19 wherein the probability of a feature vector given a segment is based in part on a mean feature vector for the segment.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2015
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034766/0509 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 26, 2007
From: SHI, YU; SOONG, FRANK KAO-PING
To: MICROSOFT CORPORATION
Reel/Frame 019061/0957 →