IP Library Granted Patent US 8,650,029
Granted Patent B2
US 8,650,029 · App. 13/035,044 · Granted Feb 11, 2014

Leveraging speech recognizer feedback for voice activity detection

Inventors: Albert Joseph Kishan Thambiratnam (Beijing, CN); Weiwu Zhu (Beijing, CN); Frank Torsten Bernd Seide (Beijing, CN)
Assignee: Microsoft Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,650,029
App. No.
13/035,044
Granted
Feb 11, 2014
Kind
B2
Abstract

A voice activity detection (VAD) module analyzes a media file, such as an audio file or a video file, to determine whether one or more frames of the media file include speech. A speech recognizer generates feedback relating to an accuracy of the VAD determination. The VAD module leverages the feedback to improve subsequent VAD determinations. The VAD module also utilizes a look-ahead window associated with the media file to adjust estimated probabilities or VAD decisions for previously processed frames.

Claims (38)

1. A method comprising:

under control of one or more computing devices comprising one or more processors,

classifying, by a VAD module, a plurality of frames of a media file into one or more speech frames and one or more non-speech frames;

receiving feedback associated with the one or more speech frames and the one or more non-speech frames, the feedback including a determination of accuracy of classifying one or more frames of the media file that are located at a predetermined time before the plurality of frames of the media file or that are selected based on a predetermined condition, the feedback being generated by a speech recognizer, the VAD module and the speech recognizer processing the media file asynchronously; and

utilizing the feedback to update a model to be used for voice activity detection (VAD) of a plurality of frames of the media file yet to be processed.

2. A method as recited in claim 1 , wherein the classifying includes identifying whether each of the plurality of frames of the media file includes speech or non-speech.

3. A method as recited in claim 1 , further comprising sending the one or more speech frames and the one or more non-speech frames to a speech recognizer.

4. A method as recited in claim 1 , further comprising classifying additional frames of the plurality of frames prior to receiving the feedback.

5. A method as recited in claim 1 , wherein the feedback includes a text transcript representing a content of the one or more speech frames.

6. A method as recited in claim 5 , wherein the text transcript is confidence-scored based at least in part on the accuracy of the classifying.

7. A method as recited in claim 6 , wherein the confidence-scored text transcript includes words or phrases of the media file that exceed a predetermined threshold of reliability.

8. A method comprising:

under control of one or more computing devices comprising one or more processors,

accessing a voice activity decision corresponding to one or more frames of a media file, the voice activity decision being generated by a VAD module;

after speech recognition of the one or more frames of the media file, receiving feedback associated with the voice activity decision that represents a relative accuracy of the voice activity decision, the media file being processed asynchronously such that each of the one or more frames is accessed prior to the generating the feedback, the feedback being generated by a speech recognizer; and

enabling use of the feedback to guide voice activity detection (VAD) for one or more subsequent frames of the media file.

9. A method as recited in claim 8 , wherein the voice activity decision indicates whether the one or more frames includes speech or non-speech.

10. A method as recited in claim 8 , further comprising converting the one or more frames into a text transcript representative of a content of the media file.

11. A method as recited in claim 8 , further comprising confidence-scoring a transcript corresponding to the media file such that words or phrases within the transcript that exceed a predetermined threshold are deemed to be confident.

12. A method as recited in claim 8 , wherein the feedback is leveraged to update models associated with a VAD module that are used for VAD.

13. A system comprising:

one or more processors;

memory communicatively coupled to the one or more processors for storing:

a voice activity detection (VAD) module configured to:

assign a probability to a first frame of a media file that represents a likelihood that the first frame includes speech; and

update the probability of the first frame based at least in part on one or more frames within a frame window;

assign a probability to a second frame within the frame window, the probability assigned to the second frame representing a likelihood that the second frame includes speech, the second frame being subsequent to the first frame; and

update the probability of the first frame based at least in part on the probability of the second frame.

14. A system as recited in claim 13 , wherein the VAD module is further configured to:

classify the first frame as including speech or non-speech;

classify a second frame within the frame window as including speech or non-speech; and

update the classifying of the first frame based at least in part on the classifying associated with the second frame.

15. A system as recited in claim 13 , wherein the VAD module is further configured to update the probability of the first frame prior to the first frame being processed by a speech recognizer.

16. A system as recited in claim 13 , wherein the VAD module is further configured to delay a voice activity detection decision associated with the first frame until the probability of the first frame is updated.

17. A system as recited in claim 13 , wherein the memory further stores a speech recognizer configured to convert one or more speech frames of the media file identified by the VAD module into a text transcript.

18. A system as recited in claim 13 , wherein the feedback includes a text transcript representing a content of one or more frames of the speech.

19. A system as recited in claim 18 , wherein the text transcript is confidence-scored.

20. A system as recited in claim 19 , wherein the confidence-scored text transcript includes words or phrases of the media file that exceed a predetermined threshold of reliability.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034544/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 12, 2011
From: THAMBIRATNAM, ALBERT JOSEPH KISHAN; ZHU, WEIWU; SEIDE, FRANK TORSTEN BERND
To: MICROSOFT CORPORATION
Reel/Frame 026111/0063 →
Continuity (1)
Related Publication 20120221330A1 · Aug 30, 2012