IP Library Granted Patent US 10,460,727
Granted Patent B2
US 10,460,727 · App. 15/602,366 · Granted Oct 29, 2019

Multi-talker speech recognizer

Inventors: James Droppo (Carnation, WA); Xuedong Huang (Bellevue, WA); Dong Yu (Bothell, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/20G10L15/063G10L17/04G10L17/18G10L21/0308G10L25/30G10L25/51H04R3/005G10L15/32
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,460,727
App. No.
15/602,366
Granted
Oct 29, 2019
Kind
B2
Abstract

Various systems and methods for multi-talker speech separation and recognition are disclosed herein. In one example, a system includes a memory and a processor to process mixed speech audio received from a microphone. In an example, the processor can also separate the mixed speech audio using permutation invariant training, wherein a criterion of the permutation invariant training is defined on an utterance of the mixed speech audio. In an example, the processor can also generate a plurality of separated streams for submission to a speech decoder.

Claims (51)

1. A system for multi-talker speech separation and recognition, comprising:

a processor; and

a memory device coupled to the processor, the memory device to store instructions that, when executed by the processor, cause the processor to:

process a mixed speech audio received from a microphone;

use permutation invariant training (PIT) to separate and trace speech streams in the mixed speech audio, wherein the PIT includes:

estimate, in a first stage, a first mask by applying a first deep learning model to the mixed speech audio; and

estimate, in a second stage, a second mask by applying a second deep learning model to the first estimated mask and the mixed speech audio;

generate a plurality of separated speech streams from the mixed speech audio for submission to a speech decoder, wherein the plurality of separated speech streams are generated based on the first estimated mask and the second estimated mask; and

decode, using the speech decoder, speech of a first separated speech stream of the plurality of separated speech streams.

2. The system of claim 1 , wherein:

the plurality of separated speech streams comprises the first separated speech stream and a second separated speech stream; and

the processor is to decode with the speech decoder the speech of the first separated speech stream based on the first separated speech stream and the second separated speech stream.

3. The system of claim 1 , wherein the processor is to trace the mixed speech audio using permutation invariant training to provide tags to a speech stream indicating a speaker identity.

4. The system of claim 1 , wherein the permutation invariant training applies a language model of prediction to separating the mixed speech audio.

5. The system of claim 1 , wherein the permutation invariant training applies an acoustic model of prediction to separating the mixed speech audio.

6. The system of claim 1 , wherein:

the microphone is a microphone array; and

the mixed speech audio is multi-channel mixed speech audio.

7. The system of claim 1 , wherein the processor is to apply a beam former to separate the mixed speech audio into different directional streams based on a direction of the mixed speech audio compared to a direction of the beam former.

8. The system of claim 1 , wherein the processor is to adapt the separation of the mixed speech audio by the permutation invariant training based on supervision of a true transcript.

9. The system of claim 1 , wherein the plurality of separated speech streams comprises senone posterior probabilities.

10. The system of claim 9 , wherein the senone posterior probabilities are optimized using a cross entropy objective function.

11. The system of claim 1 , wherein the permutation invariant training applies a model that separates and traces speech streams at the same time, where the separation and tracing uses a stacking system in which a first stage separation result is used to inform a second stage model that includes the mixed speech and the separated speech streams from the first stage.

12. A method for multi-talker speech separation and recognition, comprising:

processing a mixed speech audio received from a microphone;

using permutation invariant training (PIT) to separate and trace speech streams in the mixed speech audio, wherein the PIT includes:

estimate, in a first stage, a first mask by applying a first deep learning model to the mixed speech audio; and

estimate, in a second stage, a second mask by applying a second deep learning model to the first estimated mask and the mixed speech audio;

generating a plurality of separated speech streams from the mixed speech audio for submission to a speech decoder, wherein the plurality of separated speech streams are generated based on the first estimated mask and the second estimated mask; and

decoding, using the speech decoder, speech of a first separated speech stream of the plurality of separated speech streams.

13. The method of claim 12 , comprising:

separating the mixed speech audio into the first separated speech stream and a second separated speech stream; and

decoding with the speech decoder the speech of the first separated speech stream based on the first separated speech stream and the second separated speech stream.

14. The method of claim 12 , comprising tracing the mixed speech audio using permutation invariant training to provide tags to a speech stream indicating a speaker identity.

15. The method of claim 12 , wherein the permutation invariant training applies a language model of prediction to separating the mixed speech audio.

16. The method of claim 12 , wherein the permutation invariant training applies an acoustic model of prediction to separating the mixed speech audio.

17. The method of claim 12 , wherein:

the microphone is a microphone array; and

the mixed speech audio is multi-channel mixed speech audio.

18. The method of claim 12 , comprising applying a beam former to separate the mixed speech audio into different directional streams based on a direction of the mixed speech audio compared to a direction of the beam former.

19. The method of claim 12 , comprising adapting the separation of the mixed speech audio by the permutation invariant training based on supervision of a true transcript.

20. A storage device that stores computer-readable instructions that, in response to an execution by a processor, cause the processor to:

process a mixed speech audio received from a microphone;

use permutation invariant training (PIT) to separate and trace speech streams in the mixed speech audio, wherein the PIT includes:

estimate, in a first stage, a first mask by applying a first deep learning model to the mixed speech audio; and

estimate, in a second stage, a second mask by applying a second deep learning model to the first estimated mask and the mixed speech audio;

generate a plurality of separated speech streams from the mixed speech audio for submission to a speech decoder, wherein the plurality of separated speech streams are generated based on the first estimated mask and the second estimated mask; and

decode, using the speech decoder, speech of a first separated speech stream of the plurality of separated speech streams.

21. The storage device of claim 20 , wherein:

the plurality of separated speech streams comprises the first separated speech stream and a second separated speech stream; and

wherein the instructions instruct a processor to decode with the speech decoder the speech of the first separated speech stream based on the first separated speech stream and the second separated speech stream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2017
From: DROPPO, JAMES; HUANG, XUEDONG; YU, DONG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 042482/0342 →
Continuity (2)
Provisional Application 62466902 · Mar 3, 2017
Related Publication 20180254040A1 · Sep 6, 2018
Cited By (2)
US 12,266,347 US 12,431,155