IP Library Granted Patent US 10,424,317
Granted Patent B2
US 10,424,317 · App. 15/403,481 · Granted Sep 24, 2019

Method for microphone selection and multi-talker segmentation with ambient automated speech recognition (ASR)

Inventors: Pablo Peso Parada (Berkshire, GB); Dushyant Sharma (Burlington, MA); Patrick Naylor (Berkshire, GB)
Assignee: Nuance Communications, Inc.
G10L21/0232G10L15/04G10L17/06G10L17/16G10L21/028G10L25/03H04R1/406G10L15/00G10L25/84G10L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,424,317
App. No.
15/403,481
Granted
Sep 24, 2019
Kind
B2
Abstract

Disclosed methods and systems are directed to determining a best microphone pair and segmenting sound signals. The methods and systems may include receiving a collection of sound signals comprising speech from one or more audio sources (e.g., meeting participants) and/or background noise. The methods and systems may include calculating a TDOA and determining, based on the TDOA and via robust statistics, the best pair of microphones. The methods and systems may also include segmenting sound signals from multiple sources.

Claims (553)

1. A method comprising:

receiving a plurality of audio signals, wherein each audio signal of the plurality of audio signals is received by one or more pairs of a plurality of microphones;

determining a time delay of arrival (TDOA) for each audio signal corresponding to a difference in receipt time of the plurality of audio signals for the one or more pairs of the plurality of microphones;

clustering the TDOAs to be associated with one of an audio source and interference, resulting in clustering information, wherein at least one TDOA is associated with the audio source and at least one TDOA is associated with the interference; and

segmenting each audio signal of the plurality of audio signals received by the one or more pairs of the plurality of microphones using the clustering information resulting from clustering the TDOAs to identify the audio source.

2. The method of claim 1 , wherein the plurality of microphones comprises at least three microphones, the method further comprising:

performing the clustering for possible pairs of the at least three microphones, resulting in additional clustering information;

generating, based on the additional clustering information, a confidence measure per possible pair of the at least three microphones; and

selecting, based on the confidence measure, one of the possible pairs of microphones.

3. The method of claim 2 , wherein the confidence measure is determined by:

CM l =max P (θ i |τ l ),

where i={1, . . . , N spk }, and

where P(θ i |τ l ), is determined based on a channel selection strategy.

4. The method of claim 1 , wherein the clustering is performed using statistical models.

5. The method of claim 4 wherein the statistical models comprise a Gaussian mixture model (GMM).

6. The method of claim 5 wherein a plurality of input parameters to the GMM are determined at each small time analysis window of a plurality of small time analysis windows and wherein the GMM is determined by:

arg

max

θ

v

log

(

θ

v

|

τ

v

)

,

where

v

=

{

1

,

2

,

,

N

TDOA

-

(

N

w

-

N

o

)

N

o

}

,

τ

v

=

{

τ

(

v

-

1

)

·

(

N

w

)

+

1

,

τ

(

v

-

1

)

·

(

N

w

)

+

2

,

,

τ

(

v

-

1

)

·

(

N

w

)

+

N

w

}

,

wherein N o comprises a number of overlapped frames,

wherein v represents one of the small time analysis windows, and

wherein N w comprises a length of each small time analysis window.

7. The method of claim 4 , wherein the statistical models are trained using conditional expectation-maximization.

8. The method of claim 7 , further comprising applying linear constraints on a plurality of means, wherein at least a first mean of the plurality of means is associated with a speaker model and at least a second mean of the plurality of means is associated with a noise model, and wherein the linear constraints are determined by:

[

μ

B

μ

1

μ

2

μ

N

spk

]

=

[

1

0

0

1

0

1

0

1

]

·

[

β

1

β

2

]

+

[

0

0

C

2

C

N

spk

]

,

μ

B

=

β

1

,

μ

1

=

β

2

,

μ

2

=

β

2

+

C

2

,

μ

N

spk

=

β

N

spk

+

C

N

spk

.

9. The method of claim 8 , further comprising determining C Nspk by:

C

N

spk

=

τ

maxN

spk

-

τ

max

1

,

where

,

τ

max

1

=

arg

max

τ

{

p

(

τ

)

|

dp

(

τ

)

d

τ

=

0

}

,

τ

maxN

spk

=

arg

max

τ

{

p

(

τ

)

|

dp

(

τ

)

d

τ

=

0

and

τ

{

τ

max

1

,

τ

max

2

,

,

τ

max

(

N

spk

-

1

)

}

}

,

p

(

τ

)

=

1

N

TDOA

l

=

1

N

TDOA

1

σ

(

τ

-

τ

l

σ

)

=

1

N

TDOA

l

=

1

N

TDOA

1

(

2

πσ

2

)

1

/

2

e

-

τ

-

τ

l

2

2

σ

2

,

σ

*

=

0.9

N

-

1

/

5

·

min

(

σ

,

IQR

/

1.34

)

,

wherein σ comprises a standard deviation, and

wherein IQR comprises an interquartile range computed from input data τ.

10. The method of claim 7 , further comprising applying linear constraints to a plurality of standard deviations, wherein at least a first standard deviation of the plurality of standard deviations is associated with a speaker model and at least a second standard deviation of the plurality of standard deviations is associated with a noise model, and wherein the linear constraints are determined by:

[

1

/

σ

B

1

/

σ

1

1

/

σ

2

1

/

σ

N

spk

]

=

[

ι

B

ι

1

ι

2

ι

N

spk

]

=

[

1

0

1

1

1

1

1

1

]

·

[

Υ

1

Υ

2

]

,

σ

B

=

1

/

Υ

1

,

{

σ

1

,

σ

2

,

,

σ

N

spk

}

=

1

/

Υ

1

+

1

/

Υ

2

.

11. The method of claim 1 , wherein at least one microphone of the plurality of microphones is placed at an ad hoc location.

12. The method of claim 11 , wherein the at least one microphone of the plurality of microphones comprises a mobile device.

13. The method of claim 1 , wherein the interference comprises background noise.

14. The method of claim 5 further comprising determining a speaker index that maximizes an a posteriori probability (MAP) of each GMM based on the TDOA, and wherein the MAP is maximized by:

arg

max

i

P

(

τ

l

|

θ

i

)

·

P

(

θ

i

)

.

15. A method comprising:

receiving a plurality of audio signals, each received by one or more pairs of a plurality of microphones;

determining a time delay of arrival (TDOA) for each audio signal corresponding to a difference in receipt time of the plurality of audio signals by the one or more pairs;

determining, based on the TDOAs, one or more Gaussian mixture models (GMM) for each microphone pair of the one or more microphone pairs of the plurality of microphones;

determining a speaker index that maximizes an a posteriori probability (MAP) associated with the GMM; and

selecting, based at least in part on the MAP, a microphone pair of the one or more pairs of the plurality of microphones.

16. A method comprising:

receiving a plurality of audio signals by one or more pairs of a plurality of microphones;

determining a time delay of arrival (TDOA) for each audio signal corresponding to a difference in receipt time of the plurality of audio signals at the one or more pairs;

determining, without supervision and based on the TDOAs, a Gaussian mixture model (GMM) for each microphone pair of the one or more microphone pairs of the plurality of microphones;

determining a speaker index that maximizes an a posteriori probability (MAP) associated with the GMM; and

selecting, based at least in part on the MAP, a microphone pair of the one or more pairs of the plurality of microphones.

17. The method of claim 16 , wherein the determining the GMM further comprises clustering the TDOA to be associated with an audio source, resulting in clustering information, wherein at least one TDOA is associated with speech and at least one TDOA is associated with background noise.

18. The method of claim 17 further comprising:

performing the clustering for additional pairs of microphones, resulting in additional clustering information;

generating a confidence level for the additional pairs of microphones based on the additional clustering information; and

selecting, based at least in part on the confidence level, the microphone pair of the one or more microphone pairs of the plurality of microphones.

19. The method of claim 17 further comprising segmenting, based on the clustering, the plurality of audio signals to identify the audio source at regular intervals.

20. The method of claim 18 , wherein the confidence level is determined by:

CM l =max P (θ i |τ l ),

where i={1, . . . , N spk }, and

where P(θ i |τ l ) is determined based on a channel selection strategy.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065531/0665 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 13, 2017
From: PARADA, PABLO PESO; SHARMA, DUSHYANT; NAYLOR, PATRICK
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041234/0078 →
Continuity (2)
Provisional Application 62394286 · Sep 14, 2016
Related Publication 20180075860A1 · Mar 15, 2018