IP Library Granted Patent US 12,579,984
Granted Patent B2
US 12,579,984 · App. 17/579,750 · Granted Mar 17, 2026

Data augmentation system and method for multi-microphone systems

Inventors: Dushyant Sharma (Mountain House, CA); Ljubomir Milanovic (Vienna, AT); Philipp Salletmayr (Austria, AT); Rong Gong (Vienna, AT); Patrick A. Naylor (Reading, GB)
Assignee: Microsoft Technology Licensing, LLC.
G10L17/04G10L21/0208G10L2021/02082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,984
App. No.
17/579,750
Granted
Mar 17, 2026
Kind
B2
Abstract

A method, computer program product, and computing system for obtaining one or more speech signals from a first device, thus defining one or more first device speech signals. One or more speech signals may be obtained from a second device, thus defining one or more second device speech signals. One or more acoustic relative transfer functions mapping reverberation from the one or more first device speech signals to the one or more second device speech signals may be generated. One or more augmented second device speech signals may be generated based upon, at least in part, the one or more acoustic relative transfer functions and first device training data.

Claims (60)

1 . A computer-implemented method, executed on a computing device, comprising:

obtaining one or more first device speech signals from a first device, wherein a first speech processing machine learning (ML) model was trained to process voice audio recorded by the first device using training data associated with the first device;

obtaining one or more second device speech signals from a second device;

generating augmented training data from the training data associated with the first device by generating one or more acoustic relative transfer functions (RTFs) mapping reverberation from the one or more first device speech signals to the one or more second device speech signals and processing the training data associated with the first device based on the one or more acoustic RTFs to generate the augmented training data, the one or more acoustic RTFs including one or more static acoustic RTFs generated by extracting a single acoustic relative transfer function per segment of the first device speech signals and the second device speech signals and one or more dynamic acoustic RTFs generated by extracting acoustic RTFs from the first device speech signals and the second device speech signals at predefined time increments, wherein the one or more dynamic acoustic RTFs model speaker movements in speech signals captured by the second device; and

training a second speech processing ML model to process voice audio recorded by the second device using the augmented training data.

2 . The computer-implemented method of claim 1 , further comprising:

processing the one or more first device speech signals; and

processing the one or more second device speech signals.

3 . The computer-implemented method of claim 2 , wherein processing the one or more first device speech signals includes:

detecting one or more speech active portions from the one or more first device speech signals;

identifying a speaker associated with the one or more speech active portions from the one or more first device speech signals; and

applying signal filtering to the one or more speech active portions associated with a predefined signal bandwidth to generate one or more first device filtered speech active portions.

4 . The computer-implemented method of claim 2 , wherein processing the one or more second device speech signals includes:

detecting one or more speech active portions from the one or more second device speech signals;

identifying a speaker associated with the one or more speech active portions from the one or more second device speech signals; and

applying signal filtering to the one or more speech active portions associated with a predefined signal bandwidth to generate one or more second device filtered speech active portions.

5 . The computer-implemented method of claim 1 , wherein generating one or more acoustic RTFs includes modeling relationships between characteristics of the one or more first device speech signals and characteristics of the one or more second device speech signals utilizing an adaptive filter.

6 . The computer-implemented method of claim 1 , further comprising:

adding the one or more acoustic RTFs to a codebook of acoustic RTFs.

7 . The computer-implemented method of claim 1 , wherein the first speech processing ML model is trained to process near field microphone system (NFMS) audio recorded by the first device using NFMS training data associated with the first device.

8 . The computer-implemented method of claim 7 , wherein the second ML model is trained to process far field microphone system (FFMS) audio recorded by the second device using the augmented training data.

9 . The computer-implemented method of claim 1 , wherein the second device is a microphone array that includes multiple audio acquisition devices deployed in an acoustic environment, and wherein the first device is a near field microphone system (NFMS) that is independent from the microphone array.

10 . A computer program product residing on a non-transitory computer readable medium having programming instructions stored thereon which, when executed by one or more processors of a system, cause the system to perform the following operations:

obtaining one or more first device speech signals from a first device, wherein a first speech processing machine learning (ML) model was trained to process voice audio recorded by the first device using training data associated with the first device;

obtaining one or more second device speech signals from a second device;

generating augmented training data from the training data associated with the first device by generating one or more acoustic relative transfer functions (RTFs) mapping reverberation from the one or more first device speech signals to the one or more second device speech signals and processing the training data associated with the first device based on the one or more acoustic RTFs to generate the augmented training data, the one or more acoustic RTFs including one or more static acoustic RTFs generated by extracting a single acoustic relative transfer function per segment of the first device speech signals and the second device speech signals and one or more dynamic acoustic RTFs generated by extracting acoustic RTFs from the first device speech signals and the second device speech signals at predefined time increments, wherein the one or more dynamic acoustic RTFs model speaker movements in speech signals captured by the second device; and

training a second speech processing ML model to process voice audio recorded by the second device using the augmented training data.

11 . The computer program product of claim 10 , wherein the operations further comprise:

processing the one or more first device speech signals; and

processing the one or more second device speech signals.

12 . The computer program product of claim 11 , wherein processing the one or more first device speech signals includes:

detecting one or more speech active portions from the one or more first device speech signals;

identifying a speaker associated with the one or more speech active portions from the one or more first device speech signals; and

applying signal filtering to the one or more speech active portions associated with a predefined signal bandwidth to generate one or more first device filtered speech active portions.

13 . The computer program product of claim 11 , wherein processing the one or more second device speech signals includes:

detecting one or more speech active portions from the one or more second device speech signals;

identifying a speaker associated with the one or more speech active portions from the one or more second device speech signals; and

applying signal filtering to the one or more speech active portions associated with a predefined signal bandwidth to generate one or more second device filtered speech active portions.

14 . The computer program product of claim 10 , wherein generating one or more acoustic RTFs includes modeling relationships between characteristics of the one or more first device speech signals and characteristics of the one or more second device speech signals utilizing an adaptive filter.

15 . The computer program product of claim 10 , wherein the operations further comprise:

adding the one or more acoustic RTFs to a codebook of acoustic RTFs.

16 . A computing system comprising:

one or more processors; and

a memory storing programming instructions for execution by the one or more processors, the programming instructions, upon execution by the one more processors, causing the computing system to perform the following operations:

obtaining one or more first device speech signals from a first device, wherein a first speech processing machine learning (ML) model was trained to process voice audio recorded by the first device using training data associated with the first device;

obtaining one or more second device speech signals from a second device;

generating augmented training data from the training data associated with the first device by generating one or more acoustic relative transfer functions (RTFs) mapping reverberation from the one or more first device speech signals to the one or more second device speech signals and processing the training data associated with the first device based on the one or more acoustic RTFs to generate the augmented training data, the one or more acoustic RTFs including one or more static acoustic RTFs generated by extracting a single acoustic relative transfer function per segment of the first device speech signals and the second device speech signals and one or more dynamic acoustic RTFs generated by extracting acoustic RTFs from the first device speech signals and the second device speech signals at predefined time increments, wherein the one or more dynamic acoustic RTFs model speaker movements in speech signals captured by the second device; and

training a second speech processing ML model to process voice audio recorded by the second device using the augmented training data.

17 . The computing system of claim 16 , wherein the programming instructions further cause the computing system to perform the following operations:

process the one or more first device speech signals; and

process the one or more second device speech signals.

18 . The computing system of claim 17 , wherein processing the one or more first device speech signals includes:

detecting one or more speech active portions from the one or more first device speech signals;

identifying a speaker associated with the one or more speech active portions from the one or more first device speech signals; and

applying signal filtering to the one or more speech active portions associated with a predefined signal bandwidth to generate one or more first device filtered speech active portions.

19 . The computing system of claim 16 , wherein processing the one or more second device speech signals includes:

detecting one or more speech active portions from the one or more second device speech signals;

identifying a speaker associated with the one or more speech active portions from the one or more second device speech signals; and

applying signal filtering to the one or more speech active portions associated with a predefined signal bandwidth to generate one or more second device filtered speech active portions.

20 . The computing system of claim 16 , wherein generating one or more acoustic RTFs includes modeling relationships between characteristics of the one or more first device speech signals and characteristics of the one or more second device speech signals utilizing an adaptive filter.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2022
From: SHARMA, DUSHYANT; MILANOVIC, LJUBOMIR; SALLETMAYR, PHILIPP; GONG, RONG; NAYLOR, PATRICK A.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 058706/0696 →
Continuity (1)
Related Publication 20230230599A1 · Jul 20, 2023
References Cited (30)
US 20110033063A1 · McGrath · 2011 [cited by examiner]
US 20110300806A1 · Lindahl · 2011 [cited by applicant]
US 20150025881A1 · Carlos · 2015 [cited by examiner]
US 20160064009A1 · Every · 2016 [cited by examiner]
US 20170094421A1 · Giri · 2017 [cited by examiner]
US 20170309294A1 · Gannot · 2017 [cited by examiner]
US 20180240471A1 · Markovich Golan · 2018 [cited by examiner]
US 20180350381A1 · Bryan · 2018 [cited by applicant]
US 20190362711A1 · Nosrati · 2019 [cited by applicant]
US 20200219524A1 · Braun · 2020 [cited by examiner]
US 20200395003A1 · Sharma et al. · 2020 [cited by applicant]
US 20210201931A1 · Wang et al. · 2021 [cited by applicant]
US 20210233509A1 · Lashkari · 2021 [cited by examiner]
US 20210329388A1 · Zahedi · 2021 [cited by examiner]
US 20230230582A1 · Sharma · 2023 [cited by applicant]
Brendel, Andreas et al. “Manifold Learning-Supported Estimation of Relative Transfer Functions for Spatial Filtering.” ICASSP 2022—2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (… [cited by examiner]
Sofer, Amit et al. “Robust Relative Transfer Function Identification on Manifolds for Speech Enhancement.” 2021 29th European Signal Processing Conference (EUSIPCO) (2021): 401-405. (Year: 2021). [cited by examiner]
Talmon, Ronen and Sharon Gannot. “Relative transfer function identification on manifolds for supervised GSC beamformers.” 21st European Signal Processing Conference (EUSIPCO 2013) (2013): 1-5. (Year: 2013). [cited by examiner]
A. Brendel, J. Zeitler and W. Kellermann, “Manifold Learning-Supported Estimation of Relative Transfer Functions for Spatial Filtering,” ICASSP 2022—2022 IEEE International Conference on Acoustics, Speech and Signal Pro… [cited by examiner]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US23/060989”, Mailed Date: May 2, 2023, 9 Pages. [cited by applicant]
Final Office Action mailed on Sep. 28, 2024, in U.S. Appl. No. 17/579,806, 24 pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2023/060986, mailed on May 2, 2023, 13 Pages. [cited by applicant]
Non-Final Office Action mailed on Apr. 23, 2024, in U.S. Appl. No. 17/579,806, 22 pages. [cited by applicant]
Notice of Allowance mailed on Jan. 23, 2025, in U.S. Appl. No. 17/579,806, 16 pages. [cited by applicant]
Notice of Allowance mailed on Mar. 19, 2025, in U.S. Appl. No. 17/579,806, 16 Pages. [cited by applicant]
Schwartz, et al., “Multi-microphone speech dereverberation and noise reduction using relative early transfer functions”, IEEE ACM Transactions, vol. 23, Issue No. 2, 2015, pp. 240-251. [cited by applicant]
Notice of Allowance mailed on Jul. 2, 2025, in U.S. Appl. No. 17/579,806, 16 pages. [cited by applicant]
Extended European Search Report Received in European Patent Application No. 23743949.2, mailed on Nov. 5, 2025, 11 pages. [cited by applicant]
Communication pursuant to Rules 70(2) and 70a(2) received for EP Application No. 23743949.2, mailed on Nov. 25, 2025, 1 page. [cited by applicant]
Extended European search report received in European Application No. 23743948.4, mailed on Nov. 7, 2025, 05 pages. [cited by applicant]