IP Library Granted Patent US 12676137
Granted Patent B2
US 12676137 · App. 18/318,185 · Granted Jul 7, 2026

System and method for multi-channel speech privacy processing

Inventor: Dushyant Sharma (Tracy, CA)
Assignee: Microsoft Technology Licensing, LLC
G10L13/02G10L15/14
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12676137
App. No.
18/318,185
Granted
Jul 7, 2026
Kind
B2
Abstract

A method, computer program product, and computing system for receiving a speech signal from a single microphone. A sensitive speech component is identified from the speech signal. In response to identifying the sensitive speech component, a filtered speech signal is generated by removing the sensitive speech component from the speech signal. A voice style transfer of the speech signal is generated. Speech processing is performed on the filtered speech signal and the voice style transfer of the speech signal.

Claims (46)

1 . A method comprising:

receiving, by a speech processing system, an original speech signal recorded by a single microphone, the original speech signal containing voice characteristics of a user;

filtering the original speech signal to obtain a filtered speech signal that excludes the voice characteristics of the user;

performing voice conversion on the original speech signal to obtain a licensed voice signal containing voice characteristics of a licensed user, the voice characteristics of the user being excluded from the filtered speech signal, wherein the filtered speech signal and the licensed voice signal form a multi-channel input;

generating a single-channel representation of the multi-channel input by combining frames of the filtered speech signal with corresponding frames of the licensed voice signal in a weighted manner such that at least some frames of the filtered speech signal are assigned different weights; and

performing, via an automatic speech recognition (ASR) model, speech recognition on the single-channel representation of the multi-channel input to generate a transcript of the original speech signal.

2 . The method of claim 1 , further comprising:

randomly selecting the licensed user from a group of target speaker representations.

3 . The method of claim 1 , further comprising:

generating a synthetic speech signal by processing the transcript using a text-to-speech system.

4 . The method of claim 1 , wherein speech recognition is performed on the single-channel representation during run-time of the speech processing system.

5 . The method of claim 1 , further comprising:

selecting the licensed user based on a closest matching target speaker representation from a group of target speaker representations.

6 . The method of claim 1 , further comprising:

selecting the licensed user based on a least matching target speaker representation from a group of target speaker representations.

7 . The method of claim 1 , wherein the frames of the filtered speech signal and the corresponding frames of the licensed voice signal are time-domain frames having a predefined length.

8 . A computing system comprising:

a processor; and

a memory storing programming instructions for execution by the processor, the programming instructions, upon execution by the processor, causing the computing system to perform the following operations:

receiving, by a speech processing system, an original speech signal recorded by a single microphone, the original speech signal containing voice characteristics of a user;

filtering the original speech signal to obtain a filtered speech signal that excludes the voice characteristics of the user;

performing voice conversion on the original speech signal to obtain a licensed voice signal containing voice characteristics of a licensed user, the voice characteristics of the user being excluded from the filtered speech signal, wherein the filtered speech signal and the licensed voice signal form a multi-channel input;

generating a single-channel representation of the multi-channel input by combining frames of the filtered speech signal with corresponding frames of the licensed voice signal in a weighted manner such that at least some frames of the filtered speech signal are assigned different weights; and

performing, via an automatic speech recognition (ASR) model, speech recognition on the single-channel representation of the multi-channel input to generate a transcript of the original speech signal.

9 . The computing system of claim 8 , wherein the programming instructions further cause the computing system to perform the following operation:

generating a synthetic speech signal by processing the transcript using a text-to-speech system.

10 . The computing system of claim 8 , wherein speech recognition is performed on the single-channel representation during run-time of the speech processing system.

11 . The computing system of claim 8 , wherein the programming instructions further cause the computing system to perform the following operation:

selecting the licensed user based on a closest matching target speaker representation from a group of target speaker representations.

12 . The computing system of claim 8 , wherein the programming instructions further cause the computing system to perform the following operation:

selecting the licensed user based on a least matching target speaker representation from a group of target speaker representations.

13 . The computing system of claim 8 , wherein the frames of the filtered speech signal and the corresponding frames of the licensed voice signal are time-domain frames having a predefined length.

14 . A method comprising:

receiving, by a speech processing system, an original speech signal recorded by a single microphone, the original speech signal containing voice characteristics of a user;

filtering the original speech signal to obtain a filtered speech signal that excludes the voice characteristics of the user;

performing voice conversion on the original speech signal to obtain a licensed voice signal containing voice characteristics of a licensed user, the voice characteristics of the user being excluded from the filtered speech signal, wherein the filtered speech signal and the licensed voice signal form a multi-channel input;

generating a single-channel representation of the multi-channel input by combining bins of the filtered speech signal with corresponding bins of the licensed voice signal in a weighted manner such that at least some of the bins of the filtered speech signal are assigned different weights; and

performing, via an automatic speech recognition (ASR) model, speech recognition on the single-channel representation of the multi-channel input to generate a transcript of the original speech signal.

15 . The method of claim 14 , further comprising:

generating a synthetic speech signal by processing the transcript using a text-to-speech system.

16 . The method of claim 14 , wherein speech recognition is performed on the single-channel representation during run-time of the speech processing system.

17 . The method of claim 14 , further comprising:

selecting the licensed user based on a closest matching target speaker representation from a group of target speaker representations.

18 . The method of claim 14 , further comprising:

selecting the licensed user based on a least matching target speaker representation from a group of target speaker representations.

19 . The method of claim 14 , wherein the bins of the filtered speech signal and the corresponding bins of the licensed voice signal are frequency bins of a predefined size.