System and method for multi-channel speech privacy processing
A method, computer program product, and computing system for receiving a speech signal from a single microphone. A sensitive speech component is identified from the speech signal. In response to identifying the sensitive speech component, a filtered speech signal is generated by removing the sensitive speech component from the speech signal. A voice style transfer of the speech signal is generated. Speech processing is performed on the filtered speech signal and the voice style transfer of the speech signal.
1 . A method comprising:
receiving, by a speech processing system, an original speech signal recorded by a single microphone, the original speech signal containing voice characteristics of a user;
filtering the original speech signal to obtain a filtered speech signal that excludes the voice characteristics of the user;
performing voice conversion on the original speech signal to obtain a licensed voice signal containing voice characteristics of a licensed user, the voice characteristics of the user being excluded from the filtered speech signal, wherein the filtered speech signal and the licensed voice signal form a multi-channel input;
generating a single-channel representation of the multi-channel input by combining frames of the filtered speech signal with corresponding frames of the licensed voice signal in a weighted manner such that at least some frames of the filtered speech signal are assigned different weights; and
performing, via an automatic speech recognition (ASR) model, speech recognition on the single-channel representation of the multi-channel input to generate a transcript of the original speech signal.
2 . The method of claim 1 , further comprising:
randomly selecting the licensed user from a group of target speaker representations.
3 . The method of claim 1 , further comprising:
generating a synthetic speech signal by processing the transcript using a text-to-speech system.
4 . The method of claim 1 , wherein speech recognition is performed on the single-channel representation during run-time of the speech processing system.
5 . The method of claim 1 , further comprising:
selecting the licensed user based on a closest matching target speaker representation from a group of target speaker representations.
6 . The method of claim 1 , further comprising:
selecting the licensed user based on a least matching target speaker representation from a group of target speaker representations.
7 . The method of claim 1 , wherein the frames of the filtered speech signal and the corresponding frames of the licensed voice signal are time-domain frames having a predefined length.
8 . A computing system comprising:
a processor; and
a memory storing programming instructions for execution by the processor, the programming instructions, upon execution by the processor, causing the computing system to perform the following operations:
receiving, by a speech processing system, an original speech signal recorded by a single microphone, the original speech signal containing voice characteristics of a user;
filtering the original speech signal to obtain a filtered speech signal that excludes the voice characteristics of the user;
performing voice conversion on the original speech signal to obtain a licensed voice signal containing voice characteristics of a licensed user, the voice characteristics of the user being excluded from the filtered speech signal, wherein the filtered speech signal and the licensed voice signal form a multi-channel input;
generating a single-channel representation of the multi-channel input by combining frames of the filtered speech signal with corresponding frames of the licensed voice signal in a weighted manner such that at least some frames of the filtered speech signal are assigned different weights; and
performing, via an automatic speech recognition (ASR) model, speech recognition on the single-channel representation of the multi-channel input to generate a transcript of the original speech signal.
9 . The computing system of claim 8 , wherein the programming instructions further cause the computing system to perform the following operation:
generating a synthetic speech signal by processing the transcript using a text-to-speech system.
10 . The computing system of claim 8 , wherein speech recognition is performed on the single-channel representation during run-time of the speech processing system.
11 . The computing system of claim 8 , wherein the programming instructions further cause the computing system to perform the following operation:
selecting the licensed user based on a closest matching target speaker representation from a group of target speaker representations.
12 . The computing system of claim 8 , wherein the programming instructions further cause the computing system to perform the following operation:
selecting the licensed user based on a least matching target speaker representation from a group of target speaker representations.
13 . The computing system of claim 8 , wherein the frames of the filtered speech signal and the corresponding frames of the licensed voice signal are time-domain frames having a predefined length.
14 . A method comprising:
receiving, by a speech processing system, an original speech signal recorded by a single microphone, the original speech signal containing voice characteristics of a user;
filtering the original speech signal to obtain a filtered speech signal that excludes the voice characteristics of the user;
performing voice conversion on the original speech signal to obtain a licensed voice signal containing voice characteristics of a licensed user, the voice characteristics of the user being excluded from the filtered speech signal, wherein the filtered speech signal and the licensed voice signal form a multi-channel input;
generating a single-channel representation of the multi-channel input by combining bins of the filtered speech signal with corresponding bins of the licensed voice signal in a weighted manner such that at least some of the bins of the filtered speech signal are assigned different weights; and
performing, via an automatic speech recognition (ASR) model, speech recognition on the single-channel representation of the multi-channel input to generate a transcript of the original speech signal.
15 . The method of claim 14 , further comprising:
generating a synthetic speech signal by processing the transcript using a text-to-speech system.
16 . The method of claim 14 , wherein speech recognition is performed on the single-channel representation during run-time of the speech processing system.
17 . The method of claim 14 , further comprising:
selecting the licensed user based on a closest matching target speaker representation from a group of target speaker representations.
18 . The method of claim 14 , further comprising:
selecting the licensed user based on a least matching target speaker representation from a group of target speaker representations.
19 . The method of claim 14 , wherein the bins of the filtered speech signal and the corresponding bins of the licensed voice signal are frequency bins of a predefined size.