Signal processing device and signal processing method
To satisfactorily perform processing of increasing the sound quality of a recorded sound source obtained by picking up vocal sound and musical instrument sound in a room. An output audio signal is obtained by a sound converter performing sound conversion processing on a recorded sound source (an input audio signal) obtained by picking up vocal sound or musical instrument sound by using any microphone in any room. The sound conversion processing includes processing of removing room reverberation from the recorded sound source, processing of remove picked-up sound noise from the recorded sound source, processing of including target microphone characteristics into the recorded sound source, and processing of including the target studio characteristics into the recorded sound source.
1 . A signal processing device, comprising:
a sound converter configured to:
perform a first process on an input audio signal to obtain an output signal, wherein
the input audio signal is from a first microphone,
the first microphone picks up one of vocal sound or musical instrument sound in a first room, and
the first process is a process of removal of room reverberation from the input audio signal; and
perform a second process on the output signal to obtain an output audio signal, wherein
the second process is a process of inclusion of characteristics of a target microphone and characteristics of an anechoic room into the output signal.
2 . The signal processing device according to claim 1 , wherein the first process of removing the room reverberation is performed using a deep neural network trained to remove the room reverberation.
3 . The signal processing device according to claim 2 , wherein the deep neural network has been trained in such a manner that uses as a deep neural network input an audio signal with the room reverberation obtained by convolving a dry input with a room reverberation impulse response generated by causing a reference speaker to output sound in a room based on a Time stretched Pulse (TSP) signal and then picking up the sound with the first microphone, and feeds back a difference displacement of a deep neural network output in response to the dry input to parameters.
4 . The signal processing device according to claim 1 , wherein
the sound converter is further configured to perform a third process on the input audio signal, and
the third process is a process of removal of picked-up sound noise from the input audio signal.
5 . The signal processing device according to claim 4 , wherein the third process of removing the picked-up sound noise is performed using a deep neural network trained to remove the picked-up sound noise.
6 . The signal processing device according to claim 5 , wherein the deep neural network has been trained in such a manner that uses as a deep neural network input an audio signal obtained by adding noise picked up with the first microphone to a dry input, and feeds back a difference displacement of a deep neural network output in response to the dry input to parameters.
7 . The signal processing device according to claim 5 , wherein the deep neural network has been trained in such a manner that uses as a deep neural network input an audio signal obtained by adding the picked-up sound noise picked up with the first microphone to an audio signal with room reverberation obtained by convolving a dry input with a room reverberation impulse response generated by causing a reference speaker to output sound in a room based on a Time stretched Pulse (TSP) signal and then picking up the sound with the first microphone, and feeds back a difference displacement of a deep neural network output in response to the audio signal with the room reverberation to parameters.
8 . The signal processing device according to claim 4 , wherein simultaneously with the first process of removing the room reverberation, the third process of removing the picked-up sound noise is performed using a deep neural network trained to remove the room reverberation and the picked-up sound noise.
9 . The signal processing device according to claim 8 , wherein the deep neural network has been trained in such a manner that uses as a deep neural network input an audio signal obtained by adding the picked-up sound noise picked up with the first microphone to an audio signal with the room reverberation obtained by convolving a dry input with a room reverberation impulse response generated by causing a reference speaker to output sound in a room based on a Time stretched Pulse (TSP) signal and then picking up the sound with the first microphone, and feeds back a difference displacement of a deep neural network output in response to the dry input to parameters.
10 . The signal processing device according to claim 1 , wherein
the sound converter is further configured to perform the second process based on convolution of the output signal with an impulse response, and
the impulse response includes the characteristics of the target microphone and the characteristics of the anechoic room.
11 . The signal processing device according to claim 10 , wherein the sound converter is further configured to:
control a reference speaker to output first sound, wherein
the output of the first sound is based on a time stretched pulse (TSP) signal, and
the target microphone picks up the outputted first sound; and
generate the impulse response based on the outputted first sound.
12 . The signal processing device according to claim 1 , wherein the second process of including the characteristics of the target microphone is performed by convolving the input audio signal with an impulse response for the characteristics of the target microphone and then using a deep neural network trained to include non-linear characteristics of the target microphone.
13 . The signal processing device according to claim 12 , wherein the impulse response for the characteristics of the target microphone is generated by causing a reference speaker to output first sound based on a TSP signal and then picking up the outputted first sound with the target microphone, and
the deep neural network has been trained in such a manner that uses as a deep neural network input an audio signal obtained by convolving with the impulse response for the characteristics of the target microphone, and feeds back to parameters a difference displacement of a deep neural network output in response to the audio signal obtained by causing the reference speaker to output second sound based on a dry input and then picking up the outputted second sound with the target microphone.
14 . The signal processing device according to claim 1 , wherein the second process of including the characteristics of the target microphone is performed using a deep neural network trained to include both linear and non-linear characteristics of the target microphone into the input audio signal.
15 . The signal processing device according to claim 14 , wherein the deep neural network has been trained in such a manner that uses a dry input as a deep neural network input, and feeds back to parameters a difference displacement of a deep neural network output in response to an audio signal obtained by causing a reference speaker to output sound based on the dry input and then picking up the outputted sound with the target microphone.
16 . The signal processing device according to claim 1 , wherein
the sound converter is further configured to perform a fourth process, and
the fourth process is a process of inclusion of characteristics of a target studio into the input audio signal.
17 . The signal processing device according to claim 16 , wherein
the sound converter is further configured to perform the fourth process based on convolution of the input audio signal with an impulse response, and
the impulse response includes the characteristics of the target studio.
18 . A signal processing method, comprising:
a first process on an input audio signal to obtain an output signal, wherein
the input audio signal is from a first microphone,
the first microphone picks up one of vocal sound or musical instrument sound in a first room, and
the first process is a process of removal of room reverberation from the input audio signal; and
performing a second process on the output signal to obtain an output audio signal, wherein
the second process is a process of inclusion of characteristics of a target microphone and characteristics of an anechoic room into the output signal.
19 . A non-transitory computer-readable medium having stored thereon, computer-executable instructions which, when executed by a computer, cause the computer to execute operations, the operations comprising:
performing a first process on an input audio signal to obtain an output signal, wherein
the input audio signal is from a first microphone,
the first microphone picks up one of vocal sound or musical instrument sound in a first room, and
the first process is a process of removal of room reverberation from the input audio signal; and
perform a second process on the output signal to obtain an output audio signal, wherein
the second process is a process of inclusion of characteristics of a target microphone and characteristics of an anechoic room into the output signal.