Realistic room acoustic simulation for videoconferencing
Systems and methods for synthetic audio datasets generation for videoconferencing based on realistic room acoustic simulation are provided. For example, a room-independent recording for a source audio signal and a particular type of microphones can be generated and a target room setup for a target room can be obtained. The target room setup specifying one or more of a size of the target room, respective locations of a speaker and a microphone in the target room. Room characteristics of the target room can be generated based on the target room setup via a room acoustic model. A synthetic audio signal for the target room and the particular type of microphones can be generated by applying the room acoustic characteristics onto the room-independent recording.
1 . A method for generating a synthetic audio signal for a target room, the method comprising:
generating a room-independent recording for a source audio signal and a particular type of microphones;
obtaining a target room setup for a target room, the target room setup specifying one or more of a size of the target room, respective locations of a speaker and a microphone in the target room;
generating, via a room acoustic model, room acoustic characteristics of the target room based on the target room setup; and
generating a synthetic audio signal for the target room and the particular type of microphones by applying the room acoustic characteristics onto the room-independent recording.
2 . The method of claim 1 , wherein generating the room-independent recording comprises obtaining a near-field recording of the source audio signal using a microphone of the particular type placed in proximity of a speaker playing the source audio signal.
3 . The method of claim 1 , wherein generating the room-independent recording comprises providing the source audio signal and a microphone-speaker distance to a machine learning mode, the machine learning model being trained to generate room-independent recordings for the particular type of microphones at a given distance.
4 . The method of claim 1 , wherein the target room setup is estimated from a target room audio sample recorded in the target room.
5 . The method of claim 1 , wherein the room acoustic characteristics of the target room comprise a room impulse response (RIR).
6 . The method of claim 5 , wherein the room acoustic model comprises an improved image source model and wherein a perturbation level of noise added to each reflection of the improved image source model increases as a number of reflections increases.
7 . The method of claim 5 , wherein the room acoustic model comprises an improved image source model and wherein the RIR is adjusted using a mask reducing an amplitude of an early reverberation of the RIR and preserving a late reverberation of the RIR.
8 . The method of claim 5 , wherein generating the synthetic audio signal comprises convolving the room-independent recording with the room acoustic characteristics.
9 . The method of claim 1 , further comprising causing the synthetic audio signal to be used as training data or testing data for an audio processing model.
10 . The method of claim 1 , wherein the source audio signal comprises human speeches.
11 . A system comprising:
a non-transitory computer-readable medium; and
a processor communicatively coupled to the non-transitory computer-readable medium, the processor configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:
generate a room-independent recording for a source audio signal and a particular type of microphones;
obtain a target room setup for a target room, the target room setup specifying one or more of a size of the target room, respective locations of a speaker and a microphone in the target room;
generate, via a room acoustic model, room acoustic characteristics of the target room based on the target room setup; and
generate a synthetic audio signal for the target room and the particular type of microphones by applying the room acoustic characteristics onto the room-independent recording.
12 . The system of claim 11 , wherein generating the room-independent recording comprises obtaining a near-field recording of the source audio signal using a microphone of the particular type placed in proximity of a speaker playing the source audio signal.
13 . The system of claim 11 , wherein generating the room-independent recording comprises providing the source audio signal and a microphone-speaker distance to a machine learning mode, the machine learning model being trained to generate room-independent recordings for the particular type of microphones at a given distance.
14 . The system of claim 11 , wherein the room acoustic characteristics of the target room comprise a room impulse response (RIR), and wherein the room acoustic model comprises an improved image source model and wherein a perturbation level of noise added to each reflection of the improved image source model increases as a number of reflections increases and the RIR is adjusted using a mask reducing an amplitude of an early reverberation of the RIR and preserving a late reverberation of the RIR.
15 . The system of claim 14 , wherein generating the synthetic audio signal comprises convolving the room-independent recording with the room acoustic characteristics.
16 . A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to:
generate a room-independent recording for a source audio signal and a particular type of microphones;
obtain a target room setup for a target room, the target room setup specifying one or more of a size of the target room, respective locations of a speaker and a microphone in the target room;
generate, via a room acoustic model, room acoustic characteristics of the target room based on the target room setup; and
generate a synthetic audio signal for the target room and the particular type of microphones by applying the room acoustic characteristics onto the room-independent recording.
17 . The non-transitory computer-readable medium of claim 16 , wherein generating the room-independent recording comprises obtaining a near-field recording of the source audio signal using a microphone of the particular type placed in proximity of a speaker playing the source audio signal.
18 . The non-transitory computer-readable medium of claim 16 , wherein generating the room-independent recording comprises providing the source audio signal and a microphone-speaker distance to a machine learning mode, the machine learning model being trained to generate room-independent recordings for the particular type of microphones at a given distance.
19 . The non-transitory computer-readable medium of claim 16 , wherein the room acoustic characteristics of the target room comprise a room impulse response (RIR), and wherein the room acoustic model comprises an improved image source model and wherein a perturbation level of noise added to each reflection of the improved image source model increases as a number of reflections increases and the RIR is adjusted using a mask reducing an amplitude of an early reverberation of the RIR and preserving a late reverberation of the RIR.
20 . The non-transitory computer-readable medium of claim 19 , wherein generating the synthetic audio signal comprises convolving the room-independent recording with the room acoustic characteristics.