IP Library Granted Patent US 12,671,958
Granted Patent B1
US 12,671,958 · App. 18/425,875 · Granted Jun 30, 2026

Realistic room acoustic simulation for videoconferencing

Inventors: Yuhui Chen (San Jose, CA); Yifeng Fan (Foster City, CA); Zhaofeng Jia (Saratoga, CA)
Assignee: Zoom Communications, Inc.
H04S7/305H04S2400/13H04S2400/15H04S2420/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,671,958
App. No.
18/425,875
Filed
Jan 29, 2024
Granted
Jun 30, 2026
Kind
B1
Examiner
KIM, PAUL
Art Unit
2695
USPC
381/303
Abstract

Systems and methods for synthetic audio datasets generation for videoconferencing based on realistic room acoustic simulation are provided. For example, a room-independent recording for a source audio signal and a particular type of microphones can be generated and a target room setup for a target room can be obtained. The target room setup specifying one or more of a size of the target room, respective locations of a speaker and a microphone in the target room. Room characteristics of the target room can be generated based on the target room setup via a room acoustic model. A synthetic audio signal for the target room and the particular type of microphones can be generated by applying the room acoustic characteristics onto the room-independent recording.

Claims (34)

1 . A method for generating a synthetic audio signal for a target room, the method comprising:

generating a room-independent recording for a source audio signal and a particular type of microphones;

obtaining a target room setup for a target room, the target room setup specifying one or more of a size of the target room, respective locations of a speaker and a microphone in the target room;

generating, via a room acoustic model, room acoustic characteristics of the target room based on the target room setup; and

generating a synthetic audio signal for the target room and the particular type of microphones by applying the room acoustic characteristics onto the room-independent recording.

2 . The method of claim 1 , wherein generating the room-independent recording comprises obtaining a near-field recording of the source audio signal using a microphone of the particular type placed in proximity of a speaker playing the source audio signal.

3 . The method of claim 1 , wherein generating the room-independent recording comprises providing the source audio signal and a microphone-speaker distance to a machine learning mode, the machine learning model being trained to generate room-independent recordings for the particular type of microphones at a given distance.

4 . The method of claim 1 , wherein the target room setup is estimated from a target room audio sample recorded in the target room.

5 . The method of claim 1 , wherein the room acoustic characteristics of the target room comprise a room impulse response (RIR).

6 . The method of claim 5 , wherein the room acoustic model comprises an improved image source model and wherein a perturbation level of noise added to each reflection of the improved image source model increases as a number of reflections increases.

7 . The method of claim 5 , wherein the room acoustic model comprises an improved image source model and wherein the RIR is adjusted using a mask reducing an amplitude of an early reverberation of the RIR and preserving a late reverberation of the RIR.

8 . The method of claim 5 , wherein generating the synthetic audio signal comprises convolving the room-independent recording with the room acoustic characteristics.

9 . The method of claim 1 , further comprising causing the synthetic audio signal to be used as training data or testing data for an audio processing model.

10 . The method of claim 1 , wherein the source audio signal comprises human speeches.

11 . A system comprising:

a non-transitory computer-readable medium; and

a processor communicatively coupled to the non-transitory computer-readable medium, the processor configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:

generate a room-independent recording for a source audio signal and a particular type of microphones;

obtain a target room setup for a target room, the target room setup specifying one or more of a size of the target room, respective locations of a speaker and a microphone in the target room;

generate, via a room acoustic model, room acoustic characteristics of the target room based on the target room setup; and

generate a synthetic audio signal for the target room and the particular type of microphones by applying the room acoustic characteristics onto the room-independent recording.

12 . The system of claim 11 , wherein generating the room-independent recording comprises obtaining a near-field recording of the source audio signal using a microphone of the particular type placed in proximity of a speaker playing the source audio signal.

13 . The system of claim 11 , wherein generating the room-independent recording comprises providing the source audio signal and a microphone-speaker distance to a machine learning mode, the machine learning model being trained to generate room-independent recordings for the particular type of microphones at a given distance.

14 . The system of claim 11 , wherein the room acoustic characteristics of the target room comprise a room impulse response (RIR), and wherein the room acoustic model comprises an improved image source model and wherein a perturbation level of noise added to each reflection of the improved image source model increases as a number of reflections increases and the RIR is adjusted using a mask reducing an amplitude of an early reverberation of the RIR and preserving a late reverberation of the RIR.

15 . The system of claim 14 , wherein generating the synthetic audio signal comprises convolving the room-independent recording with the room acoustic characteristics.

16 . A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to:

generate a room-independent recording for a source audio signal and a particular type of microphones;

obtain a target room setup for a target room, the target room setup specifying one or more of a size of the target room, respective locations of a speaker and a microphone in the target room;

generate, via a room acoustic model, room acoustic characteristics of the target room based on the target room setup; and

generate a synthetic audio signal for the target room and the particular type of microphones by applying the room acoustic characteristics onto the room-independent recording.

17 . The non-transitory computer-readable medium of claim 16 , wherein generating the room-independent recording comprises obtaining a near-field recording of the source audio signal using a microphone of the particular type placed in proximity of a speaker playing the source audio signal.

18 . The non-transitory computer-readable medium of claim 16 , wherein generating the room-independent recording comprises providing the source audio signal and a microphone-speaker distance to a machine learning mode, the machine learning model being trained to generate room-independent recordings for the particular type of microphones at a given distance.

19 . The non-transitory computer-readable medium of claim 16 , wherein the room acoustic characteristics of the target room comprise a room impulse response (RIR), and wherein the room acoustic model comprises an improved image source model and wherein a perturbation level of noise added to each reflection of the improved image source model increases as a number of reflections increases and the RIR is adjusted using a mask reducing an amplitude of an early reverberation of the RIR and preserving a late reverberation of the RIR.

20 . The non-transitory computer-readable medium of claim 19 , wherein generating the synthetic audio signal comprises convolving the room-independent recording with the room acoustic characteristics.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2026
From: CHEN, YUHUI; FAN, YIFENG; JIA, ZHAOFENG
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 074861/0341 →
References Cited (2)
US 10616706B1 · Amengual Gari · 2020 [cited by examiner]
US 20080232603A1 · Soulodre · 2008 [cited by examiner]