IP Library Granted Patent US 12,581,038
Granted Patent B2
US 12,581,038 · App. 18/543,126 · Granted Mar 17, 2026

Audio processing in video conferencing system using multimodal features

Inventors: Valdemar Ettrup Larsen (Ballerup, DK); Henning Toft Schwarz (Ballerup, DK); Elias Lundgaard Pedersen (Ballerup, DK); Robert James (Ballerup, DK); Anuj Dutt (Ballerup, DK)
Assignee: GN HEARING A/S
H04N7/157H04N7/147H04N7/152
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,581,038
App. No.
18/543,126
Granted
Mar 17, 2026
Kind
B2
Abstract

The disclosure relates to a method for processing audio signals in a video conference call. At first, virtual room boundaries defining space relevant for the video conference call are determined. By at least one input transducer, audio data are obtained and by at least one camera, video data are obtained. The video data and audio data are then correlated. On the basis of the correlation and defined virtual room boundaries, the audio data are modified and modified output audio data are generated.

Claims (37)

1 . A method for processing audio signals in a video conference call, the method comprising:

determining virtual room boundaries that define space relevant for the video conference call based on an initial audio signal that includes background noise and an initial video signal,

obtaining, by at least one input transducer, audio data;

obtaining, by at least one camera, video data;

correlating the video data and audio data;

modifying the obtained audio data on the basis of the correlation and the determined virtual room boundaries to generate output audio data.

2 . The method of claim 1 , further comprising:

determining one or more speech signals present in the audio data.

3 . The method according to claim 2 , further comprising:

extracting, from the determined speech signals, directions of arrival of the speech signals.

4 . The method of claim 1 , further comprising:

determining, based on the video data, a presence of one or more meeting participants.

5 . The method of claim 4 , wherein the correlating of the video and audio data includes correlating the one or more speech signals and the presence of one or more meeting participants.

6 . The method according to claim 4 , wherein the determining of the presence of one or more meeting participants comprises:

calculating relative positions of each identified participant,

wherein the relative positions include at least one of the distance or the angle of a participant from the at least one camera.

7 . The method according to claim 6 , wherein the correlating of the video and audio data includes correlating the directions of arrival of the speech signals and the relative positions of each identified participant.

8 . The method according to claim 4 , further comprising:

identifying movements of the one or more participants within the virtual room boundaries and dynamically adapting participant related parameters based on the movements.

9 . The method according to claim 1 , further comprising:

assigning an identity number to each determined participant.

10 . The method according to claim 1 , further comprising:

attenuating any audio signal originating from outside of the virtual room boundaries.

11 . The method according to claim 1 , further comprising:

stopping one or more adaptive algorithms from adapting to speech detected outside of the virtual room boundaries.

12 . The method according to claim 1 , further comprising:

determining a biomarker parameter for each participant, and

including the biomarker parameter in correlation of the video data and audio data.

13 . The method according to claim 12 , wherein the biomarker parameter comprises at least one of a lip movement parameter, gaze parameter, a face parameter, a body orientation parameter, a face landmark parameter, or a body landmark parameter.

14 . The method according to claim 1 , wherein the method is performed by a processing unit.

15 . A video conference device for processing audio signals in a video conference call, the device comprising:

one or more input transducers for obtaining audio data,

one or more cameras for obtaining video data, and

a processing unit configured to:

determine one or more virtual room boundaries that define space relevant for a video conference call based on an initial audio signal included in the obtained audio data that includes background noise and an initial video signal included in the obtained video data,

determine a correlation between the obtained video data and the obtained audio data, and

modify the obtained audio data based on the determined correlation and determined virtual room boundaries to generate output audio data.

Assignments (3)
MERGER Recorded Mar 30, 2026
From: GN AUDIO A/S
To: GN HEARING A/S
Reel/Frame 075299/0225 →
MERGER Recorded Jan 28, 2025
From: GN AUDIO A/S
To: GN HEARING A/S
Reel/Frame 070026/0456 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2024
From: LARSEN, VALDEMAR ETTRUP; SCHWARZ, HENNING TOFT; PEDERSEN, ELIAS LUNDGAARD; JAMES, ROBERT; DUTT, ANUJ
To: GN AUDIO A/S
Reel/Frame 067167/0375 →
Continuity (1)
Related Publication 20250203045A1 · Jun 19, 2025
References Cited (16)
US 20150156598A1 · Sun et al. · 2015 [cited by applicant]
US 20170098453A1 · Wright et al. · 2017 [cited by applicant]
US 20200344278A1 · Mackell et al. · 2020 [cited by applicant]
US 20200412772A1 · Nesta · 2020 [cited by examiner]
US 20210051037A1 · Atkins · 2021 [cited by examiner]
US 20210345040A1 · Meyer · 2021 [cited by examiner]
US 20220086564A1 · Bou Daher · 2022 [cited by examiner]
US 20220210341A1 · Hwang · 2022 [cited by examiner]
US 20220382907A1 · Siohan et al. · 2022 [cited by applicant]
US 20240071356A1 · Gu · 2024 [cited by examiner]
US 20240119731A1 · Hammer · 2024 [cited by examiner]
US 20240185449A1 · Bhatt · 2024 [cited by examiner]
US 20240185876A1 · Zhang et al. · 2024 [cited by applicant]
US 20240214520A1 · Dao · 2024 [cited by examiner]
WO 2022262316A1 · 2022 [cited by applicant]
The extended European search report issued in European Application No. 24153086.4, dated Jun. 27, 2024. [cited by applicant]