IP Library Granted Patent US 12,387,738
Granted Patent B2
US 12,387,738 · App. 17/848,678 · Granted Aug 12, 2025

Distributed teleconferencing using personalized enhancement models

Inventor: Ross Cutler (Clyde Hill, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L21/02H04M3/569
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,738
App. No.
17/848,678
Granted
Aug 12, 2025
Kind
B2
Abstract

This document relates to distributed teleconferencing. Some implementations can employ personalized enhancement models to enhance microphone signals for participants in a call. Further implementations can perform proximity-based mixing, where microphone signals received from devices in a particular room can be omitted from playback signals transmitted to other devices in the same room. These techniques can allow enhanced call quality for teleconferencing sessions where co-located users can employ their own devices to participate in a call with other users.

Claims (53)

1. A method comprising:

receiving a first microphone signal from a first device during a communication session, the first microphone signal being enhanced by a first personalized enhancement model that uses self-attention processing to attenuate components of the first microphone signal including first signals other than the first microphone signal during the communication session, wherein the first personalized enhancement model comprises:

a first encoder configured to process first input features using first multi-head attention to generate first intermediate representations;

a first decoder configured to generate first masking parameters from the first intermediate representations where the first masking parameters are used to selectively enhance first speech from a first user while attenuating other first audio components;

receiving a second microphone signal from a second device that is co-located with the first device during the communication session, the second microphone signal being enhanced by a second personalized enhancement model that uses self-attention processing to attenuate components of the second microphone signal including second signals other than the second microphone signal during the communication session, wherein the second personalized enhancement model comprises:

a second encoder configured to process second input features using second multi-head attention to generate second intermediate representations;

a second decoder configured to generate second masking parameters from the second intermediate representations where the second masking parameters are used to selectively enhance second speech from a second user while attenuating other second audio components;

mixing the first microphone signal with the second microphone signal to obtain a playback signal;

sending the playback signal to a third device that is participating in a call with the first device and the second device; and

sending another playback signal to at least one of the first device or the second device, the another playback signal omitting the first microphone signal and the second microphone signal.

2. The method of claim 1 , further comprising:

detecting that the first device is co-located with the second device; and

omitting the first microphone signal and the second microphone signal from the another playback signal responsive to determining that the first device is co-located with the second device.

3. The method of claim 2 , wherein the detecting is performed using ultrasound communication between the first device and the second device to determine that both devices are in a same room.

4. The method of claim 1 , further comprising:

synchronizing a first speaker of the first device with a second speaker of the second device; and

playing the playback signal using the synchronized first speaker and second speaker.

5. A computing device comprising:

a processor; and

a storage medium storing instructions which, when executed by the processor, cause the computing device to:

receive a first microphone signal from a first device during a communication session, the first microphone signal being enhanced by a first personalized enhancement model that uses self-attention processing to attenuate components of the first microphone signal including first signals other than the first microphone signal during the communication session, wherein the first personalized enhancement model comprises:

a first encoder configured to process first input features using first multi-head attention to generate first intermediate representations;

a first decoder configured to generate first masking parameters from the first intermediate representations where the first masking parameters are used to selectively enhance first speech from a first user while attenuating other first audio components;

receive a second microphone signal from a second device that is co-located with the first device during the communication session, the second microphone signal being enhanced by a second personalized enhancement model that uses self-attention processing to attenuate components of the second microphone signal including second signals other than the second microphone signal during the communication session, wherein the second personalized enhancement model comprises:

a second encoder configured to process second input features using second multi-head attention to generate second intermediate representations;

a second decoder configured to generate second masking parameters from the second intermediate representations where the second masking parameters are used to selectively enhance second speech from a second user while attenuating other second audio components;

mix the first microphone signal with the second microphone signal to obtain a playback signal;

send the playback signal to a third device that is participating in a call with the first device and the second device; and

send another playback signal to at least one of the first device or the second device, the another playback signal omitting the first microphone signal and the second microphone signal.

6. The computing device of claim 5 , wherein the instructions, when executed by the processor, cause the computing device to:

detect that the first device is co-located with the second device; and

omit the first microphone signal and the second microphone signal from the another playback signal responsive to determining that the first device is co-located with the second device.

7. The computing device of claim 6 , wherein the detecting is performed using ultrasound communication between the first device and the second device to determine that both devices are in a same room.

8. The computing device of claim 5 , wherein the instructions, when executed by the processor, cause the computing device to:

synchronize a first speaker of the first device with a second speaker of the second device; and

play the playback signal using the synchronized first speaker and second speaker.

9. A system comprising:

means for receiving a first microphone signal from a first device during a communication session, the first microphone signal being enhanced by a first personalized enhancement model that uses self-attention processing to attenuate components of the first microphone signal including first signals other than the first microphone signal during the communication session, wherein the first personalized enhancement model comprises:

a first encoder configured to process first input features using first multi-head attention to generate first intermediate representations;

a first decoder configured to generate first masking parameters from the first intermediate representations where the first masking parameters are used to selectively enhance first speech from a first user while attenuating other first audio components;

means for receiving a second microphone signal from a second device that is co-located with the first device during the communication session, the second microphone signal being enhanced by a second personalized enhancement model that uses self-attention processing to attenuate components of the second microphone signal including second signals other than the second microphone signal during the communication session, wherein the second personalized enhancement model comprises:

a second encoder configured to process second input features using second multi-head attention to generate second intermediate representations;

a second decoder configured to generate second masking parameters from the second intermediate representations where the second masking parameters are used to selectively enhance second speech from a second user while attenuating other second audio components;

means for mixing the first microphone signal with the second microphone signal to obtain a playback signal;

means for sending the playback signal to a third device that is participating in a call with the first device and the second device; and

means for sending another playback signal to at least one of the first device or the second device, the another playback signal omitting the first microphone signal and the second microphone signal.

10. The system of claim 9 , the system further comprising:

means for detecting that the first device is co-located with the second device; and

means for omitting the first microphone signal and the second microphone signal from the another playback signal responsive to determining that the first device is co-located with the second device.

11. The system of claim 10 , wherein the detecting is performed using ultrasound communication between the first device and the second device to determine that both devices are in a same room.

12. The system of claim 9 , the system further comprising:

means for synchronizing a first speaker of the first device with a second speaker of the second device; and

means for playing the playback signal using the synchronized first speaker and second speaker.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2022
From: CUTLER, ROSS
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 061246/0192 →
Continuity (1)
Related Publication 20230421702A1 · Dec 28, 2023
References Cited (20)
US 8126129B1 · McGuire · 2012 [cited by examiner]
US 9741360B1 · Li · 2017 [cited by examiner]
US 20040116130A1 · Seligmann · 2004 [cited by examiner]
US 20220084509A1 · Sivaraman et al. · 2022 [cited by applicant]
US 20220199102A1 · Ostrand · 2022 [cited by examiner]
US 20220303502A1 · Fisher · 2022 [cited by examiner]
US 20220377117A1 · Xi · 2022 [cited by examiner]
EP 1381237A2 · 2004 [cited by applicant]
WO 2021119090A1 · 2021 [cited by applicant]
WO 2022253003A1 · 2022 [cited by applicant]
WO 2023064750A1 · 2023 [cited by applicant]
Borriello, et al., “Walrus: Wireless Acoustic Location with Room-level Resolution using Ultrasound”, In Proceedings of the 3rd International Conference on Mobile Systems, Applications, and Services, Jun. 6, 2005, pp. 19… [cited by applicant]
Hu, et al., “DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement”, In Repository of arXiv:2008.00264v1, Aug. 1, 2020, 5 Pages. [cited by applicant]
Vaswani, et al., “Attention Is All You Need”, In Proceedings of Advances in Neural Information Processing Systems, Dec. 4, 2017, 11 Pages. [cited by applicant]
Williamson, et al., “Complex Ratio Masking for Monaural Speech Separation”, In Journal of IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 24, Issue 3, Mar. 2016, pp. 483-492. [cited by applicant]
“Invitation To Pay Additional Fees Issued in PCT Application No. PCT/US23/023351”, Mailed Date: Jul. 13, 2023, 11 Pages. [cited by applicant]
Cutler, et al., “ICASSP 2022 Acoustic Echo Cancellation Challenge”, In Repository of arXiv:2202.13290v1, Feb. 27, 2022, 5 Pages. [cited by applicant]
Eskimez, et al., “Personalized Speech Enhancement: New Models And Comprehensive Evaluation”, In Repository of arXiv:2110.09625v1, Oct. 18, 2021, 5 Pages. [cited by applicant]
Ruiz, et al., “Distributed Combined Acoustic Echo Cancellation And Noise Reduction Using GEVD-Based Distributed Adaptive Node Specific Signal Estimation With Prior Knowledge”, In Proceedings of 28th European Signal Proc… [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US23/023351”, Mailed Date: Sep. 5, 2023, 18 Pages. [cited by applicant]