IP Library Granted Patent US 11,670,317
Granted Patent B2
US 11,670,317 · App. 17/182,506 · Granted Jun 6, 2023

Dynamic audio quality enhancement

Inventors: Sayan Acharya Ghosh (Singapore, SG); Prashant Jain (Singapore, SG); Ai Kiar Ang (Singapore, SG); Gary Kim Chwee Lim (Singapore, SG)
Assignee: KYNDRYL, INC.
G10L21/02G06F3/165G10L25/18G10L25/60G10L25/78H04R3/02H04R29/001H04R29/004
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,670,317
App. No.
17/182,506
Granted
Jun 6, 2023
Kind
B2
Abstract

A set of user pools can be determined based on location data associated with each device in an audio/video (A/V) conference. A key active user can be determined for each user pool of the set of user pools based on valid audio signals received from each device within each user pool. A determination can be made whether there is feedback within each user pool. Responsive to determining feedback in at least one user pool, speakers of devices within the at least one user pool can be disconnected except for the key active user device within each respective user pool.

Claims (48)

1. A method comprising:

determining, based on location data associated with a plurality of devices in an audio/video (A/V) conference, a set of user pools, the set of user pools including the plurality of devices and including, at least, one user pool and another user pool;

determining, based on active valid audio signals received from at least a portion of the plurality of devices of the set of user pools, at least, one key active user device for the one user pool and another key active user device for the another user pool;

determining that there is feedback within the one user pool, wherein the feedback is identified based on repeating frequencies amplified and propagated through a speaker to microphone feedback loop;

disconnecting, in response to determining that there is feedback in the one user pool, speakers of devices within the one user pool except the one key active user device within the one user pool; and

performing the determining and the disconnecting for, at least, the another user pool.

2. The method of claim 1 , further comprising:

calculating a time delay in which audio is received for each device within each user pool with reference to each key active user device; and

disconnecting, for each device where the time delay exceeds a time delay threshold, microphones of each device emitting audio that exceeded the time delay threshold.

3. The method of claim 2 , further comprising:

enhancing, for each device where the time delay is within the time delay threshold, audio by aligning the audio received from each device where the time delay is within the time delay threshold with the audio output by the key active user device.

4. The method of claim 3 , further comprising: playing back the enhanced audio to the A/V conference.

5. The method of claim 4 , wherein the played back audio is attenuated for users within a same user pool.

6. The method of claim 3 , wherein audio is aligned by converting analog signals representing the audio to a frequency domain using Fast Fourier Transform (FFT).

7. The method of claim 1 , wherein the one key active user device is determined based on the one key active user device having a highest elapsed time period of valid audio signals based on silence detection.

8. The method of claim 1 , wherein for each user pool of the set of user pools, based on determining that there is feedback in a given user pool, a speaker of a key active user device of the given user pool remains active and speakers of other devices of the given user pool are disconnected.

9. A system comprising:

one or more processors; and

one or more computer-readable storage media storing program instructions which, when executed by the one or more processors, are configured to cause the one or more processors to perform a method comprising:

determining, based on location data associated with a plurality of devices in an audio/video (A/V) conference, a set of user pools, the set of user pools including the plurality of devices and including, at least, one user pool and another user pool;

determining, based on active valid audio signals received from at least a portion of the plurality of devices of the set of user pools, at least, one key active user device for the one user pool and another key active user device for the another user pool;

determining that there is feedback within the one user pool, wherein the feedback is identified based on repeating frequencies amplified and propagated through a speaker to microphone feedback loop;

disconnecting, in response to determining that there is feedback in the one user pool, speakers of devices within the one user pool except the one key active user device within the one user pool; and

performing the determining and the disconnecting for, at least, the another user pool.

10. The system of claim 9 , wherein the method performed by the one or more processors further comprises:

calculating a time delay in which audio is received for each device within each user pool with reference to each key active user device; and

disconnecting, for each device where the time delay exceeds a time delay threshold, microphones of each device emitting audio that exceeded the time delay threshold.

11. The system of claim 10 , wherein the method performed by the one or more processors further comprises:

enhancing, for each device where the time delay is within the time delay threshold, audio by aligning the audio received from each device where the time delay is within the time delay threshold with the audio output by the key active user device.

12. The system of claim 11 , wherein the method performed by the one or more processors further comprises:

playing back the enhanced audio to the A/V conference.

13. The system of claim 12 , wherein the played back audio is attenuated for users within a same user pool.

14. The system of claim 11 , wherein audio is aligned by converting analog signals representing the audio to a frequency domain using Fast Fourier Transform (FFT).

15. The system of claim 9 , wherein the one key active user device is determined based on the one key active user device having a highest elapsed time period of valid audio signals based on silence detection.

16. A computer program product comprising one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions comprising instructions configured to cause one or more processors to perform a method comprising:

determining, based on location data associated with a plurality of devices in an audio/video (A/V) conference, a set of user pools, the set of user pools including the plurality of devices and including, at least, one user pool and another user pool;

determining, based on active valid audio signals received from at least a portion of the plurality of devices of the set of user pools, at least, one key active user device for the one user pool and another key active user device for the another user pool;

determining that there is feedback within the one user pool, wherein the feedback is identified based on repeating frequencies amplified and propagated through a speaker to microphone feedback loop;

disconnecting, in response to determining that there is feedback in the one user pool, speakers of devices within the one user pool except the one key active user device within the one user pool; and

performing the determining and the disconnecting for, at least, the another user pool.

17. The computer program product of claim 16 , wherein the method performed by the one or more processors further comprises:

calculating a time delay in which audio is received for each device within each user pool with reference to each key active user device; and

disconnecting, for each device where the time delay exceeds a time delay threshold, microphones of each device emitting audio that exceeded the time delay threshold.

18. The computer program product of claim 17 , wherein the method performed by the one or more processors further comprises:

enhancing, for each device where the time delay is within the time delay threshold, audio by aligning the audio received from each device where the time delay is within the time delay threshold with the audio output by the key active user device.

19. The computer program product of claim 18 , wherein the method performed by the one or more processors further comprises:

playing back the enhanced audio to the A/V conference.

20. The computer program product of claim 16 , wherein the one key active user device is determined based on the one key active user device having a highest elapsed time period of valid audio signals based on silence detection.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2021
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: KYNDRYL, INC.
Reel/Frame 058213/0912 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2021
From: GHOSH, SAYAN ACHARYA; JAIN, PRASHANT; ANG, AI KIAR; LIM, GARY KIM CHWEE
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 055367/0972 →
Continuity (1)
Related Publication 20220270628A1 · Aug 25, 2022