IP Library Granted Patent US 12,248,729
Granted Patent B2
US 12,248,729 · App. 18/227,519 · Granted Mar 11, 2025

Source-based sound quality adjustment tool

Inventor: Jeffry Copps Robert Jose (Tamil Nadu, IN)
Assignee: Adeia Guides Inc.
G06F3/165G06F3/162G06F3/167H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,729
App. No.
18/227,519
Granted
Mar 11, 2025
Kind
B2
Abstract

Systems and methods for adjusting a sound level during a conference call are disclosed herein. Conferencing application receives an audio input including a first sound and a second sound. A user selectable element is generated for each sound in an user interface, where a user selection setting a first user selectable element associated with the first sound at a user-specified level is received. The sound level for the first sound is adjusted based on the user selection and output at the user-specified level while the second sound is output at a default sound level.

Claims (61)

1. A method for adjusting a sound level during a conference call involving a plurality of devices via a communications network, the method comprising:

identifying audio to be output at a first device of the plurality of devices during the conference call;

identifying a learning model that is specific to the first device and that is trained using audio data of at least one conference call;

determining, that the audio to be output at the first device comprises a first sound and a second sound, wherein the first sound is received from a human speaker;

identifying a non-human source of the second sound using the learning model that is specific to the first device;

analyzing at least a portion of the second sound to generate a label;

generating for display a first user selectable element for a volume control of the human speaker, and generating for display a second user selectable element with the generated label for controlling a volume of the non-human source of the second sound;

receiving a user selection setting the first user selectable element associated with the first sound at a first sound level different from a default sound level;

causing adjustment of the volume of the non-human source of the second sound; and

causing outputting of the first sound at the first sound level while the second sound is being output at the adjusted volume.

2. The method of claim 1 , further comprising:

identifying second audio to be output at a second device of the plurality of devices during the conference call;

identifying a learning model that is specific to the second device and that is trained using audio data of one or more conference calls, wherein the learning model that is specific to the second device is different from the learning model that is specific to the first device;

determining that the second audio comprises the first sound and the second sound; and

identifying the non-human source of the second sound using the learning model that is specific to the second device.

3. The method of claim 2 , further comprising:

analyzing a second portion of the second sound to generate a second label;

generating for display a third user selectable element for a second volume control of the human speaker, and generating for display a fourth user selectable element with a second generated label for controlling a second volume of the non-human source of the second sound;

receiving a second user selection setting the third user selectable element associated with the first sound at a second sound level different from the default sound level;

causing adjustment of a second volume of the non-human source of the second sound; and

causing outputting of the first sound at the second sound level while the second sound is being output at the adjusted second volume.

4. The method of claim 3 , wherein the first sound is caused to be output at the first sound level at the first device while the second sound is being output at the adjusted volume at the first device, and wherein the first sound is caused to be output at the second sound level at the second device while the second sound is being output at the adjusted second volume at the second device.

5. The method of claim 1 , further comprising:

applying a corresponding weight filter to the respective first and second sounds based on a user selection setting a respective user selectable element at a user-specified sound level.

6. The method of claim 5 , further comprising:

setting a respective user selectable element at the default sound level, wherein the default sound level indicates an original sound level, and wherein the default sound level is output without applying the corresponding weight filter to the respective sound.

7. The method of claim 1 , wherein the learning model is trained at least in part using combined audio data of the first and second sounds.

8. The method of claim 1 , wherein the adjusting the volume of the non-human source of the second sound comprises at least one of enhancing or reducing the second sound.

9. The method of claim 1 , wherein outputting the first sound at the first sound level is based at least in part on statistical spectral features including at least one of flatness, perceptual spread, or shape of the first sound.

10. The method of claim 1 , wherein each of the user selectable elements is a slider that moves along a respective sliding region of the respective one of the user selectable elements.

11. A system for adjusting a sound level during a conference call involving a plurality of devices via a communications network, the system comprising:

control circuitry configured to:

identify audio to be output at a first device of the plurality of devices during the conference call;

identify a learning model that is specific to the first device and that is trained using audio data of at least one conference call;

determine that the audio to be output at the first device comprises a first sound and a second sound, wherein the first sound is received from a human speaker;

identify a non-human source of the second sound using the learning model that is specific to the first device;

analyze at least a portion of the second sound to generate a label;

generate for display a first user selectable element for a volume control of the human speaker, and generating for display a second user selectable element with the generated label for controlling a volume of the non-human source of the second sound;

receive a user selection setting the first user selectable element associated with the first sound at a first sound level different from a default sound level;

cause adjustment of the volume of the non-human source of the second sound; and

cause output of the first sound at the first sound level while the second sound is being output at the adjusted volume.

12. The system of claim 11 , wherein the control circuitry is further configured to:

a identify second audio to be output at a second device of the plurality of devices during the conference call;

identify a learning model that is specific to the second device and that is trained using audio data of one or more conference calls, wherein the learning model that is specific to the second device is different from the learning model that is specific to the first device;

determine, by control circuitry of the second device, that the second audio comprises the first sound and the second sound; and

identify the non-human source of the second sound using the learning model that is specific to the second device.

13. The system of claim 12 the control circuitry further configured to:

analyze a second portion of the second sound to generate a second label;

generate for display a third user selectable element for a second volume control of the human speaker, and generating for display a fourth user selectable element with a second generated label for controlling a second volume of the non-human source of the second sound;

receive a second user selection setting the third user selectable element associated with the first sound at a second sound level different from the default sound level;

cause adjustment of a second volume of the non-human source of the second sound; and

cause output of the first sound at the second sound level while the second sound is being output at the adjusted second volume.

14. The system of claim 13 , wherein the control circuitry is configured to cause the first sound to be output at the first sound level at the first device while the second sound is being output at the adjusted volume at the first device, and wherein the control circuitry is configured to cause the first sound at the second sound level while the second sound is being output at the adjusted second at the second device.

15. The system of claim 11 , the control circuitry further configured to:

apply a corresponding weight filter to the respective first and second sounds based on a user selection setting a respective user selectable element at a user-specified sound level.

16. The system of claim 15 , the control circuitry further configured to:

set a respective user selectable element at the default sound level, wherein the default sound level indicates an original sound level, and wherein the default sound level is output without applying the corresponding weight filter to the respective sound.

17. The system of claim 11 , wherein the learning model is trained at least in part using combined audio data of the first and second sounds.

18. The system of claim 11 , wherein the adjusting the volume of the non-human source of the second sound comprises at least one of enhancing or reducing the second sound.

19. The system of claim 11 , wherein outputting the first sound at the first sound level is based at least in part on statistical spectral features including at least one of flatness, perceptual spread, or shape of the first sound.

20. The system of claim 11 , wherein each of the user selectable elements is a slider that moves along a respective sliding region of the respective one of the user selectable elements.

Assignments (2)
CHANGE OF NAME Recorded Oct 4, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069113/0323 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2023
From: ROBERT JOSE, JEFFRY COPPS
To: ROVI GUIDES, INC.
Reel/Frame 064426/0396 →
Continuity (2)
Continuation 17180254 · Feb 19, 2021
Related Publication 20230367543A1 · Nov 16, 2023
References Cited (3)
US 20170187884A1 · Minor · 2017 [cited by examiner]
US 20210224319A1 · Ingel · 2021 [cited by examiner]
US 20220269473A1 · Robert Jose · 2022 [cited by applicant]