IP Library Granted Patent US 11,848,655
Granted Patent B1
US 11,848,655 · App. 17/475,550 · Granted Dec 19, 2023

Multi-channel volume level equalization based on user preferences

Inventors: Mohammed Khalilia (Lynnwood, WA); Naveen Sudhakaran Nair (Issaquah, WA)
Assignee: Amazon Technologies, Inc.
H03G5/025G06F3/165G06N20/00G10L17/04G10L25/63G10L25/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,848,655
App. No.
17/475,550
Granted
Dec 19, 2023
Kind
B1
Abstract

Systems, devices, and methods are provided for multi-stem volume equalization, wherein the volume levels of each stem may be adjusted non-uniformly. Audio may be diarized into a plurality of stems, including background noise separate. Mean and variance of the volume levels of the stems may be computed. Each audio stem may be automatically adjusted based on a stem-specific preference that a user may specify. View may adjust actor volume relative to the mean/variance that maintains a relative difference in volume levels between stems.

Claims (66)

1. A computer-implemented method, comprising:

diarizing multimedia content into a plurality of stems, wherein the multimedia content comprises audio content and video content;

computing mean and variance volume information for the plurality of stems;

receiving, a request to play the multimedia content for a user;

determining a first set of target volume levels for the plurality of stems;

causing, at a first time, the multimedia content to be played for the user according to the first set of target volumes;

receiving, at a second time, a volume control command from the user;

determining speech contextual information of the audio content being played at the second time;

recording user feedback information based on the received volume control command;

training a machine-learning model based on the multimedia content and the user feedback to infer a second set of target volume levels for the plurality of stems; and

performing a relative volume level adjustment between first and second stems of the plurality of stems based on the second set of target volume levels.

2. The computer-implemented method of claim 1 , further comprising:

providing, to the user, a graphical interface comprising a plurality of volume controls;

receiving, from the user, an indication to increase volume of the first stem; and

increasing a first volume level for the first stem and increasing a second volume level for the second stem, wherein a variance between the first stem and second stem is maintained.

3. The computer-implemented method of claim 1 , wherein determining the speech contextual information comprises analyzing video being played at the second time to determine emotion information.

4. The computer-implemented method of claim 1 , wherein

the machine-learning model comprises an autoencoder and a generative adversarial network; and

training the machine-learning model comprises performing reinforcement learning using the autoencoder and the generative adversarial network to determine the second set of target volume levels for the plurality of stems.

5. A system, comprising:

one or more processors; and

memory storing executable instructions that, as a result of execution by the one or more processors, cause the system to:

determine a plurality of stems for digital content comprising audio;

receive, at a first time, a volume control command from a user during playback of the digital content;

in response to the volume control command:

determining, using a machine-learning model, a first volume adjustment to a first stem of the digital content;

determining, using the machine-learning model, a second volume adjustment to a second stem of the digital content; and

wherein the first volume adjustment and the second volume adjustment maintain a relative difference in volume levels between the first stem and the second stem;

wherein the machine-learning model is trained to infer a set of target volume levels for the plurality of stems based on historical volume control commands issued by the user; and

perform a relative volume level adjustment between the first and second stems based on the set of target volume levels.

6. The system of claim 5 , wherein executable instructions include further instructions that, as a result of execution by the one or more processors, further cause the system to:

diarize the digital content into a plurality of stems;

compute mean and variance volume information for the plurality of stems; and

determine the set of target volume levels based at least in part on the computed mean and variance volume information.

7. The system of claim 5 , wherein the set of target volume levels is determined further based at least in part on paralinguistic information detected at the first time.

8. The system of claim 7 , wherein the paralinguistic information is processed based on video of the digital content at the first time.

9. The system of claim 5 , wherein:

a variance in volume between the first stem and the second stem is maintained; and

an average volume of the digital content is increased or decreased in response to the volume control command.

10. The system of claim 5 , wherein the first volume adjustment is an increase in volume of the first stem and the second volume adjustment is a decrease in volume of the second stem.

11. The system of claim 5 , wherein:

the first stem is associated with a first speaker identity and the second stem is associated with a second stem identity; and

the machine-learning model is trained further based on the first speaker identity and the second speaker identity.

12. The system of claim 11 , wherein at least one stem of the plurality of stems corresponds to the first speaker identity corresponds to a first latent representation of one or more paralinguistic characteristics.

13. A non-transitory computer-readable storage medium storing executable instructions that, as a result of being executed by one or more processors of a computer system, cause the computer system to at least:

determine a plurality of stems for digital content comprising audio;

receive, at a first time, a volume control command from a user during playback of the digital content;

in response to the volume control command:

determining, using a machine-learning model, a first volume adjustment to a first stem of the digital content;

determining, using the machine-learning model, a second volume adjustment to a second stem of the digital content; and

wherein the first volume adjustment and the second volume adjustment maintain a relative difference in volume levels between the first stem and the second stem; and

wherein the machine-learning model is trained to infer a set of target volume levels for the plurality of stems based on historical volume control commands issued by the user; and

perform a relative volume level adjustment between the first and second stems based on the set of target volume levels.

14. The non-transitory computer-readable storage medium of claim 13 , instructions, as a result of being executed by the one or more processors of the computer system, further cause the system to:

perform speaker diarization and background noise separate of the digital content to determine a plurality of stems;

compute mean and variance volume information for the plurality of stems; and

determine the set of target volume levels based at least in part on the computed mean and variance volume information.

15. The non-transitory computer-readable storage medium of claim 13 , wherein the set of target volume levels is determined further based at least in part on emotion context detected at the first time.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the emotion context is processed based on video of the multimedia content at the first time.

17. The non-transitory computer-readable storage medium of claim 13 , wherein:

a variance in volume between the first stem and the second stem is maintained; and

an average volume of the digital content is increased or decreased in response to the volume control command.

18. The non-transitory computer-readable storage medium of claim 13 , wherein the first volume adjustment is an increase in volume of the first stem and the second volume adjustment is a decrease in volume of the second stem.

19. The non-transitory computer-readable storage medium of claim 13 , the first stem is associated with a first speaker identity and the second stem is associated with a second stem identity; and

the machine-learning model is trained further based on the first speaker identity and the second speaker identity.

20. The non-transitory computer-readable storage medium of claim 13 , wherein at least one stem of the plurality of stems corresponds to a background stem.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 17, 2021
From: KHALILIA, MOHAMMED; NAIR, NAVEEN SUDHAKARAN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 057513/0650 →
Cited By (4)
US 12,361,975 US 12,399,678 US 12,437,786 US 12,499,902