IP Library Granted Patent US 12,061,840
Granted Patent B2
US 12,061,840 · App. 18/453,792 · Granted Aug 13, 2024

Methods and apparatus for dynamic volume adjustment via audio classification

Inventors: Markus Cremer (Orinda, CA); Robert Coover (Orinda, CA); Steven D. Scherf (Oakland, CA); Cameron Aubrey Summers (Oakland, CA)
Assignee: Gracenote, Inc.
G06F3/165G10L25/51G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,061,840
App. No.
18/453,792
Granted
Aug 13, 2024
Kind
B2
Abstract

Methods, apparatus, systems and articles of manufacture are disclosed for dynamic volume adjustment via audio classification. Example apparatus include at least one memory; instructions; and at least one processor to execute the instructions to: analyze, with a neural network, a parameter of an audio signal associated with a first volume level to determine a classification group associated with the audio signal; determine an input volume of the audio signal; determine a classification gain value based on the classification group; determine an intermediate gain value as an intermediate between the input volume and the classification gain value by applying a first weight to the input volume and a second weight to the classification gain value; apply the intermediate gain value to the audio signal, the intermediate gain value to modify the first volume level to a second volume level; and apply a compression value to the audio signal, the compression value to modify the second volume level to a third volume level that satisfies a target volume threshold.

Claims (37)

1. A non-transitory machine-readable medium having stored thereon instructions that, when executed, cause one or more processors to perform a set of operations comprising:

analyzing, by a neural network, a parameter of an input audio signal associated with a first input volume level to determine a classification group associated with the input audio signal;

determining an input volume of the input audio signal;

determining a classification gain value based on the classification group;

determining a target gain value as an intermediary value between the input volume and the classification gain value, wherein the target gain value is determined by applying one or more weights to the input volume and the classification gain value; and

generating a gain-adjusted volume by applying the target gain value to the input audio signal.

2. The non-transitory machine-readable medium of claim 1 , wherein the target gain value is determined by applying a first weight to the input volume and a second weight to the classification gain value.

3. The non-transitory machine-readable medium of claim 1 , wherein the set of operations further comprises applying the intermediary value to the input audio signal, wherein the intermediary value modifies the first input volume level to a second volume level.

4. The non-transitory machine-readable medium of claim 3 , wherein the set of operations further comprises applying a compression value to the input audio signal, wherein the compression value modifies the second volume level to a third volume level that satisfies a target volume threshold.

5. The non-transitory machine-readable medium of claim 1 , wherein the set of operations further comprises determining if a source of the input audio signal has changed.

6. The non-transitory machine-readable medium of claim 5 , wherein determining if the source of the input audio signal has changed is based on at least one of: (1) a comparison of a current compressor gain associated with the input audio signal to a previous compressor gain associated with the input audio signal, (2) a comparison of a RMS power associated with the input audio signal to a previous RMS power associated with the input audio signal, or (3) a comparison of a current audio sample value associated with the input audio signal to a previous audio sample value associated with the input audio signal.

7. The non-transitory machine-readable medium of claim 5 , wherein the set of operations further comprises, in response to determining the source of the input audio signal has changed, resetting the intermediary value of the input audio signal.

8. The non-transitory machine-readable medium of claim 1 , wherein the classification group comprises at least one of: (1) a genre of music represented by the input audio signal, (2) a time period of music represented by the input audio signal, or (3) a presence of an instrument in music represented by the input audio signal.

9. A method, comprising:

analyzing, by a neural network, a parameter of an input audio signal associated with a first input volume level to determine a classification group associated with the input audio signal;

determining an input volume of the input audio signal;

determining a classification gain value based on the classification group;

determining a target gain value as an intermediary value between the input volume and the classification gain value, wherein the target gain value is determined by applying one or more weights to the input volume and the classification gain value; and

generating a gain-adjusted volume by applying the target gain value to the input audio signal.

10. The method of claim 9 , wherein the target gain value is determined by applying a first weight to the input volume and a second weight to the classification gain value.

11. The method of claim 9 , wherein the method further comprises applying the intermediary value to the input audio signal, wherein the intermediary value modifies the first input volume level to a second volume level.

12. The method of claim 11 , wherein the method further comprises applying a compression value to the input audio signal, wherein the compression value modifies the second volume level to a third volume level that satisfies a target volume threshold.

13. The method of claim 9 , wherein the method further comprises determining if a source of the input audio signal has changed.

14. The method of claim 13 , wherein determining if the source of the input audio signal has changed is based on at least one of: (1) a comparison of a current compressor gain associated with the input audio signal to a previous compressor gain associated with the input audio signal, (2) a comparison of a RMS power associated with the input audio signal to a previous RMS power associated with the input audio signal, or (3) a comparison of a current audio sample value associated with the input audio signal to a previous audio sample value associated with the input audio signal.

15. The method of claim 13 , wherein the method further comprises, in response to determining the source of the input audio signal has changed, resetting the intermediary value of the input audio signal.

16. The method of claim 9 , wherein the classification group comprises at least one of: (1) a genre of music represented by the input audio signal, (2) a time period of music represented by the input audio signal, or (3) a presence of an instrument in music represented by the input audio signal.

17. An apparatus, comprising:

one or more processors; and

a non-transitory machine-readable medium having stored thereon instructions that, when executed, cause one or more processors to perform a set of operations comprising:

analyzing, by a neural network, a parameter of an input audio signal associated with a first input volume level to determine a classification group associated with the input audio signal;

determining an input volume of the input audio signal;

determining a classification gain value based on the classification group;

determining a target gain value as an intermediary value between the input volume and the classification gain value, wherein the target gain value is determined by applying one or more weights to the input volume and the classification gain value; and

generating a gain-adjusted volume by applying the target gain value to the input audio signal.

18. The apparatus of claim 17 , wherein the target gain value is determined by applying a first weight to the input volume and a second weight to the classification gain value.

19. The apparatus of claim 17 , wherein the set of operations further comprises applying the intermediary value to the input audio signal, wherein the intermediary value modifies the first input volume level to a second volume level.

20. The apparatus of claim 19 , wherein the set of operations further comprises applying a compression value to the input audio signal, wherein the compression value modifies the second volume level to a third volume level that satisfies a target volume threshold.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE APPLICATION NUMBER 18/453732 PREVIOUSLY RECORDED ON REEL 67645 FRAME 357. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Sep 30, 2024
From: CREMER, MARKUS; SCHERF, STEVEN D.; SUMMERS, CAMERON AUBREY; COOVER, ROBERT
To: GRACENOTE, INC.
Reel/Frame 068965/0592 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2024
From: CREMER, MARKUS; SCHERF, STEVEN D.; SUMMERS, CAMERON AUBREY; COOVER, ROBERT
To: GRACENOTE, INC.
Reel/Frame 067645/0357 →
Continuity (5)
Continuation 17380936 · Jul 20, 2021
Continuation 16563717 · Sep 6, 2019
Provisional Application 62745148 · Oct 12, 2018
Provisional Application 62728677 · Sep 7, 2018
Related Publication 20240045649A1 · Feb 8, 2024