IP Library › Granted Patent US 12,563,339
Granted Patent B2
US 12,563,339 · App. 18/398,821 · Granted Feb 24, 2026

Signal normalization using loudness metadata for audio processing

Inventor: Sunil Bharitkar (Stevenson Ranch, CA)
Assignee: Samsung Electronics Co., Ltd.
H04R3/00G06N3/0464H04R2430/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,563,339
App. No.
18/398,821
Granted
Feb 24, 2026
Kind
B2
Abstract

One embodiment provides a method of signal normalization. The method comprises receiving an input content with a corresponding audio signal, and extracting loudness metadata from an audio signal corresponding to the input content. The method further comprises estimating, using a machine learning model, a peak-level amplitude based on the loudness metadata. The peak-level amplitude represents a maximum linear amplitude of the audio signal over an entire duration of the input content. The method further comprises determining a gain based at least on the peak-level amplitude, and applying the gain to the audio signal. The resulting gain-scaled audio signal is provided to one or more speakers coupled to or integrated in an electronic device for audio playback.

Claims (46)

1 . A method of signal normalization, comprising:

receiving an input content with a corresponding audio signal;

extracting loudness metadata from an audio signal corresponding to the input content;

estimating, using a machine learning model, a peak-level amplitude based on the loudness metadata, wherein the peak-level amplitude represents a maximum linear amplitude of the audio signal over an entire duration of the input content;

determining a gain based at least on the peak-level amplitude; and

applying the gain to the audio signal, wherein the resulting gain-scaled audio signal is provided to one or more speakers coupled to or integrated in an electronic device for audio playback.

2 . The method of claim 1 , further comprising:

obtaining one or more desired user settings for the electronic device; and

obtaining one or more device settings representing a preset configuration of the electronic device.

3 . The method of claim 2 , wherein the gain is further based on at least one of the one or more desired user settings and the one or more device settings.

4 . The method of claim 1 , wherein the machine learning model is one of a least-squares optimal model, a linear regression model, a nonlinear regression model, or a fully-connected feed-forward neural network.

5 . The method of claim 1 , further comprising:

applying one or more audio post-processes to the gain-scaled audio signal before the audio playback.

6 . The method of claim 1 , wherein the one or more audio post-processes include at least one of perceptual bass enhancement, equalization, audio upmixing, spatial rendering, and digital-to-analog conversion.

7 . The method of claim 1 , wherein the peak-level amplitude is estimated without analyzing the entire duration of the input content.

8 . The method of claim 1 , wherein the loudness metadata is extracted from a header of the audio signal.

9 . A system of signal normalization, comprising:

at least one processor; and

a non-transitory processor-readable memory device storing instructions that when executed by the at least one processor causes the at least one processor to perform operations including:

receiving an input content with a corresponding audio signal;

extracting loudness metadata from an audio signal corresponding to the input content;

estimating, using a machine learning model, a peak-level amplitude based on the loudness metadata, wherein the peak-level amplitude represents a maximum linear amplitude of the audio signal over an entire duration of the input content;

determining a gain based at least on the peak-level amplitude; and

applying the gain to the audio signal, wherein the resulting gain-scaled audio signal is provided to one or more speakers coupled to or integrated in an electronic device for audio playback.

10 . The system of claim 9 , wherein the operations further include:

obtaining one or more desired user settings for the electronic device; and

obtaining one or more device settings representing a preset configuration of the electronic device.

11 . The system of claim 10 , wherein the gain is further based on at least one of the one or more desired user settings and the one or more device settings.

12 . The system of claim 9 , wherein the machine learning model is one of a least-squares optimal model, a linear regression model, a nonlinear regression model, or a fully-connected feed-forward neural network.

13 . The system of claim 9 , wherein the operations further include:

applying one or more audio post-processes to the gain-scaled audio signal before the audio playback.

14 . The system of claim 9 , wherein the one or more audio post-processes include at least one of perceptual bass enhancement, equalization, audio upmixing, spatial rendering, and digital-to-analog conversion.

15 . The system of claim 9 , wherein the peak-level amplitude is estimated without analyzing the entire duration of the input content.

16 . A non-transitory processor-readable medium that includes a program that when executed by a processor performs a method of signal normalization, the method comprising:

receiving an input content with a corresponding audio signal;

extracting loudness metadata from an audio signal corresponding to the input content;

estimating, using a machine learning model, a peak-level amplitude based on the loudness metadata, wherein the peak-level amplitude represents a maximum linear amplitude of the audio signal over an entire duration of the input content;

determining a gain based at least on the peak-level amplitude; and

applying the gain to the audio signal, wherein the resulting gain-scaled audio signal is provided to one or more speakers coupled to or integrated in an electronic device for audio playback.

17 . The non-transitory processor-readable medium of claim 16 , wherein the method further comprises:

obtaining one or more desired user settings for the electronic device; and

obtaining one or more device settings representing a preset configuration of the electronic device.

18 . The non-transitory processor-readable medium of claim 17 , wherein the gain is further based on at least one of the one or more desired user settings and the one or more device settings.

19 . The non-transitory processor-readable medium of claim 16 , wherein the machine learning model is one of a least-squares optimal model, a linear regression model, a nonlinear regression model, or a fully-connected feed-forward neural network.

20 . The non-transitory processor-readable medium of claim 16 , wherein the method further comprises:

applying one or more audio post-processes to the gain-scaled audio signal before the audio playback.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2023
From: BHARITKAR, SUNIL
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 065972/0436 →
Continuity (2)
Provisional Application 63444324 · Feb 9, 2023
Related Publication 20240276143A1 · Aug 15, 2024
References Cited (25)
US 8457321B2 · Daubigny · 2013 [cited by examiner]
US 9934790B2 · Baumgarte · 2018 [cited by applicant]
US 10341770B2 · Baumgarte · 2019 [cited by applicant]
US 10958229B2 · Baumgarte et al. · 2021 [cited by applicant]
US 11232807B2 · Hines · 2022 [cited by examiner]
US 11354088B2 · Alexander · 2022 [cited by examiner]
US 11533575B2 · Ward et al. · 2022 [cited by applicant]
US 11950069B2 · Ali · 2024 [cited by examiner]
US 20100250258A1 · Smithers et al. · 2010 [cited by applicant]
US 20150332685A1 · Bleidt · 2015 [cited by applicant]
US 20160323684A1 · Carroll et al. · 2016 [cited by applicant]
US 20190325886A1 · Riedmiller et al. · 2019 [cited by applicant]
US 20200314578A1 · Filos · 2020 [cited by examiner]
US 20210265966A1 · Port · 2021 [cited by examiner]
US 20210349683A1 · Cremer et al. · 2021 [cited by applicant]
US 20210367574A1 · Chon et al. · 2021 [cited by applicant]
US 20220129237A1 · Chon et al. · 2022 [cited by applicant]
US 20220165289A1 · Stone et al. · 2022 [cited by applicant]
US 20220166397A1 · Chung et al. · 2022 [cited by applicant]
US 20220264160A1 · Jo et al. · 2022 [cited by applicant]
US 20230023024A1 · Riedmiller et al. · 2023 [cited by applicant]
KR 1020220071954A · 2022 [cited by applicant]
WO 2022086196A1 · 2022 [cited by applicant]
International Search Report and Written Opinion dated Apr. 23, 2024 for International Application PCT/KR2024/001013, from Korean Intellectual Property Office, pp. 1-6, Republic of Korea. [cited by applicant]
Extended European search report dated Nov. 4, 2025 for European Application 24753499.3, from European Patent Office, pp. 1-11, Munich, Germany. [cited by applicant]