IP Library › Granted Patent US 11,714,596
Granted Patent B2
US 11,714,596 · App. 17/506,678 · Granted Aug 1, 2023

Audio signal processing method and apparatus

Inventors: Sangbae Chon (Seoul, KR); Soochul Park (Seoul, KR)
Assignee: GAUDIO LAB, INC.
G06F3/165H04R3/00H04S7/30H04R2430/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,714,596
App. No.
17/506,678
Granted
Aug 1, 2023
Kind
B2
Abstract

Disclosed is an operation method of an audio signal processing device configured to process an audio signal including a first audio signal component and a second audio signal component. The operation method includes: receiving the audio signal; normalizing loudness of the audio signal, based on a pre-designated target loudness; acquiring the first audio signal component from the audio signal having the normalized loudness, by using a machine learning model; and de-normalizing loudness of the first audio signal component, based on the pre-designated target loudness.

Claims (34)

1. An operation method of an audio signal processing device which operates in at least one process and is configured to process an audio signal comprising a first audio signal component and a second audio signal component, the method comprising:

receiving the audio signal;

normalizing loudness of the audio signal, based on a pre-designated target loudness;

acquiring the first audio signal component from the audio signal having the normalized loudness, by using a machine learning model; and

de-normalizing loudness of the first audio signal component, based on the pre-designated target loudness.

2. The method of claim 1 , wherein at least one of the first audio signal component and the second audio signal component is an audio signal component corresponding to a voice.

3. The method of claim 1 , wherein the normalizing of the loudness of the audio signal, based on the pre-designated target loudness, comprises normalizing loudness in units of contents included in the audio signal.

4. The method of claim 1 , wherein the machine learning model processes the audio signal having the normalized loudness in a frequency area.

5. The method of claim 1 , wherein the normalizing of the loudness of the audio signal, based on the pre-designated target loudness, comprises:

dividing the audio signal into multiple pre-designated time intervals, dividing loudness values in the multiple pre-designated time intervals into multiple levels, and acquiring loudness of the audio signal by using a loudness value distribution for each of the multiple levels; and

normalizing the loudness of the audio signal to target loudness.

6. The method of claim 1 , wherein the machine learning model comprises gate logic.

7. The method of claim 1 , wherein the acquiring of the first audio signal component from the audio signal having the normalized loudness, by using the machine learning model, comprises classifying a frequency bin-specific score acquired from the machine learning model, based on a pre-designated threshold value,

wherein the score indicates a degree of closeness to the first audio signal component.

8. A method for training a machine learning model which operates in at least one process and is configured to classify a first audio signal component from an audio signal comprising the first audio signal component and a second audio signal acquired from different sources, the method comprising:

receiving the audio signal;

normalizing loudness of the audio signal, based on pre-designated target loudness;

acquiring a first audio signal component from the audio signal having the normalized loudness, by using the machine learning model; and

restoring the loudness of the first audio signal component, based on the pre-designated target loudness.

9. The method of claim 8 , wherein at least one of the first audio signal component and the second audio signal component is an audio signal component corresponding to a voice.

10. The method of claim 8 , wherein the normalizing of the loudness of the audio signal, based on the pre-designated target loudness, comprises normalizing loudness in units of contents included in the audio signal.

11. The method of claim 8 , wherein the machine learning model processes the audio signal having the normalized loudness in a frequency area.

12. The method of claim 8 , wherein the normalizing of the loudness of the audio signal, based on the pre-designated target loudness, comprises:

dividing the audio signal into multiple pre-designated time intervals, dividing loudness values in the multiple pre-designated time intervals into multiple levels, and acquiring loudness of the audio signal by using a loudness value distribution for each of the multiple levels; and

normalizing the loudness of the audio signal to target loudness.

13. The method of claim 8 , wherein the machine learning model comprises gate logic.

14. The method of claim 8 , wherein the acquiring of the first audio signal component from the audio signal having the normalized loudness, by using the machine learning model, comprises classifying a frequency bin-specific score acquired from the machine learning model, based on a pre-designated threshold value,

wherein the score indicates a degree of closeness to the first audio signal component.

15. An audio signal processing device configured to process an audio signal comprising a first audio signal component and a second audio signal component, the device comprising at least one processor,

wherein the at least one processor:

receives the audio signal;

normalizes loudness of the audio signal, based on a pre-designated target loudness;

acquires the first audio signal component from the audio signal having the normalized loudness, by using a machine learning model; and

de-normalizes loudness of the first audio signal component, based on the pre-designated target loudness.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2022
From: PARK, SOOCHUL; CHON, SANGBAE
To: GAUDIO LAB, INC.
Reel/Frame 059259/0523 →
Priority Claims (1)
KR 10-2020-0137269 · Oct 22, 2020 · national
Continuity (2)
Provisional Application 63118979 · Nov 30, 2020
Related Publication 20220129237A1 · Apr 28, 2022