IP Library › Granted Patent US 11,416,742
Granted Patent B2
US 11,416,742 · App. 16/122,708 · Granted Aug 16, 2022

Audio signal encoding method and apparatus and audio signal decoding method and apparatus using psychoacoustic-based weighted error function

Inventors: Jongmo Sung (Daejeon, KR); Minje Kim (Bloomington, IN); Aswin Sivaraman (Bloomington, IN); Kai Zhen (Bloomington, IN)
Assignees: Electronics and Telecommunications Research Institute; THE TRUSTEES OF INDIANA UNIVERSITY
G06N3/08G10L19/008G10L19/032G10L25/30G10L25/69
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,416,742
App. No.
16/122,708
Granted
Aug 16, 2022
Kind
B2
Abstract

Provided is a training method of a neural network that is applied to an audio signal encoding method using an audio signal encoding apparatus, the training method including generating a masking threshold of a first audio signal before training is performed, calculating a weight matrix to be applied to a frequency component of the first audio signal based on the masking threshold, generating a weighted error function obtained by correcting a preset error function using the weight matrix, and generating a second audio signal by applying a parameter learned using the weighted error function to the first audio signal.

Claims (13)

1. A training method of a neural network, that is applied to an audio signal encoding method using an audio signal encoding apparatus, the training method comprising:

generating a masking threshold of a first audio signal before training is performed;

calculating a weight matrix to be applied to a frequency component of the first audio signal based on the masking threshold;

generating a weighted error function obtained by correcting a preset error function using the weight matrix; and

generating a second audio signal by applying a parameter learned using the weighted error function to the first audio signal.

2. The training method of claim 1 , wherein the weight matrix includes a weight to be applied to the frequency component of the first audio signal, and

the weight is set to be inversely proportional to the masking threshold of the first audio signal and proportional to a magnitude of the frequency component of the first audio signal.

3. The training method of claim 1 , further comprising:

comparing the second audio signal to the first audio signal and performing a perceptual quality evaluation.

4. The training method of claim 3 , wherein the perceptual quality evaluation includes an objective evaluation based on a perceptual evaluation of speech quality (PESQ), a perceptual objective listening quality assessment (POLQA), or a perceptual evaluation of audio quality (PEAQ).

5. The training method of claim 3 , wherein the perceptual quality evaluation includes a subjective evaluation based on a mean opinion score (MOS) or multiple stimuli with hidden reference and anchor (MUSHRA).

6. The training method of claim 3 , wherein the neural network determines whether a topology, a structural complexity, or the level of quantization to represent the model parameters, included in a model is adjustable based on the perceptual quality evaluation.

7. The training method of claim 6 , wherein when determining whether the topology is adjustable, the neural network re-learns a parameter using a complexity-increased model in response to a result of the perceptual quality evaluation not satisfying a preset quality requirement and, in response to a result of the perceptual quality evaluation satisfying a preset quality requirement, re-learns a parameter using a complexity-reduced model in the quality requirement.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2019
From: KIM, MINJE; ZHEN, KAI; SIVARAMAN, ASWIN
To: THE TRUSTEES OF INDIANA UNIVERSITY
Reel/Frame 048379/0308 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2018
From: SUNG, JONGMO; KIM, MINJE; SIVARAMAN, ASWIN; ZHEN, KAI
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE; THE TRUSTEES OF INDIANA UNIVERSITY
Reel/Frame 046795/0191 →
Priority Claims (1)
KR 10-2017-0173405 · Dec 15, 2017 · national
Continuity (2)
Provisional Application 62590488 · Nov 24, 2017
Related Publication 20190164052A1 · May 30, 2019
Cited By (1)
US 12,333,440