IP Library Granted Patent US 11,545,162
Granted Patent B2
US 11,545,162 · App. 16/652,759 · Granted Jan 3, 2023

Audio reconstruction method and device which use machine learning

Inventors: Ho-sang Sung (Seoul, KR); Jong-hoon Jeong (Hwaseong-si, KR); Ki-hyun Choo (Seoul, KR); Eun-mi Oh (Seoul, KR); Jong-youb Ryu (Hwaseong-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G10L19/02G06F3/162G06K9/6227G06N3/08G06N20/00G10L19/0017
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,545,162
App. No.
16/652,759
Granted
Jan 3, 2023
Kind
B2
Abstract

Provided are an audio reconstruction method and device for providing improved sound quality by reconstructing a decoding parameter or an audio signal obtained from a bitstream, by using machine learning. The audio reconstruction method includes obtaining a plurality of decoding parameters of a current frame by decoding a bitstream, determining characteristics of a second parameter included in the plurality of decoding parameters and associated with a first parameter, based on the first parameter included in the plurality of decoding parameters, obtaining a reconstructed second parameter by applying a machine learning model to at least one of the plurality of decoding parameters, the second parameter, and the characteristics of the second parameter, and decoding an audio signal, based on the reconstructed second parameter.

Claims (46)

1. An audio reconstruction method comprising:

obtaining a plurality of decoding parameters of a current frame by decoding a bitstream;

determining characteristics of a second parameter comprised in the plurality of decoding parameters and associated with a first parameter, based on the first parameter comprised in the plurality of decoding parameters;

obtaining a reconstructed second parameter by applying a machine learning model to at least one of the plurality of decoding parameters, the second parameter, and the characteristics of the second parameter;

obtaining a corrected second parameter by correcting the reconstructed second parameter, based on the characteristics of the second parameter; and

decoding an audio signal, based on the corrected second parameter.

2. The audio reconstruction method of claim 1 ,

wherein the determining of the characteristics of the second parameter comprises determining a range of the second parameter based on the first parameter, and

wherein the obtaining of the corrected second parameter comprises, in response to the reconstructed second parameter not being within the range of the second parameter, obtaining a range value, which is closest to the reconstructed second parameter, as the corrected second parameter.

3. The audio reconstruction method of claim 1 , wherein the determining of the characteristics of the second parameter comprises determining the characteristics of the second parameter by using a pre-trained machine learning model pre-trained based on at least one of the first or second parameters.

4. The audio reconstruction method of claim 1 , wherein the obtaining of the reconstructed second parameter comprises:

determining candidates of the second parameter based on the characteristics of the second parameter; and

selecting one of the candidates of the second parameter based on the machine learning model.

5. The audio reconstruction method of claim 1 , wherein the obtaining of the reconstructed second parameter comprises obtaining the reconstructed second parameter of the current frame based on at least one of a plurality of decoding parameters of a previous frame.

6. The audio reconstruction method of claim 1 , wherein the machine learning model is generated by machine-learning an original audio signal and at least one of the plurality of decoding parameters.

7. The audio reconstruction method of claim 1 , further comprising:

based on the decoded audio signal and at least one of the plurality of decoding parameters, selecting the machine learning model from among a plurality of machine learning models; and

reconstructing the decoded audio signal by using the selected machine learning model.

8. The audio reconstruction method of claim 7 , wherein the selecting of the machine learning model comprises:

determining a start frequency of a bandwidth extension, based on at least one of the plurality of decoding parameters; and

selecting a machine learning model of the decoded audio signal, based on the start frequency and a frequency of the decoded audio signal.

9. The audio reconstruction method of claim 7 , wherein the selecting of the machine learning model comprises:

obtaining a gain of a current frame, based on at least one of the plurality of decoding parameters;

obtaining an average of gains of the current frame and frames adjacent to the current frame;

selecting a machine learning model for a transient signal when a difference between the gain of the current frame and the average of the gains is greater than a threshold;

determining whether a window type comprised in the plurality of decoding parameters indicates short, when the difference between the gain of the current frame and the average of the gains is less than the threshold;

selecting the machine learning model for the transient signal when the window type indicates short; and

selecting a machine learning model for a stationary signal when the window type does not indicate short.

10. A non-transitory computer-readable recording medium having recorded thereon a computer program for executing the method of claim 1 .

11. An audio reconstruction device comprising:

a memory storing a received bitstream; and

at least one processor configured to:

obtain a plurality of decoding parameters of a current frame by decoding the bitstream,

determine characteristics of a second parameter comprised in the plurality of decoding parameters and associated with a first parameter,

based on the first parameter comprised in the plurality of decoding parameters, obtain a reconstructed second parameter by applying a machine learning model to at least one of the plurality of decoding parameters, the second parameter, and the characteristics of the second parameter,

obtain a corrected second parameter by correcting the reconstructed second parameter, based on the characteristics of the second parameter, and

decode an audio signal, based on the corrected second parameter.

12. The audio reconstruction device of claim 11 , wherein the at least one processor is further configured to determine the characteristics of the second parameter by using a pre-trained machine learning model pre-trained based on at least one of the first or second parameters.

13. The audio reconstruction device of claim 11 , wherein the at least one processor is further configured to:

obtain the reconstructed second parameter by determining candidates of the second parameter based on the characteristics of the second parameter, and

select one of the candidates of the second parameter based on the machine learning model.

14. The audio reconstruction device of claim 11 , wherein the at least one processor is further configured to obtain the reconstructed second parameter of the current frame based on at least one of a plurality of decoding parameters of a previous frame.

15. The audio reconstruction device of claim 11 , wherein the machine learning model is generated by machine-learning an original audio signal and at least one of the plurality of decoding parameters.

16. The audio reconstruction device of claim 11 , wherein the at least one processor is further configured to:

based on the decoded audio signal and at least one of the plurality of decoding parameters, select a first machine learning model from among a plurality of machine learning models, and

reconstruct the decoded audio signal by using the first machine learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2020
From: SUNG, HO-SANG; JEONG, JONG-HOON; CHOO, KI-HYUN; OH, EUN-MI; RYU, JONG-YOUB
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 052282/0856 →
Continuity (1)
Related Publication 20200234720A1 · Jul 23, 2020