IP Library Granted Patent US 12,475,370
Granted Patent B2
US 12,475,370 · App. 17/885,128 · Granted Nov 18, 2025

System and method for automating design of sound source separation deep learning model

Inventors: Joon-Hyuk Chang (Seoul, KR); Joo-Hyun Lee (Seoul, KR)
Assignee: IUCF-HYU (INDUSTRY-UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY)
G06N3/08G10L21/0308
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,370
App. No.
17/885,128
Granted
Nov 18, 2025
Kind
B2
Abstract

Disclosed are a system and method for automating the design of a sound source separation deep learning model. A method of automating a design of a sound source separation deep learning model, which is performed by a design automation system, may include automatically searching for a combination of hyper parameters of a separation model constructed in a sound source separation deep learning model by using a neural architecture search (NAS) algorithm and reconstructing the sound source separation deep learning model based on the retrieved combination of the hyper parameters of the separation model.

Claims (43)

1 . A method of automating a design of a sound source separation deep learning model, which is performed by a design automation system, the method comprising:

automatically searching for a combination of hyper parameters of a separation model constructed in a sound source separation deep learning model by using a neural architecture search (NAS) algorithm; and

reconstructing the sound source separation deep learning model based on the retrieved, combination of the hyper parameters of the separation model,

wherein searching for the combination comprises forming a search space of the separation model based on a kernel size P and channel expansion ratio H of the separation model,

wherein searching for the combination comprises:

constructing one layer by using candidate operations in the formed search space of the separation model and an architecture parameter that assigns a weight to each of the candidate operations;

learning a weight and the architecture parameter of the candidate operation in a model architecture generated by using the constructed one layer, by using a training set and a validation set, respectively; and

obtaining a model architecture by selecting only one candidate operation in which results of the architecture parameter are a highest as the training of the generated model architecture is completed by using the constructed one layer, and

wherein searching for the combination comprises:

extracting a feature vector for each repeat constructed in the separation model through an auxiliary loss operation;

calculating a loss by the number of plurality of repeats constructed in the separation model by calculating a loss for the extracted feature vector; and

calculating a final loss by calculating an average of results of the calculated loss.

2 . The method of claim 1 , wherein:

the sound source separation deep learning model is a fully convolutional time domain voice separation network (Conv-TasNet),

the separation model constructed in the sound source separation deep learning model comprises a plurality of repeats,

each of the plurality of repeats comprises a plurality of TCN blocks, and

the plurality of TCN blocks comprises a layer (Bottleneck) that adjusts the number of channels of an input on which convolution is to be performed by a D-CNN and a layer (depth-wise convolutional network (CNN)) that performs a convolution operation.

3 . The method of claim 1 , wherein searching for the combination comprises performing NAS by adjusting the number of layers for each repeat based on a search space of the formed separation model.

4 . The method of claim 1 , wherein searching for the combination comprises obtaining a model architecture through training using any one parameter, among the weight parameter or the architecture parameter of the candidate operation, by using a binary gate algorithm that forces only one path to be recorded on a GPU memory.

5 . A non-transitory computer-readable recording medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1 .

6 . A design automation system comprising:

a hyper parameter combination search unit configured to automatically search for a combination of hyper parameters of a separation model constructed in a sound source separation deep learning model by using a neural architecture search (NAS) algorithm; and

a model reconstruction unit configured to reconstruct the sound source separation deep learning model based on the retrieved combination of the hyper parameters of the separation model,

wherein the hyper parameter combination search unit forms a search space of the separation model based on a kernel size P and channel expansion ratio H of the separation model,

wherein the hyper parameter combination search unit

constructs one layer by using candidate operations in the formed search space of the separation model and an architecture parameter that assigns a weight to each of the candidate operations,

learns a weight and the architecture parameter of the candidate operation in a model architecture generated by using the constructed one layer, by using a training set and a validation set, respectively, and

obtains a model architecture by selecting only one candidate operation in which results of the architecture parameter are a highest as the training of the generated model architecture is completed by using the constructed one layer,

wherein the hyper parameter combination search unit

constructs one layer by using candidate operations in the formed search space of the separation model and an architecture parameter that assigns a weight to each of the candidate operations,

learns a weight and the architecture parameter of the candidate operation in a model architecture generated by using the constructed one layer, by using a training set and a validation set, respectively, and

obtains a model architecture by selecting only one candidate operation in which results of the architecture parameter are a highest as the training of the generated model architecture is completed by using the constructed one layer, and

wherein the hyper parameter combination search unit

extracts a feature vector for each repeat constructed in the separation model through an auxiliary loss operation,

calculates a loss by the number of plurality of repeats constructed in the separation model by calculating a loss for the extracted feature vector, and

calculates a final loss by calculating an average of results of the calculated loss.

7 . The design automation system of claim 6 , wherein:

the sound source separation deep learning model is a fully convolutional time domain voice separation network (Conv-TasNet),

the separation model constructed in the sound source separation deep learning model comprises a plurality of repeats,

each of the plurality of repeats comprises a plurality of TCN blocks, and

the plurality of TCN blocks comprises a layer (Bottleneck) that adjusts the number of channels of an input on which convolution is to be performed by a D-CNN and a layer (depth-wise convolutional network (CNN)) that performs a convolution operation.

8 . The design automation system of claim 6 , wherein the hyper parameter combination search unit performs NAS by adjusting the number of layers for each repeat based on a search space of the formed separation model.

9 . The design automation system of claim 6 , wherein the hyper parameter combination search unit obtains a model architecture through training using any one parameter, among the weight parameter or the architecture parameter of the candidate operation, by using a binary gate algorithm that forces only one path to be recorded on a GPU memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 10, 2022
From: CHANG, JOON-HYUK; LEE, JOO-HYUN
To: IUCF-HYU (INDUSTRY-UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY)
Reel/Frame 061143/0901 →
Priority Claims (1)
KR 10-2022-0056452 · May 9, 2022 · national
Continuity (1)
Related Publication 20230359885A1 · Nov 9, 2023
References Cited (8)
US 20190066713A1 · Mesgarani et al. · 2019 [cited by applicant]
US 20210173614A1 · Shin · 2021 [cited by examiner]
US 20210224319A1 · Ingel · 2021 [cited by examiner]
US 20240013800A1 · Mesgarani · 2024 [cited by examiner]
CN 112989107A · 2021 [cited by examiner]
KR 20190100117A · 2019 [cited by examiner]
Author: Noda et al., Title: Sound Source Separation for Robot Audition using Deep Learning, Date: Nov. 2015, Publisher: IEEE, pp. 389-394. [cited by examiner]
Liu, et al., “DARTS: Differentiable Architecture Search”, ICLR,arXiv: 1806.09055v2 [cs. LG], Apr. 23, 2019 (13 pages). [cited by applicant]