IP Library › Granted Patent US 12,300,266
Granted Patent B2
US 12,300,266 · App. 17/960,563 · Granted May 13, 2025

Computing device for attention-based joint training with noise suppression model for sound event detection technology robust against noise environment and method of thereof

Inventors: Joon-Hyuk Chang (Seoul, KR); Jin Young Son (Seoul, KR)
Assignee: IUCF-HYU (INDUSTRY-UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY)
G10L25/51G10L15/20G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,300,266
App. No.
17/960,563
Granted
May 13, 2025
Kind
B2
Abstract

Various embodiments relate to a computing device for attention-based joint training with a noise suppression model for a sound event detection (SED) technology that is robust against a noise environment and a method thereof, and are intended to improve SED performance that is robust against a noise environment by using a joint model in which a noise suppression model and an SED model have been jointed. According to various embodiments, an attention scheme and a weight freezing scheme are added to the joint model.

Claims (26)

1. A computing device for a sound event detection (SED) technology robust against a noise environment, the computing device comprising:

a memory;

an input module configured to obtain a sound event input; and

a processor connected to the memory and the input module and configured to execute at least one instruction stored in the memory,

wherein the processor is configured to obtain a joint model for performing noise suppression and SED on the sound event input and identify a type of the sound event input by using the joint model, and

a noise suppression model and an SED model are jointed as the joint model,

wherein the joint model is a model generated by further being trained after the noise suppression model and the SED model are joined,

wherein the SED model in the joint model comprises a plurality of feature extraction layers to which an output of the noise suppression model in the joint model is input, and a classification layer disposed at an end of the plurality of feature extraction layers, and

wherein weights of the classification layer are fixed and not updated during the training after the noise suppression model and the SED model are joined.

2. The computing device of claim 1 , wherein each of the noise suppression model and the SED model has been pre-trained before the noise suppression model and the SED model are joined.

3. The computing device of claim 1 , wherein the joint model comprises a plurality of attention modules connected to the plurality of feature extraction layers, respectively, and configured to incorporate the output of the noise suppression model into the plurality of feature extraction layers, respectively.

4. The computing device of claim 1 , wherein the noise suppression model is implemented based on a deep neural network (DNN),

wherein the SED model is implemented based on a convolutional recurrent neural network (CRNN),

wherein the plurality of feature extraction layers comprise convolution layers, respectively, and

wherein the classification layer comprises a fully-connected layer.

5. A method of a computing device for a sound event detection (SED) technology robust against a noise environment, the method comprising:

obtaining a sound event input;

obtaining a joint model for performing noise suppression and SED on the sound event input; and

identifying a type of the sound event input by using the joint model,

wherein a noise suppression model and an SED model are jointed as the joint model,

wherein the joint model is a model generated by further being trained after the noise suppression model and the SED model are joined,

wherein the SED model in the joint model comprises a plurality of feature extraction layers to which an output of the noise suppression model in the joint model is input, and a classification layer disposed at an end of the plurality of feature extraction layers, and

wherein weights of the classification layer are fixed and not updated during the training after the noise suppression model and the SED model are joined.

6. The method of claim 5 , wherein each of the noise suppression model and the SED model has been pre-trained before the noise suppression model and the SED model are joined.

7. The method of claim 5 ,

wherein the joint model comprises a plurality of attention modules connected to the plurality of feature extraction layers, respectively, and configured to incorporate the output of the noise suppression model into the plurality of feature extraction layers, respectively.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 5, 2022
From: CHANG, JOON-HYUK; SON, JIN YOUNG
To: IUCF-HYU (INDUSTRY-UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY)
Reel/Frame 061322/0952 →
Priority Claims (1)
KR 10-2021-0177369 · Dec 13, 2021 · national
Continuity (1)
Related Publication 20230186940A1 · Jun 15, 2023
References Cited (15)
US 11074927B2 · Liang · 2021 [cited by examiner]
US 11308939B1 · Gao · 2022 [cited by examiner]
US 11830519B2 · Gevrekci · 2023 [cited by examiner]
US 20180053087A1 · Fukuda · 2018 [cited by examiner]
US 20180261225A1 · Watanabe · 2018 [cited by examiner]
US 20220084509A1 · Sivaraman · 2022 [cited by examiner]
US 20220383887A1 · Wang · 2022 [cited by examiner]
US 20220406295A1 · Weninger · 2022 [cited by examiner]
US 20230032385A1 · Zhang · 2023 [cited by examiner]
US 20230186939A1 · Tang · 2023 [cited by examiner]
JP 2016180839A · 2016 [cited by applicant]
KR 1020710030923A · 2017 [cited by applicant]
KR 1020210098083A · 2021 [cited by applicant]
Son, Jin-Young, and Joon-Hyuk Chang. “Attention-based joint training of noise suppression and sound event detection for noise-robust classification.” Sensors 21.20 (2021): 6718. (Year: 2021). [cited by examiner]
Korean Office Action issued Sep. 23, 2024 by the Korean Patent Office corresponding to Korean patent application No. 10-2021-0177369. [cited by applicant]