IP Library Granted Patent US 12,574,681
Granted Patent B2
US 12,574,681 · App. 17/918,619 · Granted Mar 10, 2026

System for real-time recognition and identification of sound sources

Inventors: Laurent Mareuge (Marly le Roi, FR); Julien Roland (Marcq en Baroeuil, FR); Maxime Baelde (Lille, FR)
Assignees: Uby; Wavely
H04R3/04H04R1/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,574,681
App. No.
17/918,619
Granted
Mar 10, 2026
Kind
B2
Abstract

The present invention relates to a method for identifying a sound source comprising the following steps: (S1): acquisition of a sound signal; (S2): application of a frequency fitter to the acquired sound signal in order to obtain a filtered signal; (S4): extraction of a matrix of features associated with the filtered signal; (S5): identification of the source by applying a classification model to the feature matrix extracted in step (S4), the classification model having, as its output, at least one class associated with the source of the acquired sound signal.

Claims (44)

1 . A method implemented by a data processor of a noise monitoring device, comprising the steps of:

S1: obtaining a sound signal emitted at a construction site and acquired by a sound sensor;

S2: applying a frequency filter to the obtained sound signal, thereby obtaining a filtered signal;

S4: extracting features from the filtered signal;

S5: identifying a specific sound source by applying a classification model to the features extracted in step S4 and providing at least one label associated to the specific sound source, wherein the specific sound source is a parent source as defined in a hierarchic model in which each sound source is either a parent source or a child source linked to a parent source; and

a post-processing step, subsequent to step S5, which comprises the following sub-steps:

evaluating a sound level of the obtained sound signal and comparing the evaluated sound level with a first predetermined threshold, said identifying is then considered reliable if the evaluated sound level is greater than the first predetermined threshold;

comparing a value representing a level of confidence associated with the presence of the specific sound source in the obtained sound signal with a second predetermined threshold;

based on the value representing the level of confidence associated with the presence of the specific sound source in the obtained sound signal being greater than the second predetermined threshold:

selecting, from one or more child sources linked to the specific sound source, the child source having a highest level of confidence associated with the presence of the child source in the obtained sound signal;

comparing the level of confidence associated with the presence of the selected child source in the obtained sound signal with a third predetermined threshold;

based on the level of confidence associated with the presence of the selected child source in the obtained sound signal being greater than the third predetermined threshold, the specific sound source is changed to the selected child source; and

based on the level of confidence associated with the presence of the selected child source in the obtained sound signal being lower than the third predetermined threshold, the specific sound source remains the parent source; and

based on the value representing the level of confidence associated with the presence of the specific sound source in the obtained sound signal being lower than the second predetermined threshold, said identifying is considered not reliable; and

communicating an identification signal to a client device based on the identifying being reliable, and not communicating an identification signal based on the identifying being not reliable.

2 . The method of claim 1 , wherein the frequency filter comprises a frequency weighting filter and/or a high pass filter.

3 . The method of claim 1 , wherein extracting the features comprises transforming the filtered signal into a sonogram representing sound energies associated with instants and frequencies.

4 . The method of claim 3 , further comprising converting the frequencies according to a non-linear frequency scale.

5 . The method of claim 4 , wherein the non-linear frequency scale is a Mel scale.

6 . The method of claim 4 , further comprising converting the sound energies according to a logarithmic scale.

7 . The method of claim 1 , further comprising, prior to step S5, normalizing the extracted features according to statistical moments of said extracted features.

8 . The method of claim 1 , wherein the classification model used in step S5 is one of a generative model or a discriminating model.

9 . The method of claim 1 , wherein the at least one label identifying a specific sound source comprises one of the following elements: a single label, a plurality of labels each associated with a probability.

10 . The method of claim 1 , further comprising, prior to step S4, a step S3, of detecting a sound event, the steps S4 and S5 being implemented only when a sound event is detected, the detection of a sound event depending on an indicator of an energy of the obtained sound signal and/or on a reception of a signaling of a sound event.

11 . The method of claim 10 , further comprising a step of notifying a sound event when a sound event is detected and/or when a signaling is received.

12 . A system for monitoring noise at a construction site, comprising:

a sound sensor configured to acquire a sound signal; and

a data processor configured to perform the steps of:

S1: obtaining a sound signal emitted at a construction site and acquired by a sound sensor;

S2: applying a frequency filter to the obtained sound signal, thereby obtaining a filtered signal;

S4: extracting features from the filtered signal;

S5: identifying a specific sound source by applying a classification model to the features extracted in step S4 and providing at least one label associated to the a specific sound source, wherein the specific sound source is a parent source as defined in a hierarchic model in which each sound source is either a parent source or a child source linked to a parent source; and

a post-processing step, subsequent to step S5, which comprises the following sub-steps:

evaluating a sound level of the obtained sound signal and comparing the evaluated sound level with a first predetermined threshold, said evaluating is then considered reliable if the evaluated sound level is greater than the first predetermined threshold;

comparing a value representing a level of confidence associated with the presence of the specific sound source in the obtained sound signal with a second predetermined threshold;

based on the value representing the level of confidence associated with the specific sound source being greater than the second predetermined threshold:

selecting, from one or more child sources linked to the specific sound source, the child source having a highest level of confidence associated with the presence of the child source in the obtained sound signal;

comparing the level of confidence associated with the presence of the selected child source in the obtained sound signal with a third predetermined threshold;

based on the level of confidence associated with the presence of the selected child source in the obtained sound signal being greater than the third predetermined threshold, the specific sound source is changed to the selected child source; and

based on the level of confidence associated with the presence of the selected child source in the obtained sound signal being lower than the third predetermined threshold, the specific sound source remains the parent source; and

based on the value representing the level of confidence associated with the specific sound source being lower than the second predetermined threshold, said identifying is considered not reliable;

communicating an identification signal to a client device based on the identifying is being reliable, and not communicating an identification signal based on the identifying being not reliable.

13 . The system of claim 12 , further comprising a detector configured to detect a sound event based on an indicator of an energy of the obtained sound signal and/or on the reception of a signaling of a sound event.

14 . The system of claim 13 , further comprising a mobile terminal configured to signal the sound event.

Assignments (2)
CHANGE OF NAME Recorded Sep 13, 2023
From: COM'IN SAS
To: UBY
Reel/Frame 064896/0540 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2022
From: MAREUGE, LAURENT; ROLAND, JULIEN; BAELDE, MAXIME
To: COM'IN SAS; WAVELY
Reel/Frame 062141/0858 →
Priority Claims (1)
FR FR2003842 · Apr 16, 2020 · national
Continuity (1)
Related Publication 20230143027A1 · May 11, 2023
References Cited (20)
US 3770892A · Clapper · 1973 [cited by examiner]
US 6012334A · Ando · 2000 [cited by examiner]
US 6058205A · Bahl · 2000 [cited by examiner]
US 6243671B1 · Lago · 2001 [cited by examiner]
US 6243695B1 · Assaleh · 2001 [cited by examiner]
US 20080243505A1 · Barinov · 2008 [cited by examiner]
US 20100142725A1 · Goldstein · 2010 [cited by examiner]
US 20170315516A1 · Kozionov · 2017 [cited by examiner]
US 20170372242A1 · Alsubai et al. · 2017 [cited by applicant]
US 20200066257A1 · Smith et al. · 2020 [cited by applicant]
US 20200191643A1 · Davis · 2020 [cited by examiner]
CN 106055573A · 2016 [cited by examiner]
CN 107545033A · 2018 [cited by examiner]
EP 1092964A2 · 2001 [cited by applicant]
WO WO2016198877A1 · 2016 [cited by applicant]
French preliminary search report issued for FR 2003842. [cited by applicant]
International search report issued for PCT FR2021/050674. [cited by applicant]
Pan, Z., et al., “Cognitive Acoustic Analytics Service for Internet of Things,” 2017 IEEE 1st International Conference on Cognitive Computing, pp. 96-103 (2017). [cited by applicant]
Maijala, P., et al., “Environmental noise monitoring using source classification in sensors,” Applied Acoustics, 129, pp. 258-267 (2018). [cited by applicant]
Cai, Y., et al., “Sound Recognition,” Computing with Instinct, pp. 16-34 (2011). [cited by applicant]