IP Library Granted Patent US 12,504,402
Granted Patent B2
US 12,504,402 · App. 18/372,300 · Granted Dec 23, 2025

Method and system for an acoustic based anomaly detection in industrial machines

Inventors: Saurabh Sahu (Bangalore, IN); Mariswamy Girish Chandra (Bangalore, IN); Kriti Kumar (Bangalore, IN); Achanna Anil Kumar (Bangalore, IN); Angshul Majumdar (New Delhi, IN)
Assignee: Tata Consultancy Services Limited
G01N29/043G01N29/069G01N29/14G01N29/4418G01N2291/0258
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,504,402
App. No.
18/372,300
Granted
Dec 23, 2025
Kind
B2
Abstract

In industrial inspection scenarios, early detection of machine malfunction is extremely essential as it helps in preventing any significant damage and the associated economic losses. Embodiments herein provide a method and system for an acoustic based anomaly detection in industrial machines using a beamforming and a sequential transform learning. Herein, the system employs two-stage multi-channel source separation technique that uses the well-known delay and sum beamforming followed by a recent data-driven sequential transform learning (STL) approach to obtain clean sources. The STL is a solution to linear state-space model where operators/matrices are learnt from data and is used here to model the dynamics of time-varying source signals for source separation. Subsequently, a reference template matching is employed on each separated source to detect an anomaly. The numerical results obtained with the Malfunctioning Industrial Machine Investigation and Inspection (MIMII) dataset demonstrate superior performance for source separation and anomaly detection.

Claims (33)

1 . A processor-implemented method for an acoustic based anomaly detection in an industrial machine comprising:

receiving, via a microphone array, a mixed audio signal of one or more spatially distributed audio sources, wherein the mixed audio signal includes noises;

beamforming, via one or more hardware processors, the received mixed audio signal using a delay-and-sum beamforming technique to obtain a constructive superposition of the audio signal along a predefined direction;

estimating, via the one or more hardware processors, at least one clean audio source of the one or more spatially distributed audio sources using a pre-trained sequential transform learning (STL), wherein the STL is used to model dynamics of the time-varying source signals for source separation;

determining, via the one or more hardware processors, a change in the estimated at least one clean audio source of the one or more spatially distributed audio sources obtained using STL and beamforming by comparing with a template of a normal audio source of the same industrial machine; and

detecting, via the one or more hardware processors, one or more abnormalities in the industrial machine based on the change determined in the associated clean audio source that is below a predefined threshold in terms of signal to noise ratio (SNR).

2 . The processor-implemented method of claim 1 , wherein the sequential transform learning (STL) model is trained on a mixture of normal audio of each of the one or more spatially distributed audio sources.

3 . The processor-implemented method of claim 1 , wherein the mixed audio signal received by each microphone of the microphone array is associated with different delays.

4 . The processor-implemented method of claim 1 , wherein the delay depends upon the spatial location of the audio source.

5 . The processor-implemented method of claim 1 , wherein the predefined threshold in terms of signal to noise ratio (SNR) is empirically calculated for each machine.

6 . A system for an acoustic based anomaly detection in an industrial machine comprising:

an input/output interface;

one or more hardware processors;

a memory in communication with the one or more hardware processors, wherein the one or more hardware processors are configured to execute programmed instructions stored in the memory, to:

receive a mixed audio signal of one or more spatially distributed audio sources via a microphone array, wherein the mixed audio signal includes noises;

beamform the received mixed audio signal using a delay-and-sum beamforming technique to obtain a constructive superposition of the audio signal along a predefined direction;

estimate at least one clean audio source of the one or more spatially distributed audio sources using a pretrained sequential transform learning (STL), wherein the STL is used to model dynamics of the time-varying source signals for source separation;

determine a change in the estimated at least one clean audio source of the one or more spatially distributed audio sources obtained using STL and beamforming by comparing with a template of a normal audio source of the same industrial machine; and

detect one or more abnormalities in the industrial machines based on the change determined in the associated clean audio source that is below a predefined threshold in terms of signal to noise ratio (SNR).

7 . The system of claim 6 , wherein the sequential transform learning (STL) model is trained on a mixture of normal audio of each of the one or more spatially distributed audio sources.

8 . The system of claim 6 , wherein the mixed audio signal received by each microphone of the microphone array is associated with different delays.

9 . The system of claim 6 , wherein the delay depends upon the spatial location of the audio source.

10 . The system of claim 6 , wherein the predefined threshold in terms of signal to noise ratio (SNR) is empirically calculated for each machine.

11 . One or more non-transitory machine-readable information storage mediums comprising one or more instructions which when executed by one or more hardware processors cause an acoustic based anomaly detection in an industrial machine comprising:

receiving, via a microphone array, a mixed audio signal of one or more spatially distributed audio sources, wherein the mixed audio signal includes noises;

beamforming the received mixed audio signal using a delay-and-sum beamforming technique to obtain a constructive superposition of the audio signal along a predefined direction;

estimating at least one clean audio source of the one or more spatially distributed audio sources using a pre-trained sequential transform learning (STL), wherein the STL is used to model dynamics of the time-varying source signals for source separation;

determining a change in the estimated at least one clean audio source of the one or more spatially distributed audio sources obtained using STL and beamforming by comparing with a template of a normal audio source of the same industrial machine; and

detecting one or more abnormalities in the industrial machine based on the change determined in the associated clean audio source that is below a predefined threshold in terms of signal to noise ratio (SNR).

12 . The one or more non-transitory machine-readable information storage mediums of claim 11 , wherein the sequential transform learning (STL) model is trained on a mixture of normal audio of each of the one or more spatially distributed audio sources.

13 . The one or more non-transitory machine-readable information storage mediums of claim 11 , wherein the mixed audio signal received by each microphone of the microphone array is associated with different delays.

14 . The one or more non-transitory machine-readable information storage mediums of claim 11 , wherein the delay depends upon the spatial location of the audio source.

15 . The one or more non-transitory machine-readable information storage mediums of claim 11 , wherein the predefined threshold in terms of signal to noise ratio (SNR) is empirically calculated for each machine.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2023
From: SAHU, SAURABH; CHANDRA, MARISWAMY GIRISH; KUMAR, KRITI; KUMAR, ACHANNA ANIL; MAJUMDAR, ANGSHUL
To: TATA CONSULTANCY SERVICES LIMITED
Reel/Frame 065008/0856 →
Priority Claims (1)
IN 202221062963 · Nov 3, 2022 · national
Continuity (1)
Related Publication 20240151690A1 · May 9, 2024
References Cited (8)
US 10735887B1 · McElveen · 2020 [cited by examiner]
US 11933695B2 · Lavid Ben Lulu · 2024 [cited by examiner]
US 11934183B2 · Rathore · 2024 [cited by examiner]
US 20240288340A1 · Sahu · 2024 [cited by examiner]
Acoustic-Base et al. machine Anomaly Detection Using Beamforming and Sequential Transform Learning Saurabh Sahu et al., IEEE vol. 7, No. 2, Feb. 2023. (Year: 2023). [cited by examiner]
Bagchi et al., “Combining Spectral Feature Mapping and Multi-Channel Model-Based Source Separation for Noise-Robust Automatic Speech Recognition,” (2015). [cited by applicant]
Chen et al., “A multichannel learning-based approach for sound source separation in reverberant environments,” EURASIP Journal on Audio, Speech, and Music Processing (2021). [cited by applicant]
Wang et al., “Localization Based Sequential Grouping for Continuous Speech Separation,” (2022). [cited by applicant]