IP Library › Granted Patent US 12,475,882
Granted Patent B2
US 12,475,882 · App. 18/448,628 · Granted Nov 18, 2025

Method and system for automatic speech recognition (ASR) using multi-task learned (MTL) embeddings

Inventors: Ashish Panda (Thane West, IN); Sunil Kumar Kopparapu (Thane West, IN); Aditya Raikar (Thane West, IN); Meetkumar Hemakshu Soni (Mumbai, IN)
Assignee: TATA CONSULTANCY SERVICES LIMITED
G10L15/16G10L15/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,882
App. No.
18/448,628
Granted
Nov 18, 2025
Kind
B2
Abstract

State of the art Acoustic Models (AM), which are trained using data from one environment, may fail to adapt to another environment, and as a result, application is restricted. The disclosure herein generally relates to speech signal processing, and, more particularly, to a method and system for Automatic Speech Recognition (ASR) using Multi-task Learned Embeddings (MTL). In this approach, MTL embeddings are extracted from an MTL neural network that has been trained using feature vectors from a plurality of speech files. The MTL embeddings are then used for generating an acoustic model, which maybe then used for the purpose of Automatic Speech Recognition, along with the feature vectors and the MTL embeddings.

Claims (30)

1 . A processor implemented method, comprising:

receiving, via one or more hardware processors, a plurality of speech files signals as input data;

extracting, via the one or more hardware processors, a plurality of feature vectors from the input data;

training, via the one or more hardware processors, a multi-task learning (MTL) neural network, using one or more of the plurality of feature vectors, wherein the trained MTL neural network able to perform: process the speech signals and extract one or more information from the speech signals, classify a room in which one or more speakers are present from whom the speech signals are being received;

extracting, via the one or more hardware processors, MTL embeddings from the MTL neural network, wherein extracted MTL embeddings are encoded with a plurality of environmental parameters and a plurality of speaker parameters, wherein the plurality of environmental parameters comprises information on classification of the room as one of large, small and medium as identified by the MTL neural network, and a noise information comprising a Signal to Noise Ratio (SNR), wherein the MTL embeddings are extracted by: initially performing frame-level features extraction using a plurality of Time delay neural network (TDNN) layers, further output of the plurality of TDNN layers is subjected to statistical pooling before performing an utterance level feature extraction and output of the utterance level feature extraction is processed using a plurality of softmax layers to obtain a speaker label and a room size label;

generating, via the one or more hardware processors, an acoustic model by training a deep neural network with a multi-task loss function, using the feature vectors and the MTL embeddings; and

decoding a transcript for an utterance from the acoustic model using the feature vectors extracted from the plurality of speech signals and the MTL embeddings.

2 . The processor implemented method of claim 1 , wherein the input data is received using an interface or received in real-time as speech signals from an environment using sensors.

3 . The processor implemented method of claim 1 , wherein the plurality of speaker parameters comprise speaker identity, gender, and age.

4 . A system, comprising:

one or more hardware processors;

a communication interface; and

a memory storing a plurality of instructions, wherein the plurality of instructions when executed, cause the one or more hardware processors to:

receive a plurality of speech signals as input data;

extract a plurality of feature vectors from the input data;

train a multi-task learning (MTL) neural network, using one or more of the plurality of feature vectors, wherein the trained MTL neural network able to perform: process the speech signals and extract one or more information from the speech signals, classify a room in which one or more speakers are present from whom the speech signals are being received;

extract MTL embeddings from the MTL neural network, wherein extracted MTL embeddings are encoded with a plurality of environmental parameters and a plurality of speaker parameters, wherein the plurality of environmental parameters comprises information on classification of the room as one of large, small and medium as identified by the MTL neural network, and a noise information comprising a Signal to Noise Ratio (SNR), wherein the MTL embeddings are extracted by: initially performing frame-level features extraction using a plurality of Time delay neural network (TDNN) layers, further output of the plurality of TDNN layers is subjected to statistical pooling before performing an utterance level feature extraction and output of the utterance level feature extraction is processed using a plurality of softmax layers to obtain a speaker label and a room size label;

generate an acoustic model by training a deep neural network with a multi-task loss function, using the feature vectors and the MTL embeddings; and

decoding a transcript for an utterance from the acoustic model using the feature vectors extracted from the plurality of speech signals and the MTL embeddings.

5 . The system of claim 4 , wherein the one or more hardware processors are further configured to receive the input data using an interface or received in real-time as speech signals from an environment using sensors.

6 . The system of claim 4 , wherein the one or more hardware processors are configured to obtain information on speaker identity, gender, and age as the plurality of speaker parameters.

7 . One or more non-transitory machine-readable information storage mediums comprising one or more instructions which when executed by one or more hardware processors cause:

receiving a plurality of speech signals as input data;

extracting a plurality of feature vectors from the input data;

training a multi-task learning (MTL) neural network, using one or more of the plurality of feature vectors, wherein the trained MTL neural network able to perform: process the speech signals and extract one or more information from the speech signals, classify a room in which one or more speakers are present from whom the speech signals are being received;

extracting MTL embeddings from the MTL neural network, wherein extracted MTL embeddings are encoded with a plurality of environmental parameters and a plurality of speaker parameters, wherein the plurality of environmental parameters comprises information on classification of the room as one of large, small and medium as identified by the MTL neural network, and a noise information comprising a Signal to Noise Ratio (SNR), wherein the MTL embeddings are extracted by: initially performing frame-level features extraction using a plurality of Time delay neural network (TDNN) layers, further output of the plurality of TDNN layers is subjected to statistical pooling before performing an utterance level feature extraction and output of the utterance level feature extraction is processed using a plurality of softmax layers to obtain a speaker label and a room size label;

generating an acoustic model by training a deep neural network with a multi-task loss function, using the feature vectors and the MTL embeddings; and

decoding a transcript for an utterance from the acoustic model using the feature vectors extracted from the plurality of speech signals and the MTL embeddings.

8 . The one or more non-transitory machine-readable information storage mediums of claim 7 , further comprising receiving the input data using an interface or received in real-time as speech signals from an environment using sensors.

9 . The one or more non-transitory machine-readable information storage mediums of claim 7 , wherein the plurality of speaker parameters comprise speaker identity, gender, and age.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2023
From: PANDA, ASHISH; KOPPARAPU, SUNIL KUMAR; RAIKAR, ADITYA; SONI, MEETKUMAR HEMAKSHU
To: TATA CONSULTANCY SERVICES LIMITED
Reel/Frame 064587/0845 →
Priority Claims (1)
IN 202221048968 · Aug 26, 2022 · national
Continuity (1)
Related Publication 20240071373A1 · Feb 29, 2024
References Cited (17)
US 10347241B1 · Meng · 2019 [cited by examiner]
US 12148437B2 · Sharma · 2024 [cited by examiner]
US 20200312346A1 · Fazeli · 2020 [cited by examiner]
US 20220199095A1 · Chang · 2022 [cited by examiner]
US 20230260521A1 · Slocum · 2023 [cited by examiner]
US 20240005908A1 · Sharma · 2024 [cited by examiner]
Giri, Ritwik, et al. “Improving speech recognition in reverberation using a room-aware deep neural network and multi-task learning.” 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)… [cited by examiner]
Ji, Xuan, et al. “Speaker-aware target speaker enhancement by jointly learning with speaker embedding extraction.” ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE… [cited by examiner]
Khare, Aparna, et al. “Multi-modal embeddings using multi-task learning for emotion recognition.” arXiv preprint arXiv:2009.05019. Sep. 2020, pp. 1-5. (Year: 2020). [cited by examiner]
Kim, Suyoun, et al. “Environmental noise embeddings for robust speech recognition.” arXiv preprint arXiv:1601.02553, Sep. 2016, pp. 1-5. (Year: 2016). [cited by examiner]
Pironkov, Gueorgui, Stéphane Dupont, and Thierry Dutoit. “Speaker-aware multi-task learning for automatic speech recognition.” 2016 23rd International Conference on Pattern Recognition (ICPR). IEEE, Dec. 2016, pp. 1-6.… [cited by examiner]
Raikar, Aditya, et al. “Acoustic Model Adaptation In Reverberant Conditions Using Multi-task Learned Embeddings.” 2022 30th European Signal Processing Conference (EUSIPCO). IEEE, Sep. 2022, pp. 145-149. (Year: 2022). [cited by examiner]
Tao, Fei, et al. “End-to-end audiovisual speech recognition system with multitask learning.” IEEE Transactions on Multimedia 23, Feb. 2020, pp. 1-11. (Year: 2020). [cited by examiner]
Zhou, Jianfeng, et al. “Training multi-task adversarial network for extracting noise-robust speaker embedding.” ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, M… [cited by examiner]
Jain, Abhinav et al., “Improved Accented Speech Recognition Using Accent Embeddings and Multi-task Learning”, Title of the item: Interspeech 2018, Date: Sep. 2-6, 2018, Publisher: IITB Link: https://www.cse.litb.ac.in/˜… [cited by applicant]
He, Weipeng et al., “Multi-task Neural Network for Robust Multiple Speaker Embedding Extraction”, Title of the item: Interspeech 2021, Date: Aug. 30-Sep. 3, 2021, Publisher: ISCA, Link: https://www.idiap.ch/˜odobez/publ… [cited by applicant]
Li, Ke et al., “Speaker Adaptation for End-To-End CTC Models”, Title of the item: 2018 IEEE Spoken Language Technology Workshop (SLT), Date: Jan. 4, 2019, Publisher: arXiv. Link: https://arxiv.org/pdf/1901.01239.pdf. [cited by applicant]