IP Library › Granted Patent US 12,340,794
Granted Patent B2
US 12,340,794 · App. 17/874,899 · Granted Jun 24, 2025

Language identification classifier trained using encoded audio from encoder of pre-trained speech-to-text system

Inventor: Zvi Kons (Yoqneam Ilit, IL)
Assignee: International Business Machines Corporation
G10L15/063G06F40/20G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,794
App. No.
17/874,899
Granted
Jun 24, 2025
Kind
B2
Abstract

An example system includes a processor to receive encoded audio from an encoder of a pre-trained speech-to-text (STT) model. The processor is to further train a language identification (LID) classifier to detect a language of the encoded audio using training samples labeled by language.

Claims (51)

1. A system, comprising a processor to:

receive an encoded audio from an encoder of a pre-trained speech-to-text (STT) model;

train a language identification (LID) classifier to detect a language of the encoded audio using training samples labeled by language;

and the system further comprising the processor to:

receive an audio sample to be converted into text;

encode the audio sample into a second encoded audio;

classify the second encoded audio via the trained LID classifier; and

generate, in response to detecting that the second encoded audio is classified as a target language, text in the target language based on the second encoded audio and prediction from a predictor of the pre-trained STT model;

wherein the LID classifier includes weights and linear projections being multiplied in a multiplier; and

wherein the LID classifier further includes an average pooling layer as a multi-head weighted-average pooling layer.

2. The system of claim 1 , wherein the encoder comprises a recurrent neural network transducer (RNN-T) encoder.

3. The system of claim 1 , wherein the encoder is pre-trained on one language.

4. The system of claim 1 , wherein the STT model comprises a plurality of predictors dedicated to different languages, wherein the LID classifier is to classify a second encoded audio corresponding to an audio sample to be converted into text and select a corresponding dedicated predictor based on the classification.

5. The system of claim 1 , wherein the encoder of the STT model is pre-trained with the plurality of predictors for different languages.

6. The system of claim 1 , wherein the encoded audio comprises a frame-level feature vector.

7. A computer-implemented method, comprising:

receiving, via a processor, an encoded audio from an encoder of a pre-trained speech-to-text (STT) model;

training, via the processor, a language identification (LID) classifier to detect a language of the encoded audio using training samples labeled by language;

the computer-implemented method further comprising:

receiving, via the processor, an audio sample to be converted into text;

encoding, via the processor, the audio sample into a second encoded audio;

classifying, via the processor, the second encoded audio via the trained LID classifier; and

sending, via the processor, the second encoded audio to a dedicated predictor of a plurality of predictors dedicated to different languages based on the classification;

wherein the LID classifier includes weights and linear projections being multiplied in a multiplier; and

wherein the LID classifier further includes an average pooling layer as a multi-head weighted-average pooling layer.

8. The computer-implemented method of claim 7 , comprising:

receiving, via the processor, an audio sample to be converted into text;

encoding, via the processor, the audio sample into a second encoded audio; and

classifying, via the processor, the second encoded audio via the trained LID classifier.

9. The computer-implemented method of claim 8 , comprising stopping, via the processor, processing of the audio sample in response to detecting that the second encoded audio is not classified as a target language.

10. The computer-implemented method of claim 8 , comprising generating, via the processor, text in the target language based on the second encoded audio and prediction from a predictor of the pre-trained STT model in response to detecting that the second encoded audio is classified as a target language.

11. The computer-implemented method of claim 7 , comprising generating the text from the second encoded audio via the dedicated predictor.

12. The computer-implemented method of claim 7 , wherein classifying the second encoded audio comprises applying a softmax function to linear projections of pooled weighted averages and classifying the second encoded audio based on a language class with a highest decimal probability.

13. A computer program product for training language identification classifiers, the computer program product comprising a computer-readable storage medium having program code embodied therewith, the program code executable by a processor to cause the processor to:

receive encoded audio from an encoder of a pre-trained speech-to-text (STT) model;

train a language identification (LID) classifier to detect a language of the encoded audio using training samples labeled by language;

the computer program product further comprising program code executable by the processor to:

receive an audio sample to be converted into text;

encode the audio sample into a second encoded audio;

classify the second encoded audio via the trained LID classifier;

send the second encoded audio to a dedicated predictor of a plurality of predictors dedicated to different languages based on the classification; and

generate the text from the encoded audio via the dedicated predictor;

wherein the LID classifier includes weights and linear projections being multiplied in a multiplier; and

wherein the LID classifier further includes an average pooling layer as a multi-head weighted-average pooling layer.

14. The computer program product of claim 13 , further comprising program code executable by the processor to:

receive an audio sample to be converted into text;

encode the audio sample into a second encoded audio; and

classify the second encoded audio via the trained LID classifier.

15. The computer program product of claim 14 , further comprising program code executable by the processor to stop processing of the audio sample in response to detecting that the encoded audio is not classified as a target language.

16. The computer program product of claim 14 , further comprising program code executable by the processor to generate text in the target language based on the second encoded audio and prediction from a predictor of the pre-trained STT model in response to detecting that the second encoded audio is classified as a target language.

17. The computer program product of claim 13 , further comprising program code executable by the processor to apply a softmax function to linear projections of a pooled weighted averages and classify the second encoded audio based on a language class with a highest decimal probability.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 27, 2022
From: KONS, ZVI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 060644/0732 →
Continuity (1)
Related Publication 20240038216A1 · Feb 1, 2024
References Cited (18)
US 11238845B2 · Chen et al. · 2022 [cited by applicant]
US 20170011734A1 · Ganapathy et al. · 2017 [cited by applicant]
US 20180067918A1 · Bellegarda et al. · 2018 [cited by applicant]
US 20200380215A1 · Kannan · 2020 [cited by examiner]
US 20210005183A1 · Lee · 2021 [cited by examiner]
US 20230306958A1 · Zhang · 2023 [cited by examiner]
CN 117475996A · 2024 [cited by applicant]
JP 2024018989A · 2024 [cited by applicant]
WO WO2020113031A1 · 2020 [cited by examiner]
Kons, Zvi et al. “Extending RNN-T-based speech recognition systems with emotion and language classification.” arXiv preprint arXiv:2207.13965 (2022). (Year: 2022). [cited by examiner]
Joshi et al., “Mutliple Softmax Architecture for Streaming Multilingual End-to-End ASR Systems”, Microsoft Corporation, Interspeech 2021, Aug. 30, 2021-Sep. 3, 2021, 5 pages. [cited by applicant]
Punjabi et al., “Joint ASR and Language Identification Using RNN-T: An Efficient Approach to Dynamic Language Switching”, Alexa Machine Learning, Amazon, 2021, IEEE, Downloaded Feb. 9, 2022, pp. 7218-7222. [cited by applicant]
Punjabi et al., “Streaming End-to-End Bilingual ASR Systems with Joint Language Identification”, Alexa Machine earning, Amazon, Jul. 8, 2020, 5 pages. [cited by applicant]
Waters et al., “Leveraging Language ID in Multilingual End-to-End Speech Recognition”, Google Inc., USA, IEEE, 2019, Downloaded Mar. 3, 2022, pp. 928-935. [cited by applicant]
Aronowitz et al., “Towards a Common Speech Analysis Engine.”, arXiv:2203.00613, IBM Research AI, May 1, 2022, 05 pages. [cited by applicant]
Baevski et al., v“wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations”, DOI: 10.48550/ARXIV.2006.11477. URL: https://arxiv.org/abs/2006.11477, Oct. 22, 2020, 19 pages. [cited by applicant]
Hsu et al., “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units”, 2021. doi: 10.48550/ARXIV. 2106.07447. URL: https://arxiv.org/abs/2106.07447, Jun. 14, 2021, 10 pages. [cited by applicant]
Zhu et al., “Multilingual Speech Recognition with Self-Attention Structured Parameterization”, Interspeech 2020, Oct. 25-29, 2020, 05 pages, http://dx.doi.org/10.21437/Interspeech.2020-2847. [cited by applicant]