IP Library › Granted Patent US 12,444,405
Granted Patent B2
US 12,444,405 · App. 18/310,598 · Granted Oct 14, 2025

Textual knowledge transfer for improved speech recognition and understanding

Inventors: Samuel Thomas (White Plains, NY); Vishal Sunder (Columbus, OH); Hong-Kwang Kuo (Pleasantville, NY); Brian E. D. Kingsbury (Cortlandt Manor, NY); Eric Fosler-Lussier (Columbus, OH); George Andrei Saon (Stamford, CT)
Assignees: International Business Machines Corporation; Ohio State Innovation Foundation
G10L15/063G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,405
App. No.
18/310,598
Granted
Oct 14, 2025
Kind
B2
Abstract

Systems, computer-implemented methods, and computer program products to facilitate fine-grained textual knowledge transfer to improve speech recognition and understanding are provided. According to an embodiment, a system can comprise a processor that executes components stored in memory. The computer executable components comprise deriving component that can derive one or more speech-based embeddings from an utterance via a speech encoder. The computer executable components can comprise a cross-attention component that can align, at a token level, one or more large language model (LLM) based sentence embeddings with the one or more speech-based embeddings. The computer executable components can comprise a loss component that can combine an alignment loss and an automatic speech recognition (ASR) loss.

Claims (47)

1. A system comprising:

a processor that executes computer-executable components stored in a non-transitory computer-readable memory, wherein the computer-executable components comprise:

a deriving component that derives one or more speech-based embeddings from an utterance via a speech encoder;

a cross-attention component that aligns, at a token level, one or more Large Language Model (LLM) based sentence embeddings with the one or more speech-based embeddings;

a loss component that combines an alignment loss and an Automatic Speech Recognition (ASR) loss; and

a training component that trains an ASR system with an end-to-end framework using the loss component and the cross-attention component to produce one or more enriched embeddings.

2. The system of claim 1 , wherein the cross-attention component determines a contrastive loss between the one or more LLM based sentence embeddings and the one or more speech-based embeddings.

3. The system of claim 1 , wherein the cross-attention component uses non-contextual (NC) embeddings as queries to align the one or more LLM based sentence embeddings with the one or more speech-based embeddings.

4. The system of claim 1 , wherein the ASR system is adapted to perform a Spoken Language Understanding (SLU) task.

5. The system of claim 4 , wherein the cross-attention component creates speech-based embeddings that are aligned with the one or more LLM based sentence embeddings, and the speech-based embeddings are fused with one or more ASR based embeddings.

6. The system of claim 5 , wherein the cross-attention component creates a speech-based summary token to approximate the one or more LLM based sentence embeddings to determine an SLU loss.

7. The system of claim 6 , wherein a gating mechanism integrates the speech-based summary token along with other embeddings of an end-to-end ASR model to improve a final training loss.

8. A computer-implemented method comprising:

deriving, by a system operatively coupled to a processor, one or more speech-based embeddings from an utterance via a speech encoder;

aligning at a token level, by the system, one or more Large Language Model (LLM) based sentence embeddings with the one or more speech-based embeddings;

combining, by the system, an alignment loss and an Automatic Speech Recognition (ASR) loss; and

training, by the system, an ASR system with an end-to-end framework using alignment loss, the ASR loss, and the one or more LLM based sentence embeddings aligned with the one or more speech-based embeddings to produce one or more enriched embeddings.

9. The computer-implemented method of claim 8 , further comprising:

determining, by the system, a contrastive loss between the one or more LLM based sentence embeddings and the one or more speech-based embeddings.

10. The computer-implemented method of claim 8 , further comprising:

using, by the system, non-contextual embeddings as queries to align the one or more LLM based sentence embeddings with the one or more speech-based embeddings.

11. The computer-implemented method of claim 8 , further comprising:

adapting, by the system, the ASR system to perform a Spoken Language Understanding (SLU) task.

12. The computer-implemented method of claim 11 , further comprising:

creating, by the system, speech-based embeddings aligned with the one or more LLM based sentence embeddings, and

fusing, by the system, the speech-based embeddings with one or more ASR based embeddings.

13. The computer-implemented method of claim 12 , further comprising:

creating, by the system, a speech-based summary token to approximate the one or more LLM based sentence embeddings to determine an SLU loss.

14. The computer-implemented method of claim 13 , further comprising:

integrating, by the system, the speech-based summary token along with other embeddings of an end-to-end ASR model to improve a final training loss.

15. A computer program product comprising a non-transitory computer readable memory having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to:

derive, by the processor, one or more speech-based embeddings from an utterance via a speech encoder;

align at a token level, by the processor, one or more Large Language Model (LLM) based sentence embeddings with the one or more speech-based embeddings;

combine, by the processor, an alignment loss and an Automatic Speech Recognition (ASR) loss; and

train, by the processor, an Automatic Speech Recognition (ASR) system, with an end-to-end framework using the alignment loss, the ASR loss, and the one or more LLM based sentence embeddings aligned with the one or more speech-based embeddings to produce one or more enriched embeddings.

16. The computer program product of claim 15 , wherein the program instructions are further executable to cause the processor to:

determine, by the processor, a contrastive loss between the one or more LLM based sentence embeddings and the one or more speech-based embeddings.

17. The computer program product of claim 15 , wherein the program instructions are further executable to cause the processor to:

use, by the processor, non-contextual embeddings as queries to align the one or more LLM based sentence embeddings with the one or more speech-based embeddings.

18. The computer program product of claim 15 , wherein the program instructions are further executable to cause the processor to:

adapt, by the processor, the ASR system to perform a Spoken Language Understanding (SLU) task; and

create, by the processor, speech-based embeddings aligned with the one or more LLM based sentence embeddings, and

fusing, by the system, the speech-based embeddings with one or more ASR based embeddings.

19. The computer program product of claim 18 , wherein the program instructions are further executable to cause the processor to:

create, by the processor, a speech-based summary token to approximate the one or more LLM based sentence embeddings to determine an SLU loss.

20. The computer program product of claim 19 , wherein the program instructions are further executable to cause the processor to:

integrate, by the processor, the speech-based summary token along with other embeddings of an end-to-end ASR model to improve a final training loss.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2023
From: THOMAS, SAMUEL; SUNDER, VISHAL; KUO, HONG-KWANG; KINGSBURY, BRIAN E. D.; SAON, GEORGE ANDREI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 063502/0229 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2023
From: FOSLER-LUSSIER, ERIC
To: OHIO STATE INNOVATION FOUNDATION
Reel/Frame 063502/0319 →
Continuity (1)
Related Publication 20240371361A1 · Nov 7, 2024
References Cited (22)
US 11562735B1 · Gupta · 2023 [cited by examiner]
US 20210209139A1 · Wu · 2021 [cited by examiner]
US 20220230625A1 · Zhu et al. · 2022 [cited by applicant]
US 20220319494A1 · Thomas et al. · 2022 [cited by applicant]
US 20220382979A1 · Klein et al. · 2022 [cited by applicant]
US 20230169281A1 · Zheng · 2023 [cited by examiner]
US 20230214718A1 · Hong · 2023 [cited by examiner]
US 20230223018A1 · Xing · 2023 [cited by examiner]
US 20230386450A1 · Eby · 2023 [cited by examiner]
US 20240045925A1 · Savvides · 2024 [cited by examiner]
US 20240135238A1 · Sattigeri · 2024 [cited by examiner]
US 20240256948A1 · Betthauser · 2024 [cited by examiner]
US 20240274286A1 · Chen · 2024 [cited by examiner]
US 20240311579A1 · Dong · 2024 [cited by examiner]
US 20240331235A1 · Smock · 2024 [cited by examiner]
US 20240354319A1 · Dinu · 2024 [cited by examiner]
US 20240362272A1 · Lee · 2024 [cited by examiner]
US 20240386015A1 · Crabtree · 2024 [cited by examiner]
CN 114242071A · 2022 [cited by applicant]
Sunder, et al., “Fine-Grained Textual Knowledge Transfer to Improve RNN Transducers for Speech Recognition and Understanding,” May 5, 2022. [cited by applicant]
Lai, Cheng-I, et al., “Towards Semi-Supervised Semantics Understanding from Speech,” arXiv preprint arXiv:2011.06195, arXiv:2011.06195v1 [cs.CL] Nov. 11, 2020. [cited by applicant]
Sunder, et al., “Tokenwise Contrastive Pretraining for Finer Speech-to-BERT Alignment in End-to-End Speech-to-Intent Systems,” Jul. 1, 2022, arXiv:2204.05188v2 [cs.CL]. [cited by applicant]