IP Library › Granted Patent US 11,908,454
Granted Patent B2
US 11,908,454 · App. 17/539,752 · Granted Feb 20, 2024

Integrating text inputs for training and adapting neural network transducer ASR models

Inventors: Samuel Thomas (White Plains, NY); Hong-Kwang Kuo (Pleasantville, NY); Brian E. D. Kingsbury (Cortlandt Manor, NY); George Andrei Saon (Stamford, CT); Gakuto Kurata (Tokyo, JP)
Assignee: International Business Machines Corporation
G10L15/063G06N3/08G10L21/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,908,454
App. No.
17/539,752
Filed
Dec 1, 2021
Granted
Feb 20, 2024
Kind
B2
Art Unit
2691
USPC
704/200
Abstract

A processor-implemented method trains an automatic speech recognition system using speech data and text data. A computer device receives speech data, and generates a spectrogram based on the speech data. The computing device receives text data associated with an entire corpus of text data, and generates a textogram based upon the text data. The computing device trains an automatic speech recognition system using the spectrogram and the textogram.

Claims (53)

1. A method of using a computing device to train an automatic speech recognition system using speech data and text data, the method comprising:

receiving, by a computing device, speech data;

generating, by the computing device, a spectrogram based on the speech data;

receiving, by the computing device, text data associated with an entire corpus of text data;

generating, by the computing device, a textogram based upon the text data; and

training, by the computing device, an automatic speech recognition system using the spectrogram and the textogram.

2. The method of claim 1 , further comprising:

adapting, by the computing device, an automatic speech recognition model into a spoken language understanding model using only text data without a corresponding spectrogram; and

utilizing, by the computing device, the spoken language understanding model to interpret the speech data.

3. The method of claim 1 , wherein the speech data is first speech data, and wherein the method further comprises:

modifying, by the computing device, the automatic speech recognition model with a second speech data that is different from the first speech data.

4. The method of claim 1 , wherein the speech data is first speech data, and wherein the method further comprises:

modifying, by the computing device, the automatic speech recognition model with text data from a second speech data that is different from the first speech data.

5. The method of claim 1 , wherein the speech data is in a first speech language, and wherein the method further comprises:

modifying, by the computing device, the automatic speech recognition model with text data from a second speech language that is different from the first speech language.

6. The method of claim 1 , wherein the automatic speech recognition system is based on an automatic speech recognition model of the speech data, and wherein the method further comprises:

modifying, by the computing device, the automatic speech recognition model using independent text data that is different from text data associated with the speech data.

7. A computer program product for training an automatic speech recognition system using speech data and text data, wherein the computer program product comprises a non-transitory computer readable storage device having program instructions embodied therewith, the program instructions readable and executable by a computer to perform a method comprising:

receiving speech data;

generating a spectrogram based on the speech data;

receiving text data associated with an entire corpus of text data;

generating a textogram based upon the text data; and

training an automatic speech recognition system using the spectrogram and the textogram.

8. The computer program product of claim 7 , wherein the method further comprises:

adapting an automatic speech recognition model into a spoken language understanding model using only text data without a corresponding spectrogram; and

utilizing the spoken language understanding model to interpret the speech data.

9. The computer program product of claim 7 , wherein the speech data is first speech data, and wherein the method further comprises:

modifying the automatic speech recognition model with a second speech data that is different from the first speech data.

10. The computer program product of claim 7 , wherein the speech data is first speech data, and wherein the method further comprises:

modifying the automatic speech recognition model with text data from a second speech data that is different from the first speech data.

11. The computer program product of claim 7 , wherein the speech data is in a first speech language, and wherein the method further comprises:

modifying the automatic speech recognition model with text data from a second speech language that is different from the first speech language.

12. The computer program product of claim 7 , wherein the automatic speech recognition system is based on an automatic speech recognition model of the speech data, and wherein the method further comprises:

modifying the automatic speech recognition model using independent text data that is different from text data associated with the speech data.

13. The computer program product of claim 7 , wherein the program code is provided as a service in a cloud environment.

14. A computer system comprising one or more processors, one or more computer readable memories, and one or more computer readable non-transitory storage mediums, and program instructions stored on at least one of the one or more computer readable non-transitory storage mediums for execution by at least one of the one or more processors via at least one of the one or more computer readable memories, the stored program instructions executed to perform a method comprising:

receiving speech data;

generating a spectrogram based on the speech data;

receiving text data associated with an entire corpus of text data;

generating a textogram based upon the text data; and

training an automatic speech recognition system using the spectrogram and the textogram.

15. The computer system of claim 14 , wherein the method further comprises:

adapting an automatic speech recognition model into a spoken language understanding model using only text data without a corresponding spectrogram; and

utilizing the spoken language understanding model to interpret the speech data.

16. The computer system of claim 14 , wherein the speech data is first speech data, and wherein the method further comprises:

modifying the automatic speech recognition model with a second speech data that is different from the first speech data.

17. The computer system of claim 14 , wherein the speech data is first speech data, and wherein the method further comprises:

modifying the automatic speech recognition model with text data from a second speech data that is different from the first speech data.

18. The computer system of claim 14 , wherein the speech data is in a first speech language, and wherein the method further comprises:

modifying the automatic speech recognition model with text data from a second speech language that is different from the first speech language.

19. The computer system of claim 14 , wherein the automatic speech recognition system is based on an automatic speech recognition model of the speech data, and wherein the method further comprises:

modifying the automatic speech recognition model using independent text data that is different from text data associated with the speech data.

20. The computer system of claim 14 , wherein the program code is provided as a service in a cloud environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2021
From: THOMAS, SAMUEL; KUO, HONG-KWANG; KINGSBURY, BRIAN E.D.; SAON, GEORGE ANDREI; KURATA, GAKUTO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058259/0435 →
Continuity (1)
Related Publication 20230169954A1 · Jun 1, 2023
Cited By (2)
US 12,424,206 US 12,711,950