IP Library › Granted Patent US 10,235,991
Granted Patent B2
US 10,235,991 · App. 15/672,486 · Granted Mar 19, 2019

Hybrid phoneme, diphone, morpheme, and word-level deep neural networks

Inventors: Jintao Jiang (Great Falls, VA); Hassan Sawaf (Los Gatos, CA); Mudar Yaghi (McLean, VA)
Assignee: AppTek, Inc.
G10L15/063G10L15/02G10L15/16G10L15/187G10L15/32G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,235,991
App. No.
15/672,486
Granted
Mar 19, 2019
Kind
B2
Abstract

A hybrid frame, phone, diphone, morpheme, and word-level Deep Neural Networks (DNN) in model training and applications-is based on training a regular ASR system, which can be based on Gaussian Mixture Models (GMM) or DNN. All the training data (in the format of features) are aligned with the transcripts in terms of phonemes and words with the timing information and new features are formed in terms of phonemes, diphones, morphemes, and up to words. Regular ASR produces a result lattice with timing information for each word. A feature is then extracted and sent to the word-level DNN for scoring Phoneme features are sent to corresponding DNNs for training. Scores are combined to form the word level scores, a rescored lattice and a new recognition result.

Claims (22)

1. A system for processing audio, comprising:

a memory, including program instructions for training DNN models, preparing features and aligning units of at least one of phonemes, diphones, morphemes, and words to audio independent of frame boundaries; and

a processor, coupled to the memory, that is capable of executing the program instructions to generate a DNN in the memory, receive the audio and assign corresponding aligned units of data to levels of the DNN, and process training data to create frame, phoneme, diphone, morpheme, and word-level DNN models separately; and

wherein the processor is further capable of executing the program instructions, and processing an audio file to create scores for annotating a lattice and rescoring the lattice based on the phoneme, diphone, morpheme, and word-level DNN models.

2. The system according to claim 1 , further wherein:

the memory further includes a program for combining the phoneme DNN scores into a word score.

3. The system according to claim 2 , wherein the processor further executes the programs to combine phoneme or phoneme, diphone, and/or morpheme scores into a word score.

4. The system according to claim 1 , wherein memory further includes a program for combining phoneme, diphone, and morpheme scores into word scores.

5. The system according to claim 4 , wherein:

the memory further includes a program for combining word-level DNN scores with traditional confidence scores to form new confidence scores.

6. The system according to claim 1 , wherein the processor further executes the programs to combine phoneme or phoneme, diphone, and/or morpheme scores into a word score.

7. The system according to claim 6 , wherein the processor further executes the programs to align respective features with at least two of phonemes, diphones, morphemes, and words and apply zero padding when a duration of the respective feature is less than a predetermined amount associated with each respective feature.

8. The system according to claim 6 , wherein:

the memory further includes a program for combining word-level DNN scores with traditional confidence scores to form new confidence scores.

9. A method of training speech recognition systems, comprising:

training a DNN system using a traditional ASR tool based on audio and transcript data;

aligning features with at least two selected ones of phonemes, diphones, morphemes, and words independent of frames;

preparing new features and alignments for the respective selected ones of the phonemes, diphones, morphemes, and words;

normalizing the new features;

training new DNN models based on the new features and alignments separately for each of the respective selected ones of the phonemes, diphones, morphemes, and words;

post processing traditional ASR result lattices; and

applying the newly trained DNN models to the result lattices by rescoring words based on combinations of the selected phoneme, diphone, and morpheme scores associated with each rescored word.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2018
From: JIANG, JINTAO; SAWAF, HASSAN; YAGHI, MUDAR
To: APPTEK, INC.
Reel/Frame 047296/0928 →
Continuity (2)
Provisional Application 62372539 · Aug 9, 2016
Related Publication 20180047385A1 · Feb 15, 2018