IP Library › Granted Patent US 10,964,326
Granted Patent B2
US 10,964,326 · App. 15/435,259 · Granted Mar 30, 2021

System and method for audio-visual speech recognition

Inventor: Ian Richard Lane (Sunnyvale, CA)
Assignee: CARNEGIE MELLON UNIVERSITY, a Pennsylvania Non-Profit Corporation
G10L15/25G06K9/00268G06K9/6267G06T7/11G10L15/02G10L15/16G06T2207/10004G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,964,326
App. No.
15/435,259
Granted
Mar 30, 2021
Kind
B2
Abstract

Disclosed herein is method of performing speech recognition using audio and visual information, where the visual information provides data related to a person's face. Image preprocessing identifies regions of interest, which is then combined with the audio data before being processed by a speech recognition engine.

Claims (24)

1. A method of performing speech recognition at a distance comprising:

obtaining audio information from a plurality of microphones positioned at a distance from a speaker;

obtaining visual information from an image capture device;

pre-processing the visual information, wherein pre-processing comprises:

using a recurrent deep neural network model to identify relevant image features by classifying a pixel in a first frame of the visual information using features from a pixel in the same location from a prior frame of the visual information;

identifying a region of interest in the visual information; and

using the region of interest, aligning the visual information with the acoustic information and classifying individual frames in the visual information into context-dependent phonetic states;

combining the audio information and visual information within a single deep neural network classifier; and

performing a speech recognition process on the combined audio information and visual information, wherein the speech recognition process comprises:

generating observation probabilities for context-dependent phonetic states using a joint audio-visual observation model, and

conducting a search in a standard speech recognition engine using the observation probabilities.

2. The method of claim 1 , wherein pre-processing further comprises:

determining whether a speaker is present in the visual information.

3. The method of claim 1 , wherein the standard speech recognition engine is a WFST-based speech recognition engine.

4. The method of claim 1 , wherein pre-processing further comprises:

defining the region of interest by generating a probability distribution of classes for each pixel in the visual information.

5. The method of claim 1 , further comprising:

identifying a location of the speaker relative to the image capture device.

6. The method of claim 1 , wherein pre-processing further comprises:

scaling the region of interest.

7. The method of claim 1 , wherein combining the audio information and visual information comprises:

combining the audio information and the visual information in an output layer of the single deep neural network classifier.

8. The method of claim 1 , wherein combining the audio information and visual information comprises:

combining the audio information and the visual information in a second hidden layer of the single deep neural network classifier.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2018
From: LANE, IAN RICHARD
To: CARNEGIE MELLON UNIVERSITY
Reel/Frame 047166/0388 →
Continuity (2)
Provisional Application 62389061 · Feb 16, 2016
Related Publication 20170236516A1 · Aug 17, 2017