IP Library Granted Patent US 11,017,779
Granted Patent B2
US 11,017,779 · App. 16/277,136 · Granted May 25, 2021

System and method for speech understanding via integrated audio and visual based speech recognition

Inventors: Nishant Shukla (Hermosa Beach, CA); Ashwin Dharne (Irvine, CA)
Assignee: DMAI, INC.
G10L15/32G06K9/00G10L15/02G10L15/22G10L15/25G10L15/30G10L13/00G10L2015/025G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,017,779
App. No.
16/277,136
Granted
May 25, 2021
Kind
B2
Abstract

The present teaching relates to method, system, medium, and implementations for speech recognition. An audio signal is received that represents a speech of a user engaged in a dialogue. A visual signal is received that captures the user uttering the speech. A first speech recognition result is obtained by performing audio based speech recognition based on the audio signal. Based on the visual signal, lip movement of the user is detected and a second speech recognition result is obtained by performing lip reading based speech recognition. The first and the second speech recognition results are then integrated to generate an integrated speech recognition result.

Claims (60)

1. A method implemented on at least one machine including at least one processor, memory, and communication platform capable of connecting to a network for speech recognition, the method comprising:

receiving an audio signal representing a speech of a user engaged in a dialogue and a visual signal capturing the user uttering the speech;

obtaining a first speech recognition result by

identifying phonemes from the audio signal,

performing audio based speech recognition based on the phonemes to generate the first speech recognition result, and

converting the phonemes into a first set of visemes;

obtaining a second speech recognition result by

identifying a second set of visemes based on at least one of lip shape and movement of the user detected from the visual signal, and

performing lip reading based speech recognition to generate the second speech recognition result based on the second set of visemes; and

generating an integrated speech recognition result based on the first and second set of visemes, wherein

the first speech recognition result is selected as the integrated speech recognition result if a similarity between the first and second speech recognition results satisfies a first criterion, and

the second speech recognition result is selected as the integrated speech recognition result if the similarity satisfies a second criterion.

2. The method of claim 1 , wherein

the first speech recognition result includes a first plurality of words obtained via the phonemes; and

the second speech recognition result includes a second plurality of words obtained via the second set of visemes.

3. The method of claim 1 , further comprising synchronizing the first speech recognition result and the second speech recognition result, wherein each of the phonemes in the first speech recognition result is synchronized with a corresponding viseme in the second set of visemes of the second speech recognition result.

4. The method of claim 3 , wherein the similarity between the first and second speech recognition results is determined by:

determining, pairs of visemes, each of which includes one from the first set of visemes converted from the phonemes and a corresponding one from the second set of visemes that is synchronized with a corresponding phoneme from which the one viseme in the pair is converted;

assessing the similarity between the first and second speech recognition results based on the pairs of visemes.

5. The method of claim 1 , wherein the integrated speech recognition is given a confidence score, determined based on whether the first or second speech recognition result is selected.

6. Machine readable and non-transitory medium having information recorded thereon for speech recognition, wherein the information, when read by the machine, causes the machine to perform:

receiving an audio signal representing a speech of a user engaged in a dialogue and a visual signal capturing the user uttering the speech;

obtaining a first speech recognition result by

identifying phonemes from the audio signal,

performing audio based speech recognition based on the phonemes to generate the first speech recognition result, and

converting the phonemes into a first set of visemes;

obtaining a second speech recognition result by

identifying a second set of visemes based on at least one of lip shape and movement of the user detected from the visual signal, and

performing lip reading based speech recognition to generate the second speech recognition result based on the second set of visemes; and

generating an integrated speech recognition result based on the first and second set of visemes, wherein

the first speech recognition result is selected as the integrated speech recognition result if a similarity between the first and second speech recognition results satisfies a first criterion, and

the second speech recognition result is selected as the integrated speech recognition result if the similarity satisfies a second criterion.

7. The medium of claim 6 , wherein

the first speech recognition result includes a first plurality of words obtained via the phonemes; and

the second speech recognition result includes a second plurality of words obtained via the second set of visemes.

8. The medium of claim 6 , further comprising synchronizing the first speech recognition result and the second speech recognition result, wherein each of the phonemes in the first speech recognition result is synchronized with a corresponding viseme in the second set of visemes of the second speech recognition result.

9. The medium of claim 8 , wherein the similarity between the first and second speech recognition results is determined by:

determining, pairs of visemes, each of which includes one from the first set of visemes converted from the phonemes and a corresponding one from the second set of visemes that is synchronized with a corresponding phoneme from which the one viseme in the pair is converted;

assessing the similarity between the first and second speech recognition results based on the pairs of visemes.

10. The medium of claim 6 , wherein the integrated speech recognition is given a confidence score, determined based on whether the first or second speech recognition result is selected.

11. A system for speech recognition, comprising:

a sensor data collection unit configured for receiving an audio signal representing a speech of a user engaged in a dialogue and a visual signal capturing the user uttering the speech;

an audio based speech recognition unit configured for obtaining a first speech recognition result by

identifying phonemes from the audio signal,

performing audio based speech recognition based on the phonemes to generate the first speech recognition result, and

converting the phonemes into a first set of visemes;

a lip reading based speech recognizer configured for obtaining a second speech recognition result by

identifying a second set of visemes based on at least one of lip shape and movement of the user detected from the visual signal, and

performing lip reading based speech recognition to generate the second speech recognition result based on the second set of visemes; and

an audio-visual speech recognition integrator configured for generating an integrated speech recognition result based on the first and second set of visemes, wherein

the first speech recognition result is selected as the integrated speech recognition result if a similarity between the first and second speech recognition results satisfies a first criterion, and

the second speech recognition result is selected as the integrated speech recognition result if the similarity satisfies a second criterion.

12. The system of claim 11 , wherein

the first speech recognition result includes a first plurality of words obtained via the phonemes; and

the second speech recognition result includes a second plurality of words obtained via the second set of visemes.

13. The system of claim 11 , further comprising a synchronization unit configured for synchronizing the first speech recognition result and the second speech recognition result, wherein each of the phonemes in the first speech recognition result is synchronized with a corresponding viseme in the second set of visemes of the second speech recognition result.

14. The system of claim 13 , wherein the audio-visual speech recognition integrator is configured for:

determining, pairs of visemes, each of which includes one from the first set of visemes converted from the phonemes and a corresponding one from the second set of visemes that is synchronized with a corresponding phoneme from which the one viseme in the pair is converted;

assessing the similarity between the first and second speech recognition results based on the pairs of visemes.

15. The system of claim 11 , wherein the integrated speech recognition is given a confidence score, determined based on whether the first or second speech recognition result is selected.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2023
From: DMAL, INC.
To: DMAI (GUANGZHOU) CO.,LTD.
Reel/Frame 065489/0038 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2019
From: SHUKLA, NISHANT; DHARNE, ASHWIN
To: DMAI, INC.
Reel/Frame 048345/0448 →
Continuity (2)
Provisional Application 62630976 · Feb 15, 2018
Related Publication 20190279642A1 · Sep 12, 2019