IP Library Granted Patent US 8,635,066
Granted Patent B2
US 8,635,066 · App. 12/759,907 · Granted Jan 21, 2014

Camera-assisted noise cancellation and speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,635,066
App. No.
12/759,907
Granted
Jan 21, 2014
Kind
B2
Abstract

Methods, system, and articles are described herein for receiving an audio input and a facial image sequence for a period of time, in which the audio input includes speech input from multiple speakers. The audio input is extracted based on the received facial image sequence to extract a speech input of a particular speaker.

Claims (66)

1. A computer-implemented method, comprising:

receiving an audio input and a facial image sequence for a period of time at an electronic device, the audio input including a decipherable portion and an indecipherable portion;

converting the decipherable portion of the audio input into a first symbol sequence portion, the converting including:

detecting separations between spoken words in the decipherable segment based on the facial image sequence, and

processing the decipherable portion into the first symbol sequence portion based in part on the detected separations;

processing a portion of the facial image sequence that corresponds temporally to the indecipherable portion of the audio input into a second symbol sequence portion; and

integrating the first symbol sequence portion and the second symbol sequence portion in temporal order to form a symbol sequence.

2. The computer-implemented method of claim 1 , further comprising:

identifying portions of the facial image sequence that indicate a particular speaker is silent; and

filtering out portions of the audio input that correspond to the portions of the facial image sequence that indicate the particular speaker is silent.

3. The computer-implemented method of claim 2 , wherein the identifying includes identifying portions of the facial image sequence that indicate the particular speaker is silent based on facial features of the particular speaker shown in the portions of the facial image sequence.

4. The computer-implemented method of claim 1 , further comprising transmitting the symbol sequence to another electronic device or storing the symbol sequence in a data storage on the electronic device.

5. The computer-implemented method of claim 1 , wherein the receiving includes receiving the audio input via a microphone of an electronic device and receiving the facial image sequence via a camera of the electronic device.

6. The computer-implemented method of claim 1 , further comprising determining that speech input is initiated when the facial image sequence indicates that a particular speaker begins to utter sounds.

7. The computer-implemented method of claim 6 , further comprising determining that the speech input is terminated when a facial image of the particular speaker moves out of the view of a camera of the electronic device or when the facial image sequence indicates that the particular speaker has not uttered sounds for a predetermined period of time.

8. A computer-implemented method, comprising:

receiving an audio input and a facial image sequence for a period of time at an electronic device, the audio input including speech inputs from multiple speakers;

extracting a speech input of a particular speaker from the audio input based on the received facial image sequence, including selecting portions of the audio input that correspond to speech movement indicated by the facial image sequence, at least one selected portion including a decipherable segment and an indecipherable segment; and

processing the extracted speech input into a symbol sequence, the processing including:

converting the decipherable segment into a first sub symbol sequence portion, the converting including:

detecting separations between spoken words in the decipherable segment based on the facial image sequence, and

processing the decipherable portion into the first sub symbol sequence portion based in part on the detected separations;

processing a portion of the facial image sequence that corresponds temporally to the indecipherable segment into a second sub symbol sequence portion; and

integrating the first sub symbol sequence portion and the second sub symbol sequence portion in temporal order to form one of the corresponding symbol sequence portions.

9. The computer-implemented method of claim 8 , further comprising converting the symbol sequence into text for display by the electronic device.

10. The computer-implemented method of claim 8 , furthering comprising matching the symbol sequence to a command that causes the electronic device to perform a function.

11. The computer-implemented method of claim 8 , wherein the receiving includes receiving the audio input via a microphone of an electronic device and receiving the facial image sequence via a camera of the electronic device.

12. The computer-implemented method of claim 8 , wherein the processing includes:

converting each of the selected audio input portions into a corresponding symbol sequence portion; and

assembling the symbol sequence portions into a symbol sequence.

13. The computer-implemented method of claim 8 , wherein the processing includes:

obtaining a first symbol sequence portion and a corresponding audio transformation confidence score for an audio input portion;

obtaining a second symbol sequence portion and a corresponding visual transformation confidence score for each facial image sequence that corresponds to the audio input portion;

comparing the audio transformation confidence score and the visual transformation confidence score of the audio portion;

selecting the first symbol sequence portion for assembly into the symbol sequence when the audio transformation confidence score is higher than the visual transformation confidence score; and

selecting the second symbol sequence portion for assembly into the symbol when the visual transformation confidence score is higher than the audio transformation confidence score.

14. The computer-implemented method of claim 13 , further comprising selecting the first symbol sequence portion for assembly into the symbol sequence when the audio transformation confidence score is equal to the visual transformation confidence score and the audio transformation confidence score is equal to or higher than a predetermined audio confidence score threshold.

15. The computer-implemented method of claim 13 , further comprising selecting the first symbol sequence portion or the second symbol sequence portion for assembly into the symbol sequence when the audio transformation confidence score is equal to the visual transformation confidence score.

16. The computer-implemented method of claim 8 , wherein the indecipherable segment is masked by ambient noise or an audio input portion in which data is missing.

17. The computer-implemented method of claim 8 , further comprising determining that a segment of one of the selected portions is an indecipherable segment when a transformation confidence score of a sub symbol sequence portion obtained from the segment is below a predetermined audio confidence score threshold.

18. The computer-implemented method of claim 8 , further comprising determining that the audio input is initiated when the facial image sequence indicates that a speaker begins to utter sounds.

19. The computer-implemented method of claim 8 , further comprising determining that the audio input is terminated when a facial image of a speaker moves out of the view of a camera of the electronic device or when the facial image sequence indicates that the speaker has not uttered sounds for a predetermined period of time.

20. The computer-implemented method of claim 8 , wherein the audio input contains one or more phonemes and the facial image sequence includes one or more visemes.

21. An article of manufacture comprising:

a storage medium; and

computer-readable programming instructions stored on the storage medium and configured to program a computing device to perform operations including:

receiving an audio input and a facial image sequence for a period of time at an electronic device, wherein the audio input includes a decipherable portion and an indecipherable portion;

converting the decipherable portion of the audio input into a first symbol sequence portion, the converting including:

detecting separations between spoken words in the decipherable segment based on the facial image sequence, and

processing the decipherable portion into the first symbol sequence portion based in part on the detected separations;

processing a portion of the facial image sequence that corresponds temporally to the indecipherable portion of the audio input into a second symbol sequence portion; and

integrating the first symbol sequence portion and the second symbol sequence portion in temporal order to form a symbol sequence.

22. The article of claim 21 , wherein the operations further include converting the symbol sequence into text for display by the electronic device.

23. The article of claim 21 , wherein the operations further includes determining that a portion of the audio input is an indecipherable portion when a transformation confidence score of a symbol sequence obtained from the portion is below a predetermined audio confidence score threshold.

24. The article of claim 21 , wherein the operations further include further determining that the audio input is initiated when the facial image sequences indicate that a speaker begins to utter sounds, and determining that the audio input is terminated when a facial image of a speaker moves out of the view of a camera of the electronic device or when the facial image sequence indicates that the speaker has not uttered sounds for a predetermined period of time.

25. The article of claim 21 , wherein the receiving includes receiving the audio input via a microphone of an electronic device and receiving the facial image sequence via a camera of the electronic device.

26. A device comprising:

a microphone to receive an audio input from an environment;

a camera to receive a plurality of facial image sequences, the camera configured to automatically track a face of a speaker associated with the facial image sequences to maintain its view of the face;

a processor;

a memory that stores a plurality of modules that comprises:

a visual interpretation module to process a portion of an facial image sequence into a symbol sequence, wherein the facial image sequence corresponds temporally to an indecipherable portion of an audio input;

a speech recognition module to convert a decipherable portion of the audio input into another symbol sequence, and to integrate the symbol sequences in temporal order to form an integrated symbol sequence; and

a command module to cause the device to perform a function in response at least to the symbol sequence.

27. The device of claim 26 , wherein the speech recognition module is to further convert the symbol sequence into text for display on the device.

28. The device of claim 26 , wherein the command module is to further cause the device to perform a function in response to the integrated symbol sequence.

Assignments (7)
RELEASE OF SECURITY INTEREST Recorded Aug 23, 2022
From: DEUTSCHE BANK TRUST COMPANY AMERICAS
To: IBSV LLC; LAYER3 TV, LLC; PUSHSPRING, LLC; T-MOBILE CENTRAL LLC; T-MOBILE USA, INC.; ASSURANCE WIRELESS USA, L.P.; BOOST WORLDWIDE, LLC; CLEARWIRE COMMUNICATIONS LLC; CLEARWIRE IP HOLDINGS LLC; SPRINTCOM LLC; SPRINT COMMUNICATIONS COMPANY L.P.; SPRINT INTERNATIONAL INCORPORATED; SPRINT SPECTRUM LLC
Reel/Frame 062595/0001 →
SECURITY AGREEMENT Recorded Apr 2, 2020
From: T-MOBILE USA, INC.; ISBV LLC; T-MOBILE CENTRAL LLC; LAYER3 TV, INC.; PUSHSPRING, INC.; BOOST WORLDWIDE, LLC; CLEARWIRE COMMUNICATIONS LLC; CLEARWIRE IP HOLDINGS LLC; CLEARWIRE LEGACY LLC; SPRINT COMMUNICATIONS COMPANY L.P.; SPRINT INTERNATIONAL INCORPORATED; SPRINT SPECTRUM L.P.; ASSURANCE WIRELESS USA, L.P.
To: DEUTSCHE BANK TRUST COMPANY AMERICAS
Reel/Frame 053182/0001 →
RELEASE OF SECURITY INTEREST Recorded Apr 1, 2020
From: DEUTSCHE TELEKOM AG
To: T-MOBILE USA, INC.; IBSV LLC
Reel/Frame 052969/0381 →
RELEASE OF SECURITY INTEREST Recorded Apr 1, 2020
From: DEUTSCHE BANK AG NEW YORK BRANCH
To: T-MOBILE USA, INC.; IBSV LLC; METROPCS COMMUNICATIONS, INC.; METROPCS WIRELESS, INC.; T-MOBILE SUBSIDIARY IV CORPORATION; LAYER3 TV, INC.; PUSHSPRING, INC.
Reel/Frame 052969/0314 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 30, 2016
From: T-MOBILE USA, INC.
To: DEUTSCHE TELEKOM AG
Reel/Frame 041225/0910 →
SECURITY AGREEMENT Recorded Nov 17, 2015
From: T-MOBILE USA, INC.; METROPCS COMMUNICATIONS, INC.; T-MOBILE SUBSIDIARY IV CORPORATION
To: DEUTSCHE BANK AG NEW YORK BRANCH, AS ADMINISTRATIVE AGENT
Reel/Frame 037125/0885 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2010
From: MORRISON, ANDREW R.
To: T-MOBILE USA, INC.
Reel/Frame 024230/0819 →