IP Library Granted Patent US 7,319,955
Granted Patent B2
US 7,319,955 · App. 10/307,164 · Granted Jan 15, 2008

Audio-visual codebook dependent cepstral normalization

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,319,955
App. No.
10/307,164
Granted
Jan 15, 2008
Kind
B2
Abstract

An arrangement for yielding enhanced audio features towards the provision of enhanced audio-visual features for speech recognition. Input is provided in the form of noisy audio-visual features and noisy audio features related to the noisy audio-visual features.

Claims (34)

1. An apparatus for enhancing speech for speech recognition, said apparatus comprising:

a first input medium which obtains noisy audio-visual features;

a second input medium which obtains noisy audio features related to the noisy audio-visual features; and

a cepstral speech function output arrangement for combining the first and second inputs to yield enhanced audio features that are re-combined with visual features to yield enhanced audio-visual features used for speech recognition.

2. The apparatus according to claim 1 , wherein said arrangement for yielding enhanced audio features is adapted to yield estimated clean speech features.

3. The apparatus according to claim 1 , wherein said arrangement for yielding enhanced audio features comprises an arrangement for determining a posterior distribution based on the noisy audio-visual features.

4. The apparatus according to claim 3 , wherein said arrangement for determining a posterior distribution is adapted to determine a posterior distribution based additionally on Gaussian parameters which model a probability density function related to the noisy audio-visual features.

5. The apparatus according to claim 1 , wherein said arrangement for yielding enhanced audio features comprises an arrangement for estimating audio compensation codewords.

6. The apparatus according to claim 1 , wherein said arrangement for yielding enhanced audio features comprises an arrangement for determining the difference between the noisy audio features and modified noisy audio-visual features.

7. The apparatus according to claim 1 , wherein said arrangement for yielding enhanced audio features comprises:

an arrangement for determining a posterior distribution based on the noisy audio-visual features; and

an arrangement for estimating audio compensation codewords.

8. The apparatus according to claim 7 , wherein said arrangement for yielding enhanced audio features comprises an arrangement for effecting a multiplication of the posterior distribution with the estimated audio compensation codewords.

9. The apparatus according to claim 8 , wherein said arrangement for yielding enhanced audio features comprises an arrangement for determining the difference between the noisy audio features and the multiplication of the posterior distribution with the estimated audio compensation codewords.

10. The apparatus according to claim 1 , wherein said first input medium is adapted to accept noisy audio-visual features which have resulted from the processing of normalized audio features and normalized videofeatures.

11. A method of enhancing speech for speech recognition, said method comprisingthe steps of:

obtaining noisy audio-visual features;

obtaining noisy audio features related to the noisy audio-visual features; and

using a cepstral speech function operating on the noisy audio features and the noisy audio-visual features to yield enhanced audio features that are re-combined with visual features to yield enhanced audio-visual features used for speech recognition.

12. The method according to claim 11 , wherein said step of yielding enhanced audio features comprises yielding estimated clean speech features.

13. The method according to claim 11 , wherein said step of yielding enhanced audio features comprises determining a posterior distribution based on the noisy audio-visual features.

14. The method according to claim 13 , wherein said step of determining a posterior distribution comprises determining a posterior distribution based additionally on Gaussian parameters which model a probability density function related to the noisy audio-visual features.

15. The method according to claim 11 , wherein said step of yielding enhanced audio features comprises estimating audio compensation codewords.

16. The method according to claim 11 , wherein said step of yielding enhanced audio features comprises determining the difference between the noisy audio features and modified noisy audio-visual features.

17. The method according to claim 11 , wherein said step of yielding enhanced audio features comprises:

determining a posterior distribution based on the noisy audio-visual features; and

estimating audio compensation codewords.

18. The method according to claim 17 , wherein said step of yielding enhanced audio features comprises effecting a multiplication of the posterior distribution with the estimated audio compensation codewords.

19. The method according to claim 18 , wherein said step of yielding enhanced audio features comprises determining the difference between the noisy audio features and the multiplication of the posterior distribution with the estimated audio compensation codewords.

20. The method according to claim 11 , wherein said step of obtaining noisy audio-visual features comprises obtaining noisy audio-visual features which have resulted from the processing of normalized audio features and normalized video features.

21. A program storage device readable by machine, tangibly embodying a program of instructions executable by the machine to perform method steps for enhancing speech for speech recognition, said method comprising the steps of:

obtaining noisy audio-visual features;

obtaining noisy audio features related to the noisy audio-visual features; and

using a cepstral speech function operating on the noisy audio features and the noisy audio-visual features to yield enhanced audio features that are re-combined with visual features to yield enhanced audio-visual features used for speech recognition.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022354/0566 →