IP Library › Granted Patent US 11,398,239
Granted Patent B1
US 11,398,239 · App. 17/109,445 · Granted Jul 26, 2022

ASR-enhanced speech compression

Inventor: David Garrod (Leadville, CO)
Assignee: Medallia, Inc.
G10L19/04G06F17/18G06F40/279G10L15/18H03M7/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,398,239
App. No.
17/109,445
Granted
Jul 26, 2022
Kind
B1
Abstract

A process for compressing an audio speech signal utilizes ASR processing to generate a corresponding text representation and, depending on confidence in the corresponding text representation, selectively applies more, less, or no compression to the audio signal. The result is a compressed audio signal, with corresponding text, that is compact and well suited for searching, analytics, or additional ASR processing.

Claims (40)

1. A computer-implemented process for digitally compressing an audio signal that includes human speech utterances, the process comprising at least the following computer-implemented acts:

(1) receiving a digitally encoded audio signal that includes human speech utterances in a first, uncompressed format;

(2) identifying portions of the digitally encoded audio signal that correspond to speech utterances and, for each speech utterance, forming a corresponding uncompressed audio utterance;

(3) performing automatic speech recognition (ASR) processing to produce, for each speech utterance, at least a corresponding (i) text representation and (ii) ASR confidence that represents a likelihood that the text representation accurately captures all spoken words contained in its corresponding uncompressed audio utterance;

(4) for each speech utterance, if its ASR confidence exceeds a predetermined threshold value, then forming a corresponding compressed audio utterance in a highly compressed format;

(5) forming an output stream that includes, for each speech utterance, at least:

(i) its corresponding text representation;

(ii) its corresponding ASR confidence; and

(iii) either (a) its corresponding uncompressed audio utterance or (b) its corresponding compressed audio utterance, but not both (a) and (b), wherein the output stream contains (a) if the utterance's corresponding ASR confidence is less than or equal to the predetermined threshold value and (b) if the utterance's corresponding ASR confidence exceeds the predetermined threshold value.

2. A process, as defined in claim 1 , wherein for each speech utterance, the output stream further includes metadata computed from the corresponding uncompressed audio utterance.

3. A process, as defined in claim 1 , wherein the metadata includes one or more of: identity of the speaker, gender, approximate age, and/or emotion.

4. A process, as defined in claim 1 , wherein the ASR confidence values are derived from normalized likelihood scores.

5. A process, as defined in claim 1 , wherein the ASR confidence values are computed using an N-best homogeneity analysis.

6. A process, as defined in claim 1 , wherein the ASR confidence values are computed using an acoustic stability analysis.

7. A process, as defined in claim 1 , wherein the ASR confidence values are computed using a word graph hypothesis density analysis.

8. A process, as defined in claim 1 , wherein the ASR confidence values are derived from associated state, phoneme, or word durations.

9. A process, as defined in claim 1 , wherein the ASR confidence values are derived from language model (LM) scores or LM back-off behaviors.

10. A process, as defined in claim 1 , wherein the ASR confidence values are computed using a posterior probability analysis.

11. A process, as defined in claim 1 , wherein the ASR confidence values are computed using a log-likelihood-ratio analysis.

12. A process, as defined in claim 1 , wherein the ASR confidence values are computed using a neural net that includes word identity and aggregated words as predictors.

13. A computer-implemented process for digitally compressing an audio signal that includes human speech utterances, the process comprising at least the following computer-implemented acts:

(1) receiving a digitally encoded audio signal that includes human speech utterances in a first, lightly compressed format;

(2) identifying portions of the digitally encoded audio signal that correspond to speech utterances and, for each speech utterance, forming a corresponding lightly compressed audio utterance;

(3) performing automatic speech recognition (ASR) processing to produce, for each speech utterance, at least a corresponding (i) text representation and (ii) ASR confidence that represents a likelihood that the text representation accurately captures all spoken words contained in its corresponding lightly compressed audio utterance;

(4) for each speech utterance, if its ASR confidence exceeds a predetermined threshold value, then forming a corresponding heavily compressed audio utterance in a highly compressed format;

(5) forming an output stream that includes, for each speech utterance, at least:

(i) its corresponding text representation;

(ii) its corresponding ASR confidence; and

(iii) either (a) its corresponding lightly compressed audio utterance or (b) its corresponding heavily compressed audio utterance, but not both (a) and (b), wherein the output stream contains (a) if the utterance's corresponding ASR confidence is less than or equal to the predetermined threshold value and (b) if the utterance's corresponding ASR confidence exceeds the predetermined threshold value.

14. A computer-implemented process for digitally compressing an audio signal that includes human speech utterances, the process comprising at least the following computer-implemented acts:

(1) receiving a digitally encoded audio signal that includes human speech utterances in a first, lightly compressed format;

(2) identifying portions of the digitally encoded audio signal that correspond to speech utterances and, for each speech utterance, forming a corresponding lightly compressed audio utterance;

(3) performing automatic speech recognition (ASR) processing to produce, for each speech utterance, at least a corresponding (i) text representation and (ii) ASR confidence that represents a likelihood that the text representation accurately captures all spoken words contained in its corresponding lightly compressed audio utterance;

(4) for each speech utterance, if its ASR confidence exceeds a predetermined threshold value, then forming a corresponding heavily compressed audio utterance in a highly compressed format;

(5) forming an output stream that, for each speech utterance, consists essentially of:

(i) its corresponding text representation;

(ii) its corresponding ASR confidence; and

(iii) either (a) its corresponding lightly compressed audio utterance or (b) its corresponding heavily compressed audio utterance, but not both (a) and (b), wherein the output stream contains (a) if the utterance's corresponding ASR confidence is less than or equal to the predetermined threshold value and (b) if the utterance's corresponding ASR confidence exceeds the predetermined threshold value.

15. A process, as defined in claim 1 , wherein for each speech utterance, the output stream further consists essentially of metadata computed from the corresponding lightly compressed audio utterance.

16. A process, as defined in claim 1 , wherein the metadata includes one or more of: identity of the speaker, gender, approximate age, and/or emotion.

Assignments (2)
SECURITY INTEREST IN PATENT RIGHTS Recorded Jul 30, 2026
From: MEDALLIA, INC., AS THE GRANTOR
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS THE COLLATERAL AGENT
Reel/Frame 076083/0290 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2020
From: GARROD, DAVID, DR.
To: MEDALLIA, INC.
Reel/Frame 054556/0165 →
Continuity (1)
Continuation 16371014 · Mar 31, 2019
Cited By (1)
US 12,608,427