IP Library Granted Patent US 7,856,355
Granted Patent B2
US 7,856,355 · App. 11/172,965 · Granted Dec 21, 2010

Speech quality assessment method and system

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,856,355
App. No.
11/172,965
Granted
Dec 21, 2010
Kind
B2
Abstract

In one embodiment, distortion in a received speech signal is estimated using at least one model trained based on subjective quality assessment data. A speech quality assessment for the received speech signal is then determined based on the estimated distortion.

Claims (44)

1. A speech quality assessment method, comprising:

estimating, by a computer, objective frame distortion in a received speech signal at a portion of a speech quality system, the estimated objective frame distortion based on inputting at least one feature vector for each frame of the received speech signal into at least a first model trained based on subjective quality assessment data, the at least one feature vector including articulation power components;

estimating mute distortion caused by mutes in the received speech signal using a second model trained based on the subjective quality assessment data;

determining a speech quality assessment for the received speech signal based on the estimated objective frame distortion and the estimated mute distortion; and

storing the determined speech quality assessment at a portion of the speech quality system,

retraining the first model based on the estimated objective frame distortion and the estimated mute distortion.

2. The method of claim 1 , wherein the objective frame distortion estimating step, further includes:

estimating objective speech distortion in the received speech signal using the first model trained based on subjective quality assessment data.

3. The method of claim 2 , wherein the objective frame distortion estimating step, further includes:

estimating objective background noise distortion in the received speech signal using the first model trained based on subjective quality assessment data.

4. The method of claim 3 , further comprising:

determining an average articulation power and an average non-articulation power from the received speech signal; and wherein

the estimating objective speech distortion step estimates the speech distortion using the determined average articulation power, the determined average non-articulation power and the first model; and

the estimating objective background noise distortion step estimates the background noise distortion using the determined average articulation power, the determined average non-articulation power and the first model.

5. The method of claim 3 , wherein the first model models subjective determination of distortion in speech signals.

6. The method of claim 5 , wherein the first model is a neural network.

7. The method of claim 1 , wherein the estimating mute distortion step, further includes:

detecting mutes in the received speech signal; and

estimating distortion caused by the detected mutes.

8. The method of claim 7 , wherein,

the detecting step detects locations and durations of mutes in the received speech signal; and

the estimating distortion caused by the detected mutes step estimates the mute distortion based on the detected locations and durations.

9. The method of claim 1 , wherein the estimating distortion caused by the detected mutes step estimates the mute distortion such that mutes later in the received speech signal have a greater impact than mutes earlier in the received speech signal.

10. The method of claim 1 , wherein the first model models subjective determination of distortion in speech signals lacking mute distortion and the second model models subjective determination of mute distortion in speech signals.

11. The method of claim 1 , wherein the determining step maps the estimated objective frame distortion to a subjective quality assessment metric.

12. A processing device for speech quality assessment, comprising:

at least one processor to estimate objective frame distortion in a received speech signal, the estimated objective frame distortion based on inputting at least one feature vector for each frame of the received speech signal into at least a first model trained based on subjective quality assessment data, the at least one feature vector including articulation power components, the at least one processor to further estimate mute distortion in the received speech signal using a second model trained based on the subjective quality assessment data; and

a mapping unit mapping the estimated objective frame distortion to a speech quality metric,

wherein the first model is retrained based on the estimated objective frame distortion and the estimated mute distortion.

13. A method of estimating objective frame distortion, comprising:

receiving, by a computer, at least one feature vector for each frame of a received signal at a portion of a speech quality system, the at least one feature vector including articulation power components;

detecting mutes in the received signal;

estimating speech distortion in the received signal using the at least one feature vector for each frame and a first model trained based on subjective quality assessment data;

estimating background noise distortion in the received signal using the at least one feature vector for each frame and the first model trained based on subjective quality assessment data;

estimating mute distortion caused by the detected mutes such that mutes later in the received speech signal have a greater impact than mutes earlier in the received speech signal, using a second model trained based on the subjective quality assessment data;

combining the estimated speech distortion and the estimated background noise distortion to obtain an objective frame distortion estimate;

storing the objective frame distortion estimate at a receiver; and

storing the estimated mute distortion at the receiver,

retraining the first model based on the estimated objective frame distortion and the estimated mute distortion.

14. A method of training a speech quality assessment system, comprising:

training, by a computer, a first distortion estimation path of the system while excluding impact from a second distortion estimation path of the system using first subjective quality assessment data, the first subjective quality assessment data including first speech signals and first associated subjective quality metrics, the first speech signals lacking in mute distortion;

training the second distortion estimating path of the system using second subjective quality assessment data, the second subjective quality assessment data including second speech signals and second associated subjective quality metrics, the second speech signals including mute distortion,

wherein the first distortion estimating path is configured to estimate frame distortion and the second distortion estimating path is configured to estimate mute distortion; and

retraining the first distortion path while including the impact of the second distortion path using the first and second quality assessment data.

Assignments (9)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2021
From: PROVENANCE ASSET GROUP LLC
To: RPX CORPORATION
Reel/Frame 059352/0001 →
RELEASE OF SECURITY INTEREST Recorded Nov 30, 2021
From: NOKIA US HOLDINGS INC.
To: PROVENANCE ASSET GROUP HOLDINGS LLC; PROVENANCE ASSET GROUP LLC
Reel/Frame 058363/0723 →
RELEASE OF SECURITY INTEREST Recorded Nov 30, 2021
From: CORTLAND CAPITAL MARKETS SERVICES LLC
To: PROVENANCE ASSET GROUP HOLDINGS LLC; PROVENANCE ASSET GROUP LLC
Reel/Frame 058983/0104 →
ASSIGNMENT AND ASSUMPTION AGREEMENT Recorded Feb 14, 2019
From: NOKIA USA INC.
To: NOKIA US HOLDINGS INC.
Reel/Frame 048370/0682 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2017
From: NOKIA TECHNOLOGIES OY; NOKIA SOLUTIONS AND NETWORKS BV; ALCATEL LUCENT SAS
To: PROVENANCE ASSET GROUP LLC
Reel/Frame 043877/0001 →
SECURITY INTEREST Recorded Sep 13, 2017
From: PROVENANCE ASSET GROUP HOLDINGS, LLC; PROVENANCE ASSET GROUP LLC
To: NOKIA USA INC.
Reel/Frame 043879/0001 →
SECURITY INTEREST Recorded Sep 13, 2017
From: PROVENANCE ASSET GROUP HOLDINGS, LLC; PROVENANCE ASSET GROUP, LLC
To: CORTLAND CAPITAL MARKET SERVICES, LLC
Reel/Frame 043967/0001 →
MERGER Recorded Nov 2, 2010
From: LUCENT TECHNOLOGIES INC.
To: ALCATEL-LUCENT USA INC.
Reel/Frame 025230/0835 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2005
From: KIM, DOH-SUK
To: LUCENT TECHNOLOGIES INC.
Reel/Frame 016843/0760 →