IP Library Granted Patent US 12,087,319
Granted Patent B1
US 12,087,319 · App. 17/079,082 · Granted Sep 10, 2024

Joint estimation of acoustic parameters from single-microphone speech

Inventors: David Looney (Atlanta, GA); Nikolay Gaubitch (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L25/30G06N3/048G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,087,319
App. No.
17/079,082
Granted
Sep 10, 2024
Kind
B1
Abstract

Embodiments described herein provide for end-to-end joint determination of degradation parameter scores for certain types of degradation. Degradation parameters include degradation describing additive noise and multiplicative noise such as Signal-to-Noise Ratio (SNR), reverberation time (T60), and Direct-to-Reverberant Ratio (DRR). Various neural network architectures are described such that the inherent interplay between the degradation parameters is considered in both the degradation parameter score and degradation score determination. The neural network architectures are trained according to computer generated audio datasets.

Claims (35)

1. A computer-implemented method for end-to-end acoustic degradation estimation from an audio signal comprising:

training, by a computer, a neural network architecture by applying the neural network architecture on a plurality of simulated audio signals having one or more types of degradation, the neural network architecture comprising a plurality of sets of one or more layers to generate a corresponding degradation parameter score;

receiving, by the computer, an input audio signal originating from a speaker; and

generating, by the computer, a plurality of degradation parameter scores for the input audio signal based upon a plurality of degradation parameters corresponding to a type of degradation by applying the neural network architecture to the input audio signal,

wherein a first degradation parameter score of the plurality of degradation parameter scores is output by a first set of one or more layers of the neural network architecture and fed to a second set of one or more layers of the neural network architecture to generate a second degradation parameter score of the plurality of degradation parameter scores.

2. The method according to claim 1 , wherein training the neural network architecture further comprises applying the neural network architecture on one or more clean audio signals.

3. The method according to claim 1 , further comprising generating, by the computer, the plurality of simulated audio signals using a clean audio signal by applying at least one type of degradation on the clean audio signal.

4. The method according to claim 1 , wherein the neural network architecture is trained according to the plurality of simulated audio signals and corresponding labels, each respective label indicating a degradation parameter for the type of degradation applied to produce the respective simulated audio signal.

5. The method according to claim 1 , further comprising determining, by the computer executing the neural network architecture, a degradation score based upon the plurality of degradation parameter scores.

6. The method according to claim 1 , further comprising determining, by the computer, that an inbound call originated from a registered caller in response to determining that the plurality of degradation parameter scores satisfy one or more thresholds.

7. The method according to claim 1 , further comprising:

determining, by the computer, a confidence value for the input audio signal based upon the plurality of degradation parameter scores; and

executing, by the computer, a downstream operation based upon the confidence value.

8. The method according to claim 1 , further comprising determining, by the computer, a degradation mitigation action based upon the plurality of degradation parameter scores.

9. The method according to claim 1 , wherein the neural network architecture comprises a plurality of estimators comprising a corresponding set of one or more layers for generating a degradation score for the type of degradation, and

wherein the neural network architecture generates the plurality of degradation parameter scores by successively applying each estimator, and wherein a second estimator uses as input a degradation parameter score generated from a first estimator.

10. The method according to claim 1 , wherein a type of degradation parameter includes at least one of: a reverberation time, an early to late reverberation ratio, and a signal-to-noise ratio.

11. A system comprising:

a computing device comprising a processor and a non-transitory storage medium configured to store a plurality of computer program instructions that when executed by the processor:

train a neural network architecture by applying the neural network architecture on a plurality of simulated audio signals having one or more types of degradation, the neural network architecture comprising a plurality of sets of one or more layers trained to generate a corresponding degradation parameter score;

receive an input audio signal originating from a speaker; and

generate a plurality of degradation parameter scores for the input audio signal based upon a plurality of degradation parameters corresponding to a type of degradation by applying the neural network architecture to the input audio signal,

wherein a first degradation parameter score of the plurality of degradation parameter scores is output by a first set of one or more layers of the neural network architecture and fed to a second set of one or more layers of the neural network architecture to generate a second degradation parameter score of the plurality of degradation parameter scores.

12. The system according to claim 11 , wherein the neural network architecture is further trained by applying the neural network architecture on one or more clean audio signals.

13. The system according to claim 11 , wherein the processor of the computing device is configured to generate the plurality of simulated audio signals using a clean audio signal by applying at least one type of degradation on the clean audio signal.

14. The system according to claim 11 , wherein the neural network architecture is trained according to the plurality of simulated audio signals and corresponding labels, each respective label indicating a degradation parameter for the type of degradation applied to produce the respective simulated audio signal.

15. The system according to claim 11 , wherein the processor of the computing device is configured to determine a degradation score based upon the plurality of degradation parameter scores.

16. The system according to claim 11 , wherein the processor of the computing device is configured to determine that an inbound call originated from a registered caller in response to determining that the plurality of degradation parameter scores satisfy one or more thresholds.

17. The system according to claim 11 , wherein the processor of the computing device is configured to:

determine a confidence value for the input audio signal based upon the plurality of degradation parameter scores; and

execute a downstream operation based upon the confidence value.

18. The system according to claim 11 , wherein the processor of the computing device is configured to determine a degradation mitigation action based upon the plurality of degradation parameter scores.

19. The system according to claim 11 , wherein the neural network architecture comprises a plurality of estimators comprising a corresponding set of one or more layers for generating a degradation score for the type of degradation, and

wherein the neural network architecture generates the plurality of degradation parameter scores by successively applying each estimator, and wherein a second estimator uses as input a degradation parameter score generated from a first estimator.

20. The system according to claim 11 , wherein the computing device is at least one of an Internet of Things device, a server, and a caller device.

Assignments (4)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
RELEASE OF SECURITY INTEREST Recorded Jun 26, 2024
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: PINDROP SECURITY, INC.
Reel/Frame 069477/0962 →
SECURITY INTEREST Recorded Jul 31, 2023
From: PINDROP SECURITY, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064443/0584 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2020
From: LOONEY, DAVID; GAUBITCH, NIKOLAY D.
To: PINDROP SECURITY, INC.
Reel/Frame 054154/0053 →
Continuity (1)
Provisional Application 62925349 · Oct 24, 2019
Cited By (6)
US 12,444,423 US 12,592,220 US 12,592,239 US 12,676,137 US 12,676,144 US 12,706,098