IP Library Granted Patent US 12,190,905
Granted Patent B2
US 12,190,905 · App. 17/408,281 · Granted Jan 7, 2025

Speaker recognition with quality indicators

Inventors: Hrishikesh Rao (Atlanta, GA); Kedar Phatak (Atlanta, GA); Elie Khoury (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L25/60G06N20/20G10L15/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,905
App. No.
17/408,281
Granted
Jan 7, 2025
Kind
B2
Abstract

Embodiments described herein provide for a machine-learning architecture for modeling quality measures for enrollment signals. Modeling these enrollment signals enables the machine-learning architecture to identify deviations from expected or ideal enrollment signal in future test phase calls. These differences can be used to generate quality measures for the various audio descriptors or characteristics of audio signals. The quality measures can then be fused at the score-level with the speaker recognition's embedding comparisons for verifying the speaker. Fusing the quality measures with the similarity scoring essentially calibrates the speaker recognition's outputs based on the realities of what is actually expected for the enrolled caller and what was actually observed for the current inbound caller.

Claims (53)

1. A computer-implemented method comprising:

extracting from an inbound audio signal for an inbound speaker, by a computer, a feature vector for one or more acoustic features;

generating, by the computer, one or more quality measures and an overall quality measure for the inbound audio signal, by applying a first machine-learning architecture to the feature vector for the one or more acoustic features, the one or more quality measures corresponding to_a similarity between one or more expected quality descriptors and one or more quality descriptors for the call audio of the inbound audio signal;

extracting, by the computer, an inbound speaker embedding for the inbound speaker from the one or more acoustic features for the inbound audio signal, by applying a second machine-learning architecture to the feature vector for the one or more acoustic features of the inbound audio signal;

generating, by the computer, a first similarity score for the inbound speaker based upon the inbound speaker embedding and an enrolled voiceprint for an enrolled speaker, by applying the second machine-learning architecture;

generating, by the computer, a second similarity score for verifying the inbound speaker, the second similarity score generated based upon the one or more quality measures and the first similarity score; and

verifying, by the computer, the inbound speaker as the enrolled speaker based upon comparing the second similarity score against a verification threshold.

2. The method according to claim 1 , wherein generating the one or more quality measures for the inbound audio signal includes generating, by the computer, the overall quality measure based upon each of the quality measures.

3. The method according to claim 1 , wherein generating the one or more quality measures includes:

generating, by the computer, a plurality of speech segments from the inbound audio signal; and

determining, by the computer, a total duration of speech based upon the plurality of speech segments.

4. The method according to claim 1 , wherein the first machine-learning architecture generates a quality embedding corresponding to each respective quality descriptor.

5. The method according to claim 4 , wherein the quality descriptor includes at least one of an audio event descriptor, a codec descriptor, a microphone type descriptor, a device type, and a network type.

6. The method according to claim 1 , wherein generating a quality measure includes determining, by the computer, a similarity between the inbound speaker embedding and a corresponding enrolled speaker embedding for an enrolled audio signal, wherein the quality measure is based upon the similarity.

7. The method according to claim 1 , further comprising:

receiving, by the computer, one or more clean enrollment audio signals for the enrolled speaker;

generating, by the computer, one or more degraded enrollment audio signals corresponding to the one or more clean enrollment audio signals according to a type of degradation; and

extracting, by the computer, one or more enrolled quality embeddings for the enrolled speaker by applying the first machine-learning architecture on the one or more clean enrollment audio signals and the one or more degraded enrollment audio signals.

8. A system comprising:

a database configured store an enrolled voiceprint for an enrolled speaker; and

a server comprising a processor configured to:

extract from an inbound audio signal for an inbound speaker a feature vector for one or more acoustic features;

generate one or more quality measures and an overall quality measure for the inbound audio signal, by applying a first machine-learning architecture to the feature vector for the one or more acoustic features, the one or more quality measures corresponding to a similarity between one or more expected quality descriptors and one or more quality descriptors for the call audio of the inbound audio signal;

extract an inbound speaker embedding for the inbound speaker from the one or more acoustic features for the inbound audio signal, by applying a second machine-learning architecture to the feature vector for the one or more acoustic features of the inbound audio signal;

generate a first similarity score for the inbound speaker based upon the inbound speaker embedding and the enrolled voiceprint for the enrolled speaker, by applying the second machine-learning architecture;

generate a second similarity score for verifying the inbound speaker, the second similarity score generated based upon the one or more quality measures and the first similarity score; and

verify the inbound speaker as the enrolled speaker based upon comparing the second similarity score against a verification threshold.

9. The system according to claim 8 , wherein when generating the one or more quality measures for the inbound audio signal, the server is further configured to generate the overall quality measure based upon each of the quality measures.

10. The system according to claim 8 , wherein when generating the one or more quality measures, the server is further configured to:

generate a plurality of speech segments from the inbound audio signal; and

determine a total duration of speech based upon the plurality of speech segments.

11. The system according to claim 8 , wherein when generating a quality measure the server is configured to:

determine a similarity between the inbound speaker embedding and a corresponding enrolled speaker embedding for an enrolled audio signal, wherein the quality measure is based upon the similarity.

12. The system according to claim 8 , wherein the server is further configured to:

receive one or more clean enrollment audio signals for the enrolled speaker;

generate one or more degraded enrollment audio signals corresponding to the one or more clean enrollment audio signals according to a type of degradation; and

extract one or more enrolled quality embeddings for the enrolled speaker by applying the first machine-learning architecture on the one or more clean enrollment audio signals and the one or more degraded enrollment audio signals.

13. A non-transitory computer-readable medium comprising a non-transitory storage memory configured to store machine-readable instructions that when executed by a processor instruct the processor to:

extract from an inbound audio signal for an inbound speaker a feature vector for one or more acoustic features;

generate one or more quality measures and an overall quality measure for the inbound audio signal, by applying a first machine-learning architecture to the feature vector for the one or more acoustic features;

extract an inbound speaker embedding for the inbound speaker from the one or more acoustic features for the inbound audio signal, by applying a second machine-learning architecture to the feature vector for the one or more acoustic features of the inbound audio signal, the one or more quality measures corresponding to a similarity between one or more expected quality descriptors and one or more quality descriptors for the call audio of the inbound audio signal;

generate a first similarity score for the inbound speaker based upon the inbound speaker embedding and an enrolled voiceprint for an enrolled speaker, by applying the second machine-learning architecture;

generate a second similarity score for verifying the inbound speaker, the second similarity score generated based upon the one or more quality measures and the first similarity score; and

verify the inbound speaker as the enrolled speaker based upon comparing the second similarity score against a verification threshold.

14. The computer-readable medium of claim 13 , the processor further instructed to:

generate a plurality of speech segments from the inbound audio signal; and

determine a total duration of speech based upon the plurality of speech segments.

15. The computer-readable medium of claim 13 , wherein the first machine-learning architecture generates a quality embedding corresponding to each respective quality descriptor.

16. The computer-readable medium of claim 13 , the processor further instructed to determine a similarity between the inbound speaker embedding and a corresponding enrolled speaker embedding for an enrolled audio signal, wherein a quality measure is based upon the similarity.

17. The computer-readable medium of claim 13 , the processor further instructed to:

receive one or more clean enrollment audio signals for the enrolled speaker;

generate one or more degraded enrollment audio signals corresponding to the one or more clean enrollment audio signals according to a type of degradation; and

extract one or more enrolled quality embeddings for the enrolled speaker by applying the first machine-learning architecture on the one or more clean enrollment audio signals and the one or more degraded enrollment audio signals.

Assignments (4)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
RELEASE OF SECURITY INTEREST Recorded Jun 26, 2024
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: PINDROP SECURITY, INC.
Reel/Frame 069477/0962 →
SECURITY INTEREST Recorded Jul 31, 2023
From: PINDROP SECURITY, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064443/0584 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 20, 2021
From: RAO, HRISHIKESH; PHATAK, KEDAR; KHOURY, ELIE
To: PINDROP SECURITY, INC.
Reel/Frame 057247/0432 →
Continuity (2)
Provisional Application 63068685 · Aug 21, 2020
Related Publication 20220059121A1 · Feb 24, 2022
References Cited (27)
US 9824692B1 · Khoury et al. · 2017 [cited by applicant]
US 10141009B2 · Khoury et al. · 2018 [cited by applicant]
US 10403291B2 · Moreno · 2019 [cited by examiner]
US 10692502B2 · Khoury et al. · 2020 [cited by applicant]
US 10923111B1 · Fan · 2021 [cited by examiner]
US 11095572B1 · Stafford et al. · 2021 [cited by applicant]
US 11238843B2 · Arik · 2022 [cited by examiner]
US 11322157B2 · Vaquero · 2022 [cited by examiner]
US 11437027B1 · Guo · 2022 [cited by examiner]
US 20140379332A1 · Rodriguez et al. · 2014 [cited by applicant]
US 20150301796A1 · Visser et al. · 2015 [cited by applicant]
US 20150356974A1 · Tani · 2015 [cited by examiner]
US 20180293221A1 · Finkelstein · 2018 [cited by examiner]
US 20190333522A1 · Lesso · 2019 [cited by examiner]
US 20200184967A1 · Gupta · 2020 [cited by examiner]
US 20200312313A1 · Maddali · 2020 [cited by examiner]
US 20200312337A1 · Stafylakis · 2020 [cited by examiner]
US 20200322377A1 · Lakhdhar · 2020 [cited by examiner]
US 20210233541A1 · Chen et al. · 2021 [cited by applicant]
US 20210241776A1 · Sivaraman et al. · 2021 [cited by applicant]
US 20210280171A1 · Phatak et al. · 2021 [cited by applicant]
US 20220030345A1 · Gong · 2022 [cited by examiner]
WO WO2019145708A1 · 2019 [cited by applicant]
“Robustness of Quality-based Score Calibration of Speaker Recognition Systems with respect to low-SNR and short-duration conditions”, Biometrics and Internet Security Research Group, Hochschule Darmstadt, Germany, Odyss… [cited by examiner]
“Quality_Measure_Functions_for_Calibration_of_Speaker_Recognition_Systems_in_Various_Duration_Conditions”, IEEE Transactions on Audio, Speech, and Language Processing, vol. 21, No. 11, Nov. 2013 (Year: 2013). [cited by examiner]
International Search Report and Written Opinion for PCT/US2021/046901 dated Dec. 6, 2021 (13 pages). [cited by applicant]
International Preliminary Report on Patentability for PCT App. PCT/US2021/046901 dated Feb. 16, 2023 (11 pages). [cited by applicant]