IP Library Granted Patent US 10,950,244
Granted Patent B2
US 10,950,244 · App. 16/385,511 · Granted Mar 16, 2021

System and method for speaker authentication and identification

Inventor: Milind Borkar (Plano, TX)
Assignee: ILLUMA Labs LLC.
G10L17/16G06F17/16G10L17/04G10L17/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,950,244
App. No.
16/385,511
Granted
Mar 16, 2021
Kind
B2
Abstract

A system and method for enrolling a speaker in a speaker authentication and identification system (AIS), the method comprising: generating a user account, the user account comprising: a user identifier based on one or more metadata elements associated with an audio input received from an end device; generating a first i-vector from an audio frame of the audio input, a trained T-matrix, and a Universal Background Model (UBM), wherein the first i-vector generation comprises an optimized computation; and associating the user account with the first i-vector.

Claims (62)

1. A method for enrolling a speaker in a speaker authentication and identification system (AIS), the method comprising:

generating a user account, the user account comprising: a user identifier based on one or more metadata elements associated with an audio input received from an end device;

generating a first i-vector from an audio frame of the audio input, a trained T-matrix, and a Universal Background Model (UBM), wherein the first i-vector generation comprises an optimized computation; and

associating the user account with the first i-vector, and wherein the optimized computation comprises:

generating a feature matrix based on a feature vector;

generating a Gaussian Mixture Model (GMM) mean matrix based on a plurality of mean vectors associated with a plurality of GMM components; and

generating a delta matrix based on the feature matrix and the GMM mean matrix.

2. The method of claim 1 , wherein the optimized computation comprises:

generating a first multi-dimensional array comprising a plurality of duplicated matrices, each matrix comprised of a plurality of GMM mean vectors;

generating a multi-dimensional feature matrix, comprising a plurality of feature matrices, each feature matrix corresponding to a feature vector of a single audio frame; and

generating a multi-dimensional delta array based on the first multi-dimensional array and the multi-dimensional feature matrix.

3. The method of claim 1 , wherein the optimized computation comprises:

detecting diagonal matrices; and

performing only those computations that involve diagonal elements.

4. The method of claim 1 , wherein the optimized computation comprises:

detecting computations in an intermediate result that generate an off-diagonal element of a matrix which is diagonalized; and

eliminating the computation of the intermediate result.

5. The method of claim 1 , wherein the optimized computation comprises:

detecting a recurring computation;

precomputing the recurring computation; and

storing the precomputed result in a cache of a processor.

6. The method of claim 1 , wherein the optimized computation comprises:

detecting a computation between a first matrix and a second matrix; and

replacing the detected computation with an element by element computation between the first matrix and the second matrix in response to determining that the replacement will result in the same output.

7. The method of claim 1 , further comprising extracting channel features from the audio input.

8. The method of claim 7 , further comprising generating the first i-vector further based on the extracted channel features.

9. The method of claim 7 , further comprising: associating the extracted channel features with the generated user account.

10. The method of claim 1 , wherein the metadata elements are any of: phone number, ANI, personal identifier number, IMEI, IMSI, end device make, and end device model.

11. The method of claim 1 , wherein the first i-vector is generated by: the end device, the AIS, or any combination thereof.

12. The method of claim 1 , further comprising:

receiving a second i-vector generated from a second end device, and associated metadata; and

generating a match score between the second i-vector and the first i-vector, in response to matching at least one metadata element value between the metadata of the first i-vector and metadata of the second i-vector.

13. The method of claim 12 , further comprising:

generating a notification that the speaker is authenticated, in response to the match score exceeding a first threshold value.

14. The method of claim 12 , further comprising:

generating a notification that the speaker is not authenticated, in response to the match score being below a first threshold.

15. The method of claim 12 , wherein the second end device is a first end device.

16. The method of claim 1 , further comprising:

receiving a second i-vector generated from a second end device, and associated metadata; and

generating a match score between the second i-vector and each of a plurality of first i-vectors.

17. The method of claim 16 , further comprising:

generating a notification for each match score exceeding a second threshold value, as a possible identification.

18. The method of claim 17 , wherein the notification further includes metadata associated with each of the first i-vectors for which a match exceeded the second threshold.

19. The method of claim 16 , further comprising:

generating a notification based on the first i-vector of the plurality of i-vectors which has the highest match score with the second i-vector.

20. The method of claim 1 , wherein generating a user account is performed in response to determining that the first i-vector is not enrolled.

21. A non-transitory computer readable medium having stored thereon instructions for causing a processing circuitry to perform a process, the process comprising:

generating a user account, the user account comprising: a user identifier based on one or more metadata elements associated with an audio input received from an end device;

generating a first i-vector from an audio frame of the audio input, a trained T-matrix, and a Universal Background Model (UBM), wherein the first i-vector generation comprises an optimized computation; and

associating the user account with the first i-vector, and wherein the optimized computation comprises:

generating a feature matrix based on a feature vector;

generating a Gaussian Mixture Model (GMM) mean matrix based on a plurality of mean vectors associated with a plurality of GMM components; and

generating a delta matrix based on the feature matrix and the GMM mean matrix.

22. A system for enrolling a speaker in a speaker authentication and identification system (AIS), comprising:

a processing circuitry; and

a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to:

generate a user account, the user account comprising: a user identifier based on one or more metadata elements associated with an audio input received from an end device;

generate a first i-vector from an audio frame of the audio input, a trained T-matrix, and a Universal Background Model (UBM), wherein the first i-vector generation comprises an optimized computation; and

associate the user account with the first i-vector, wherein the system is further configured to:

generate a feature matrix based on a feature vector;

generate a Gaussian Mixture Model (GMM) mean matrix based on a plurality of mean vectors associated with a plurality of GMM components; and

generate a delta matrix based on the feature matrix and the GMM mean matrix.

Assignments (3)
SECURITY INTEREST Recorded Dec 27, 2024
From: ILLUMA LABS INC.
To: STIFEL BANK
Reel/Frame 069692/0400 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2024
From: BORKAR, MILIND
To: ILLUMA LABS INC.
Reel/Frame 067787/0102 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2019
From: BORKAR, MILIND
To: ILLUMA LABS INC.
Reel/Frame 048896/0552 →
Continuity (8)
Continuation In Part 16290399 · Mar 1, 2019
Continuation In Part 16203077 · Nov 28, 2018
Continuation In Part 16385511
Continuation In Part 16203077 · Nov 28, 2018
Provisional Application 62658123 · Apr 16, 2018
Provisional Application 62638086 · Mar 3, 2018
Provisional Application 62592156 · Nov 29, 2017
Related Publication 20190244622A1 · Aug 8, 2019