IP Library Granted Patent US 12,057,111
Granted Patent B2
US 12,057,111 · App. 17/325,886 · Granted Aug 6, 2024

System and method for voice biometrics authentication

Inventors: Natan Katz (Tel Aviv, IL); Ori Akstein (Petach Tikva, IL); Tal Haguel (Petach Tikva, IL)
Assignee: Nice Ltd.
G10L15/16G06N3/084G10L15/063G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,057,111
App. No.
17/325,886
Granted
Aug 6, 2024
Kind
B2
Abstract

A system and method for authenticating an identity may include generating a first generic representation representing a stored audio content, generating a second generic representation representing input audio content, and, providing the first and second generic representations to a voice biometrics unit adapted to authenticate an identity based on the first and second generic representations.

Claims (34)

1. A computer-implemented method of generating a system for enrollment, the method comprising:

creating a model by training a neural network to generate generic representations of audio content, wherein the audio content is compressed via at least two different compression methods using different codecs, and wherein layer information for the neural network is calculated based on bitrate of the compressed audio content and the sample rate of the compressed audio content;

including the model in a generic representation generation (GRG) unit;

generating, by the GRG unit, at least a first generic representation representing a stored audio content;

receiving input audio content, wherein the input audio content is compressed;

generating, by the GRG unit, a second generic representation representing the input audio content, wherein the first and second generic representations are usable, by a voice biometrics (VB) unit, to authenticate an identity associated with the input audio content;

associating, by the GRG unit, the first generic representation with a score value: and

providing the first and second generic representations to the VB unit if the score is within a predefined range, wherein the score is calculated using a loss function, wherein the loss function is a function that measures how much an output of the neural network deviates from an expected output.

2. The method of claim 1 , comprising providing the first and second generic representations to a second VB unit.

3. The method of claim 1 , wherein the first generic representation is generated based on a bitrate of the stored audio content and a sample rate of the stored audio content.

4. The method of claim 1 , wherein a generic representation includes a byte array of features extracted from audio content, and wherein the size of the byte array of features is set based on a bitrate of the stored audio content and a sample rate of the audio content.

5. The method of claim 1 , comprising:

generating, by the GRG unit, a third generic representation of a second stored audio content; and

providing the second and third generic representations to the VB unit.

6. The method of claim 1 , comprising generating a reconstructed audio content based on a generic representation and providing the reconstructed audio content to the VB unit.

7. The method of claim 1 , comprising:

if the score is not within the predefined range then using at least one of: the input audio content and the second generic representation to retrain the GRG unit.

8. The method of claim 7 , comprising:

after retraining the GRG unit, generating, by the retrained GRG unit, a new generic representation of the stored audio content; and

replacing the first generic representation by the newly created generic representation.

9. A system for enrollment, the system comprising:

a memory; and a generic representation generation (GRG) unit executed by a computer processor, the GRG unit including a model generated by training a neural network to generate generic representations of audio content, wherein the audio content is compressed via at least two different compression methods using different codecs, and wherein layer information of the neural network is calculated based on bitrate of the compressed audio content and the sample rate of the compressed audio content;

wherein the GRG unit is adapted to generate a first generic representation of audio stored in a storage system,

receiving input audio content, wherein the input audio content is compressed;

the GRG unit also configured to:

generate a second generic representation representing the input audio content, wherein the first and second generic representations are usable, by a voice biometrics (VB) unit, to authenticate an identity associated with the input audio content associate the first generic representation with a score value; and

provide the first and second generic representations to the VB unit if the score is within a predefined range, wherein the score is calculated using a loss function, wherein the loss function is a function that measures how much an output of the neural network deviates from an expected output.

10. The system of claim 9 , wherein a generic representation is generated based on at least one of: a bitrate of the stored audio content and a sample rate of the stored audio content.

11. The system of claim 9 , wherein a generic representation includes a byte array of features extracted from the compressed audio content, and wherein the size of the byte array of features is set based on a bitrate of the compressed audio content and a sample rate of the compressed audio content.

12. The system of claim 9 , wherein the GRG unit is further adapted to generate a reconstructed audio content based on a generic representation and provide the reconstructed audio content to the VB unit.

13. The system of claim 9 , wherein the GRG unit is further adapted to: if the score is not within the predefined range then using at least one of: the input audio content and the second generic representation to retrain the GRG unit.

14. The system of claim 13 , wherein the GRG unit is further adapted to:

generate, after a retraining, a new generic representation of the stored audio content; and

replace the first generic representation by the newly created generic representation.

Assignments (2)
SECURITY INTEREST Recorded Feb 26, 2026
From: NICE LTD; NICE SYSTEMS INC.; NICE SYSTEMS TECHNOLOGIES INC.; INCONTACT, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074986/0208 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2021
From: KATZ, NATAN; AKSTEIN, ORI; HAGUEL, TAL
To: NICE LTD.
Reel/Frame 057551/0030 →
Continuity (1)
Related Publication 20220375461A1 · Nov 24, 2022