IP Library › Granted Patent US 12,726,477
Granted Patent B1
US 12,726,477 · App. 18/382,277 · Granted Sep 1, 2026

Audio authentication for digital telephony

Inventors: Derrick John Fitzgerald (Woodstock, GA); David Chei Seong Yap (San Jose, CA)
Assignee: Zoom Communications, Inc.
H04L63/0861G10L15/02G10L17/04G10L17/18H04L63/0428H04L65/65H04M3/493H04M7/006
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,726,477
App. No.
18/382,277
Filed
Oct 20, 2023
Granted
Sep 1, 2026
Kind
B1
Art Unit
2434
USPC
713/186
Abstract

Techniques for audio authentication for digital telephony are disclosed. In an example method, a computing device receives a user voice sample. The computing device determines one or more user voice features from the user voice sample and generates a user voice model based on the one or more user voice features. Next, the computing device receives, from a client device, first authentication information including an authentication voice sample. The computing device again determines one or more authentication voice features from the authentication voice sample and generates an authentication voice model based on the one or more authentication voice features. The computing device then determines a difference between the user voice model and the authentication voice model. In response to the difference being less than a predetermined threshold, the computing device outputs, to the client device, an authentication grant.

Claims (76)

1 . A method, comprising:

receiving a user voice sample;

determining user voice features from the user voice sample;

determining that a number of user voice features exceeds a threshold number of features;

responsive to determining that the number of user voice features exceeds the threshold number of features, generating a user voice model based on the one or more user voice features;

caching the user voice model in an in-memory cache;

receiving, from a client device, first authentication information comprising an authentication voice sample;

determining one or more authentication voice features from the authentication voice sample;

generating an authentication voice model based on the one or more authentication voice features;

retrieving the user voice model from the in-memory cache;

determining a difference between the user voice model and the authentication voice model, comprising:

determining an embedded representation of the user voice sample and an embedded representation of the authentication voice sample, wherein the embedded representation of the user voice sample and the embedded representation of the authentication voice sample each comprise a multi-dimensional vector;

determining a metric characterizing a vector space relationship between the embedded representation of the user voice sample and the embedded representation of the authentication voice sample; and

comparing the metric to a predetermined threshold that is mapped to a numerical probability that the user voice model does not differ from the authentication voice sample; and

responsive to the difference being less than the predetermined threshold, outputting, to the client device, an authentication grant.

2 . The method of claim 1 , wherein the user voice sample and the authentication voice sample are encoded using a codec.

3 . The method of claim 1 , wherein the client device is an Internet Protocol (IP) telephony device.

4 . The method of claim 3 , wherein the user voice sample and the authentication voice sample are received using an interactive voice response (IVR) system.

5 . The method of claim 1 , wherein determining the difference between the user voice model and the authentication voice model further comprises:

inputting the user voice model and the authentication voice model to a trained machine learning model, the trained machine learning model trained to generate a probability that an input voice model pair are based on the user's voice; and

receiving, from the trained machine learning model, a classification of the user voice model and authentication voice model pair.

6 . The method of claim 5 , wherein the trained machine learning model is further trained based on the classification of the user voice model and the authentication voice model pair.

7 . The method of claim 1 , wherein the one or more user voice features and the one or more authentication voice features comprise one or more of pitch, tone, timbre, rhythm, or pronunciation.

8 . The method of claim 1 , wherein outputting, to the client device, the authentication grant is further responsive to receiving, from the client device, valid multi-factor authentication information.

9 . The method of claim 1 , wherein the user voice sample and the authentication voice sample are encrypted using a cryptographic transport protocol.

10 . The method of claim 9 , wherein the cryptographic transport protocol is the Secure Real-time Transport Protocol (SRTP).

11 . A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations including:

receiving a user voice sample;

determining user voice features from the user voice sample;

determining that a number of user voice features exceeds a threshold number of features;

responsive to determining that the number of user voice features exceeds the threshold number of features, generating a user voice model based on the one or more user voice features;

caching the user voice model in an in-memory cache;

receiving, from a client device, first authentication information comprising an authentication voice sample;

determining one or more authentication voice features from the authentication voice sample;

generating an authentication voice model based on the one or more authentication voice features;

retrieving the user voice model from the in-memory cache;

determining a difference between the user voice model and the authentication voice model, comprising:

determining an embedded representation of the user voice sample and an embedded representation of the authentication voice sample, wherein the embedded representation of the user voice sample and the embedded representation of the authentication voice sample each comprise a multi-dimensional vector;

determining a metric characterizing a vector space relationship between the embedded representation of the user voice sample and the embedded representation of the authentication voice sample; and

comparing the metric to a predetermined threshold that is mapped to a numerical probability that the user voice model does not differ from the authentication voice sample; and

responsive to the difference being less than the predetermined threshold, outputting, to the client device, an authentication grant.

12 . The non-transitory computer-readable medium of claim 11 , wherein the client device is an Internet Protocol (IP) telephony device.

13 . The non-transitory computer-readable medium of claim 11 , wherein determining the difference between the user voice model and the authentication voice model further comprises:

inputting the user voice model and the authentication voice model to a trained machine learning model, the trained machine learning model trained to generate a probability that an input voice model pair are based on the user's voice; and

receiving, from the trained machine learning model, a classification of the user voice model and authentication voice model pair.

14 . The non-transitory computer-readable medium of claim 13 , wherein the trained machine learning model is further trained based on the classification of the user voice model and the authentication voice model pair.

15 . A system comprising:

one or more processors; and

one or more computer-readable storage media storing instructions which, when executed by the one or more processors, cause the one or more processors to perform operations including:

receiving a user voice sample;

determining user voice features from the user voice sample;

determining that a number of user voice features exceeds a threshold number of features;

responsive to determining that the number of user voice features exceeds the threshold number of features, generating a user voice model based on the one or more user voice features;

caching the user voice model in an in-memory cache;

receiving, from a client device, first authentication information comprising an authentication voice sample;

determining one or more authentication voice features from the authentication voice sample;

generating an authentication voice model based on the one or more authentication voice features;

retrieving the user voice model from the in-memory cache;

determining a difference between the user voice model and the authentication voice model, comprising:

determining an embedded representation of the user voice sample and an embedded representation of the authentication voice sample, wherein the embedded representation of the user voice sample and the embedded representation of the authentication voice sample each comprise a multi-dimensional vector;

determining a metric characterizing a vector space relationship between the embedded representation of the user voice sample and the embedded representation of the authentication voice sample; and

comparing the metric to a predetermined threshold that is based on a numerical probability that the user voice model does not differ from the authentication voice sample; and

responsive to the difference being less than the predetermined threshold, outputting, to the client device, an authentication grant.

16 . The system of claim 15 , wherein the client device is an Internet Protocol (IP) telephony device.

17 . The system of claim 15 , wherein determining the difference between the user voice model and the authentication voice model further comprises:

inputting the user voice model and the authentication voice model to a trained machine learning model, the trained machine learning model trained to generate a probability that an input voice model pair are based on the user's voice; and

receiving, from the trained machine learning model, a classification of the user voice model and authentication voice model pair.

18 . The system of claim 15 , wherein the one or more user voice features and the one or more authentication voice features comprise one or more of pitch, tone, timbre, rhythm, or pronunciation.

19 . The method of claim 1 , wherein the metric is one of the Euclidean distance, the cosine similarity, or the dot product between the embedded representation of the user voice sample and the embedded representation of the authentication voice sample.

20 . The method of claim 4 , wherein:

receiving the user voice sample comprises configuring the IVR system, comprising:

providing a first voice prompt requesting the user voice sample via the IVR system; and

receiving the user voice sample in response to the first voice prompt; and

receiving, from the client device, the first authentication information comprising the authentication voice sample comprises:

providing a second voice prompt requesting the authentication voice sample via the IVR system; and

receiving the authentication voice sample in response to the second voice prompt.

Assignments (2)
CHANGE OF NAME Recorded Jul 24, 2026
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 076071/0671 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2023
From: FITZGERALD, DERRICK JOHN; YAP, DAVID CHEI SEONG
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 065295/0852 →
References Cited (26)
US 9940934B2 · Sachdev · 2018 [cited by examiner]
US 10593336B2 · Boyadjiev · 2020 [cited by examiner]
US 10652238B1 · Edwards · 2020 [cited by examiner]
US 10867021B1 · Shelton · 2020 [cited by examiner]
US 10956907B2 · Parnell · 2021 [cited by examiner]
US 11100739B1 · Mathew · 2021 [cited by examiner]
US 12189739B2 · Williams · 2025 [cited by examiner]
US 20050096906A1 · Barzilay · 2005 [cited by examiner]
US 20120204035A1 · Camenisch · 2012 [cited by examiner]
US 20120253810A1 · Sutton · 2012 [cited by examiner]
US 20150249664A1 · Talhami · 2015 [cited by examiner]
US 20160277439A1 · Rotter · 2016 [cited by examiner]
US 20160353282A1 · Richards · 2016 [cited by examiner]
US 20180004925A1 · Petersen · 2018 [cited by examiner]
US 20180096354A1 · Kohli · 2018 [cited by examiner]
US 20210193174A1 · Enzinger · 2021 [cited by examiner]
US 20210328801A1 · Sly · 2021 [cited by examiner]
US 20220116388A1 · Johnson · 2022 [cited by examiner]
US 20220270611A1 · Tuo · 2022 [cited by examiner]
US 20220277070A1 · Robert Jose · 2022 [cited by examiner]
US 20230041266A1 · Earman · 2023 [cited by examiner]
US 20230081988A1 · Chung · 2023 [cited by examiner]
US 20230161853A1 · Nair · 2023 [cited by examiner]
US 20240236086A1 · Traywick · 2024 [cited by examiner]
US 20240428101A1 · Smith · 2024 [cited by examiner]
US 20250005123A1 · Goel · 2025 [cited by examiner]