IP Library › Granted Patent US 12,573,406
Granted Patent B2
US 12,573,406 · App. 17/977,443 · Granted Mar 10, 2026

Voice authentication based on acoustic and linguistic machine learning models

Inventors: Stéphane B. Martin (Lausanne, CH); Erwan Barry Tarik Zerhouni (Zürich, CH)
Assignee: CISCO TECHNOLOGY, INC.
G10L17/14G06F40/30G10L17/02G10L17/04G10L17/10G10L17/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,406
App. No.
17/977,443
Granted
Mar 10, 2026
Kind
B2
Abstract

In one example embodiment, acoustic characteristics of a user voice are analyzed by a first machine learning model of a processor. Linguistic patterns in the user voice are analyzed by a second machine learning model of the processor. The user is authenticated with respect to an authorized user by the processor based on analysis of the acoustic characteristics and the linguistic patterns of the user voice by the first and second machine learning models.

Claims (41)

1 . A method comprising:

analyzing acoustic characteristics of a user voice of a user by a first machine learning model of a processor, wherein the first machine learning model generates a first value indicating correspondence of the acoustic characteristics of the user voice with an authorized user;

analyzing linguistic patterns in the user voice by a second machine learning model of the processor, wherein the linguistic patterns include words specific to particular geographic regions and vocabulary used by social groups, and the second machine learning model generates a second value indicating correspondence of the linguistic patterns in the user voice with the authorized user;

combining the first value and the second value by the processor to produce an authenticity value indicating authenticity of the user with respect to the authorized user;

determining correspondence by the processor between speed of talking and pauses in the user voice and speed of talking and pauses in speech of the authorized user; and

authenticating the user with respect to the authorized user by the processor based on the authenticity value satisfying a threshold and the speed of talking and pauses in the user voice corresponding to the speech of the authorized user.

2 . The method of claim 1 , wherein the linguistic patterns further include one or more from a group of use of vocabulary and phrases, preferences in speech register, syntactic constructions, length and complexity of sentences and clauses, and linguistic attributes.

3 . The method of claim 1 , wherein the first machine learning model processes audio signals of the user voice, and wherein the second machine learning model processes text produced from the audio signals of the user voice and generates the second value indicating correspondence of the linguistic patterns of the text with the authorized user.

4 . The method of claim 1 , further comprising:

training the first and second machine learning models by the processor using voice samples of authorized users and voice samples of unauthorized users.

5 . The method of claim 1 , further comprising:

training the first and second machine learning models by the processor using an adversarial network including a generator model that generates voice samples varying from voice samples of authorized users.

6 . The method of claim 5 , wherein the generator model includes a grammatical generator model to generate textual data and an acoustic generator model to generate audio samples from the textual data, and wherein the audio samples are provided by the generator model for training the first and second machine learning models.

7 . The method of claim 5 , wherein the generator model adapts the acoustic characteristics and the linguistic patterns of a voice sample of an authorized user to resemble the acoustic characteristics and the linguistic patterns of another individual.

8 . An apparatus comprising:

a computing system comprising one or more processors, wherein the one or more processors are configured to perform operations including:

analyzing acoustic characteristics of a user voice of a user by a first machine learning model, wherein the first machine learning model generates a first value indicating correspondence of the acoustic characteristics of the user voice with an authorized user;

analyzing linguistic patterns in the user voice by a second machine learning model, wherein the linguistic patterns include words specific to particular geographic regions and vocabulary used by social groups, and the second machine learning model generates a second value indicating correspondence of the linguistic patterns in the user voice with the authorized user;

combining the first value and the second value to produce an authenticity value indicating authenticity of the user with respect to the authorized user;

determining correspondence between speed of talking and pauses in the user voice and speed of talking and pauses in speech of the authorized user; and

authenticating the user with respect to the authorized user based on the authenticity value satisfying a threshold and the speed of talking and pauses in the user voice corresponding to the speech of the authorized user.

9 . The apparatus of claim 8 , wherein the linguistic patterns further include one or more from a group of use of vocabulary and phrases, preferences in speech register, syntactic constructions, length and complexity of sentences and clauses, and linguistic attributes.

10 . The apparatus of claim 8 , wherein the first machine learning model processes audio signals of the user voice, and wherein the second machine learning model processes text produced from the audio signals of the user voice and generates the second value indicating correspondence of the linguistic patterns of the text with the authorized user.

11 . The apparatus of claim 8 , wherein the one or more processors are further configured to perform operations including:

training the first and second machine learning models using an adversarial network including a generator model that generates voice samples varying from voice samples of authorized users.

12 . The apparatus of claim 11 , wherein the generator model includes a grammatical generator model to generate textual data and an acoustic generator model to generate audio samples from the textual data, and wherein the audio samples are provided by the generator model for training the first and second machine learning models.

13 . The apparatus of claim 11 , wherein the generator model adapts the acoustic characteristics and the linguistic patterns of a voice sample of an authorized user to resemble the acoustic characteristics and the linguistic patterns of another individual.

14 . One or more non-transitory computer readable storage media encoded with processing instructions that, when executed by one or more processors, cause the one or more processors to perform operations including:

analyzing acoustic characteristics of a user voice of a user by a first machine learning model, wherein the first machine learning model generates a first value indicating correspondence of the acoustic characteristics of the user voice with an authorized user;

analyzing linguistic patterns in the user voice by a second machine learning model, wherein the linguistic patterns include words specific to particular geographic regions and vocabulary used by social groups, and the second machine learning model generates a second value indicating correspondence of the linguistic patterns in the user voice with the authorized user;

combining the first value and the second value to produce an authenticity value indicating authenticity of the user with respect to the authorized user;

determining correspondence between speed of talking and pauses in the user voice and speed of talking and pauses in speech of the authorized user; and

authenticating the user with respect to the authorized user based on the authenticity value satisfying a threshold and the speed of talking and pauses in the user voice corresponding to the speech of the authorized user.

15 . The one or more non-transitory computer readable storage media of claim 14 , wherein the linguistic patterns further include one or more from a group of use of vocabulary and phrases, preferences in speech register, syntactic constructions, length and complexity of sentences and clauses, and linguistic attributes.

16 . The one or more non-transitory computer readable storage media of claim 14 , wherein the first machine learning model processes audio signals of the user voice, and wherein the second machine learning model processes text produced from the audio signals of the user voice and generates the second value indicating correspondence of the linguistic patterns of the text with the authorized user.

17 . The one or more non-transitory computer readable storage media of claim 14 , wherein the processing instructions further cause the one or more processors to perform operations including:

training the first and second machine learning models using voice samples of authorized users and voice samples of unauthorized users.

18 . The one or more non-transitory computer readable storage media of claim 14 , wherein the processing instructions further cause the one or more processors to perform operations including:

training the first and second machine learning models using an adversarial network including a generator model that generates voice samples varying from voice samples of authorized users.

19 . The one or more non-transitory computer readable storage media of claim 18 , wherein the generator model includes a grammatical generator model to generate textual data and an acoustic generator model to generate audio samples from the textual data, and wherein the audio samples are provided by the generator model for training the first and second machine learning models.

20 . The one or more non-transitory computer readable storage media of claim 18 , wherein the generator model adapts the acoustic characteristics and the linguistic patterns of a voice sample of an authorized user to resemble the acoustic characteristics and the linguistic patterns of another individual.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2022
From: MARTIN, STÉPHANE B.; ZERHOUNI, ERWAN BARRY TARIK
To: CISCO TECHNOLOGY, INC.
Reel/Frame 061608/0312 →
Continuity (1)
Related Publication 20240144935A1 · May 2, 2024
References Cited (37)
US 9686275B2 · Chari · 2017 [cited by examiner]
US 10027662B1 · Mutagi · 2018 [cited by examiner]
US 10277590B2 · Chari · 2019 [cited by examiner]
US 10910105B2 · McCloskey · 2021 [cited by examiner]
US 10957318B2 · Mishra · 2021 [cited by examiner]
US 11238845B2 · Chen · 2022 [cited by examiner]
US 11551699B2 · Huh · 2023 [cited by examiner]
US 11587569B2 · Ye · 2023 [cited by examiner]
US 11823684B2 · Sharifi · 2023 [cited by examiner]
US 11869511B2 · Siyavudeen · 2024 [cited by examiner]
US 11929078B2 · Tuo · 2024 [cited by examiner]
US 11942094B2 · Chojnacka · 2024 [cited by examiner]
US 12015637B2 · Lakhdhar · 2024 [cited by examiner]
US 12093827B2 · Kursun · 2024 [cited by examiner]
US 20080059176A1 · Ravi et al. · 2008 [cited by applicant]
US 20180146370A1 · Krishnaswamy · 2018 [cited by examiner]
US 20200118544A1 · Lee et al. · 2020 [cited by applicant]
US 20210174813A1 · Huh et al. · 2021 [cited by applicant]
US 20210326421A1 · Khoury · 2021 [cited by examiner]
US 20220004904A1 · Stemmer · 2022 [cited by examiner]
US 20220059121A1 · Rao · 2022 [cited by examiner]
US 20220121868A1 · Chen · 2022 [cited by examiner]
US 20240127825A1 · Carroll · 2024 [cited by examiner]
WO 2021216299A1 · 2021 [cited by applicant]
WO 2022065879A1 · 2022 [cited by applicant]
Cai, Danwei, Zexin Cai, and Ming Li, “Identifying Source Speakers for Voice Conversion based Spoofing Attacks on Speaker Verification Systems”, Jun. 2022, arXiv preprint arXiv:2206.09103. (Year: 2022). [cited by examiner]
Laptev, Aleksandr, Roman Korostik, Aleksey Svischev, Andrei Andrusenko, Ivan Medennikov, and Sergey Rybin, “You Do Not Need More Data: Improving End-to-End Speech Recognition by Text-to-Speech Data Augmentation”, Oct. 2… [cited by examiner]
Zhao, Yuanjun, Roberto Togneri, and Victor Sreeram, “Replay anti-spoofing countermeasure based on data augmentation with post selection”, May 2020, Computer Speech and Language, vol. 64, No. 101115, pp. 1-19. (Year: 202… [cited by examiner]
Ceaparu, Marian, Stefan-Adrian Toma, Svetlana Segarceanu, George Suciu, and Inge Gavat, “Multifactor Voice-Based Authentication System”, Feb. 2020, Journal of Engineering Science and Technology Review, Special Issue on … [cited by examiner]
Safavi, Saeid, Hock Gan, Iosif Mporas, and Reza Sotudeh, “Fraud Detection in Voice-based Identity Authentication Applications and Services”, Dec. 2016, 2016 IEEE 16th International Conference on Data Mining Workshops (I… [cited by examiner]
Bhattacharya, G. et al., “Generative Adversarial Speaker Embedding Networks for Domain Robust End-To-End Speaker Verification,” McGill University, Computer Research Institute of Montreal, https://arxiv.org/pdf/1811.0306… [cited by applicant]
Sriram, A. et al., “Robust Speech Recognition Using Generative Adversarial Networks,” Baidu Research, Sunnyvale, CA, USA, https://arxiv.org/pdf/1711.01567.pdf, Nov. 5, 2017, 5 pages. [cited by applicant]
Khanjani, Z. et al., “How Deep Are the Fakes? Focusing on Audio Deepfake: A Survey,” University of Maryland Baltimore County, Information System department, USA, https://arxiv.org/pdf/2111.14203.pdf, Nov. 28, 2021, 27 p… [cited by applicant]
Xu, W. et al., “DRB-GAN: A Dynamic ResBlock Generative Adversarial Network for Artistic Style Transfer,” https://arxiv.org/pdf/2108.07379.pdf, Oct. 2, 2021, 20 pages. [cited by applicant]
Azure Cognitive Services | Microsoft Learn, “What is speaker recognition?,” https://learn.microsoft.com/en-us/azure/cognitive-services/speech-service/speaker-recognition-overview, Aug. 25, 2022, 4 pages. [cited by applicant]
Text Similarity API | Twinword, “How to Use APIs,” Evaluate the similarity of two words, sentences, or paragraphs., https://www.twinword.com/api/text-similarity.php, retrieved Oct. 31, 2022, 5 pages. [cited by applicant]
PyTorch, “From Research to Production,” 1.13 Core blog: PyTorch 1.13 release, including beta versions of functorch and improved support for Apple's new M1 chips., https://pytorch.org/, retrieved Oct. 31, 2022, 3 pages. [cited by applicant]