IP Library › Granted Patent US 12,210,606
Granted Patent B1
US 12,210,606 · App. 18/629,085 · Granted Jan 28, 2025

Methods and systems for enhancing the detection of synthetic speech

Inventors: Raphael A Rodriguez (Marco Island, FL); Olena Mizynchuk (Marco Island, FL); Davyd Mizynchuk (Kyiv, UA)
Assignee: Daon Technology
G06F21/32G10L17/02G10L17/04G10L17/06G10L17/26H04N21/4415G06Q20/40G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,210,606
App. No.
18/629,085
Granted
Jan 28, 2025
Kind
B1
Abstract

A method for enhancing detection of synthetic speech is provided that includes the step of receiving, by an electronic device, voice biometric data of a user captured while the user was speaking and analyzing the context in which the received voice biometric data was captured. The context includes environmental and situational factors. Moreover, the method includes the steps of analyzing characteristics of the received voice biometric data for anomalies associated with synthetic speech, generating a risk score based on the results of the analysis, and comparing the risk score against a threshold value. In response to determining the risk score fails to satisfy the threshold score, the method includes a step of determining the captured voice biometric data includes anomalies associated with synthetic speech and initiating an alert protocol.

Claims (88)

1. A method for enhancing detection of synthetic speech comprising the steps of:

receiving, by an electronic device, voice biometric data of a user captured while the user was speaking;

analyzing the context in which the received voice biometric data was captured, wherein the context includes environmental and situational factors;

analyzing characteristics of the received voice biometric data for anomalies associated with synthetic speech, wherein the anomalies include a lack of variability in duration;

generating a risk score based on the results of the context analysis and the characteristics analysis;

comparing the risk score against a threshold value; and

in response to determining the risk score fails to satisfy the threshold score, determining the captured voice biometric data includes anomalies associated with synthetic speech and initiating an alert protocol.

2. The method according to claim 1 , wherein the alert protocol comprises:

manually reviewing the received voice biometric data; and

categorizing the received voice biometric data as potentially synthetic speech based on the risk score.

3. The method according to claim 1 , wherein the characteristics comprise:

range of pitch;

timbre;

intensity;

prosody; and

pace, rhythm, and nature of speech.

4. The method according to claim 1 , wherein the anomalies comprise:

a narrow range of pitch;

a lack of expected complexity, unusual harmonic structures, and erratic formant movements;

variations in loudness not corresponding with an expressed or expected emotion;

consistent speech rate;

abnormal pauses;

inconsistencies in stress patterns;

intonation curves unusual for the context;

hesitations or rushed speech;

lack of natural pitch variation across sentences;

an unexpected pitch contour within a phrase; and

unusually long or short durations.

5. The method according to claim 1 , further comprising:

determining a frequency range for synthetic speech and a frequency range for genuine speech;

determining the frequency range of the received voice biometric data;

determining whether the received voice biometric data frequency range is within the synthetic or genuine frequency ranges;

in response to determining the received voice biometric data frequency range is within the synthetic frequency range, subjecting the received voice biometric data to additional review; and

in response to determining the received voice biometric data frequency range is within the genuine frequency range, determining the received voice biometric data is genuine.

6. The method according to claim 5 , said determining steps comprising operating, by the electronic device, a machine learning model trained to recognize and differentiate between synthetic and genuine speech.

7. The method according to claim 6 , further comprising updating the machine learning model using data from authentication transactions to enhance the accuracy of the model in recognizing and differentiating between synthetic and genuine speech.

8. An electronic device for enhancing detection of synthetic speech comprising:

a processor; and

a memory configured to store data, said electronic device being associated with a network and said memory being in communication with said processor and having instructions stored thereon which, when read and executed by said processor, cause said electronic device to:

receive voice biometric data of a user captured while the user was speaking;

analyze the context in which the received voice biometric data was captured, wherein the context includes environmental and situational factors;

analyze characteristics of the received voice biometric data for anomalies associated with synthetic speech, wherein the anomalies include a lack of variability in duration;

generate a risk score based on the results of the context analysis and the analysis of the received voice biometric data for anomalies associated with synthetic speech;

compare the risk score against a threshold value; and

in response to determining the risk score fails to satisfy the threshold score, determine the received voice biometric data includes anomalies associated with synthetic speech and initiating an alert protocol.

9. The electronic device according to claim 8 , wherein the instructions when read and executed by said processor, cause said electronic device to:

prompt a manual review of the received voice biometric data; and

categorize the received voice biometric data as potentially synthetic based on the risk score.

10. The electronic device according to claim 8 , wherein the characteristics comprise:

range of pitch;

timbre;

intensity;

prosody; and

pace, rhythm, and nature of speech.

11. The electronic device according to claim 8 , wherein the anomalies comprise:

a narrow range of pitch;

a lack of expected complexity, unusual harmonic structures, and erratic formant movements;

variations in loudness not corresponding with an expressed or expected emotion;

consistent speech rate;

abnormal pauses;

inconsistencies in stress patterns;

intonation curves unusual for the context;

hesitations or rushed speech;

lack of natural pitch variation across sentences;

an unexpected pitch contour within a phrase; and

unusually long or short durations.

12. The electronic device according to claim 8 , wherein the instructions when read and executed by said processor, cause said electronic device to:

determine a frequency range for synthetic speech and a frequency range for genuine speech;

determine the frequency range of the received voice biometric data;

determine whether the received voice biometric data frequency range is within the synthetic or genuine frequency ranges;

in response to determining the received voice biometric data frequency range is within the synthetic frequency range, subject the received voice biometric data to additional review; and

in response to determining the received voice biometric data frequency range is within the genuine frequency range, determine the received voice biometric data is genuine.

13. The electronic device according to claim 8 , wherein the instructions when read and executed by said processor, cause said electronic device to operate a machine learning model trained to recognize and differentiate between synthetic and genuine speech.

14. The electronic device according to claim 13 , wherein the instructions when read and executed by said processor, cause said electronic device to update the machine learning model using data from authentication transactions to enhance the accuracy of the model in recognizing and differentiating between synthetic and genuine speech.

15. A non-transitory computer-readable recording medium in an electronic device for authenticating users, the non-transitory computer-readable recording medium storing instructions which when executed by a hardware processor cause the non-transitory recording medium to perform steps comprising:

receiving voice biometric data of a user captured while the user was speaking;

analyzing the context in which the received voice biometric data was captured, wherein the context includes environmental and situational factors;

analyzing characteristics of the received voice biometric data for anomalies associated with synthetic speech, wherein the anomalies include a lack of variability in duration;

generating a risk score based on the results of the context analysis and the characteristics analysis;

comparing the risk score against a threshold value; and

in response to determining the risk score fails to satisfy the threshold score, determining the captured voice biometric data includes anomalies associated with synthetic speech and initiating an alert protocol.

16. The non-transitory computer-readable recording medium according to claim 15 , wherein the instructions when read and executed by said processor, further cause said non-transitory computer-readable recording medium to perform a step comprising operating a machine trained model trained to recognize and differentiate between synthetic and genuine speech.

17. The non-transitory computer-readable recording medium according to claim 15 , wherein the instructions when read and executed by said processor, further cause said non-transitory computer-readable recording medium to perform the steps of:

determining a frequency range for synthetic speech and a frequency range for genuine speech;

determining the frequency range of the received voice biometric data;

determining whether the received voice biometric data frequency range is within the synthetic or genuine frequency ranges;

in response to determining the received voice biometric data frequency range is within the synthetic frequency range, subjecting the received voice biometric data to additional review; and

in response to determining the received voice biometric data frequency range is within the genuine frequency range, determining the received voice biometric data is genuine.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2024
From: RODRIGUEZ, RAPHAEL A., MR.; MIZYNCHUK, OLENA, MS.; MIZYNCHUK, DAVYD, MR.
To: DAON TECHNOLOGY
Reel/Frame 067175/0614 →
References Cited (12)
US 9865253B1 · De Leon · 2018 [cited by examiner]
US 10943604B1 · Bone · 2021 [cited by examiner]
US 11955122B1 · Ahmadi · 2024 [cited by examiner]
US 20170084295A1 · Tsiartas · 2017 [cited by examiner]
US 20200321009A1 · Khoury · 2020 [cited by examiner]
US 20210193174A1 · Enzinger · 2021 [cited by examiner]
US 20220036904A1 · Traynor · 2022 [cited by examiner]
US 20220328050A1 · Hennig · 2022 [cited by examiner]
US 20230107624A1 · Keith, Jr. · 2023 [cited by examiner]
US 20230411008A1 · Khanzada · 2023 [cited by examiner]
Mireia F. Cabeceran, Fusing Prosodic and Acoustic Information for Speaker Recognition, PhD Dissertation, TALP Research Center, Barcelona (Year: 2008). [cited by examiner]
B. Yegnanarayana, Combining Evidence from Source Suprasegmental and Spectral Features for a Fixed-Text Speaker Verification System, IEEE (Year: 2005). [cited by examiner]