IP Library › Granted Patent US 12,198,718
Granted Patent B2
US 12,198,718 · App. 18/446,623 · Granted Jan 14, 2025

Self-supervised speech representations for fake audio detection

Inventors: Joel Shor (Mountain View, CA); Alanna Foster Slocum (San Francisco, CA)
Assignee: Google LLC
G10L25/69G10L15/02G10L15/063G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,718
App. No.
18/446,623
Granted
Jan 14, 2025
Kind
B2
Abstract

A method for determining synthetic speech includes receiving audio data characterizing speech in audio data obtained by a user device. The method also includes generating, using a trained self-supervised model, a plurality of audio features vectors each representative of audio features of a portion of the audio data. The method also includes generating, using a shallow discriminator model, a score indicating a presence of synthetic speech in the audio data based on the corresponding audio features of each audio feature vector of the plurality of audio feature vectors. The method also includes determining whether the score satisfies a synthetic speech detection threshold. When the score satisfies the synthetic speech detection threshold, the method includes determining that the speech in the audio data obtained by the user device comprises synthetic speech.

Claims (40)

1. A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving audio data characterizing speech obtained by a user device;

generating, using a shallow discriminator model, a score indicating a presence of synthetic speech in the audio data; and

based on determining that the score satisfies a synthetic speech detection threshold, determining that the speech in the audio data obtained by the user device comprises synthetic speech,

wherein the shallow discriminator model is trained on a plurality of mixed training samples, each particular training sample of the plurality of mixed training samples comprising;

a corresponding human-originated portion being first human-originated speech representing a first portion of a corresponding training utterance; and

a corresponding synthetic speech portion being synthetic speech representing a second portion of the corresponding training utterance, the corresponding synthetic speech portion replacing second human-originated speech that represents the second portion of the corresponding training utterance for the particular training sample.

2. The computer-implemented method of claim 1 , wherein the shallow discriminator model comprises an intelligent pooling layer.

3. The method of claim 2 , wherein the intelligent pooling layer is configured to aggregate a plurality of embeddings of the audio data.

4. The computer-implemented method of claim 2 , wherein the operations further comprise:

generating, using the intelligent pooling layer of the shallow discriminator model, a single final audio feature vector based on the audio data,

wherein generating the score indicating the presence of the synthetic speech in the audio data is based on the single final audio feature vector.

5. The computer-implemented method of claim 4 , wherein the shallow discriminator model further comprises a single fully-connected layer.

6. The computer-implemented method of claim 5 , wherein the single fully-connected layer is configured to receive, as input, the single final audio feature vector and generate, as output, the score.

7. The computer-implemented method of claim 1 , wherein the shallow discriminator model comprises a logistic regression model.

8. The computer-implemented method of claim 1 , wherein the shallow discriminator model comprises a linear discriminant analysis model.

9. The computer-implemented method of claim 1 , wherein the shallow discriminator model comprises a random forest model.

10. The computer-implemented method of claim 1 , wherein the data processing hardware resides on the user device.

11. The computer-implemented method of claim 1 , wherein the shallow discriminator model comprises a single intelligent pooling layer and a single fully-connected layer.

12. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations comprising:

receiving audio data characterizing speech obtained by a user device;

generating, using a shallow discriminator model, a score indicating a presence of synthetic speech in the audio data; and

based on determining that the score satisfies a synthetic speech detection threshold, determining that the speech in the audio data obtained by the user device comprises synthetic speech,

wherein the shallow discriminator model is trained on a plurality of mixed training samples, each particular training sample of the plurality of mixed training samples comprising:

a corresponding human-originated portion being first human-originated speech representing a first portion of a corresponding training utterance; and

a corresponding synthetic speech portion being synthetic speech representing a second portion of the corresponding training utterance, the corresponding synthetic speech portion replacing second human-originated speech that represents the second portion of the corresponding training utterance for the particular training sample.

13. The system of claim 12 , wherein the shallow discriminator model comprises an intelligent pooling layer.

14. The system of claim 13 , wherein the intelligent pooling layer is configured to aggregate a plurality of embeddings of the audio data.

15. The system of claim 13 , wherein the operations further comprise:

generating, using the intelligent pooling layer of the shallow discriminator model, a single final audio feature vector based on the audio data,

wherein generating the score indicating the presence of the synthetic speech in the audio data is based on the single final audio feature vector.

16. The system of claim 15 , wherein the shallow discriminator model further comprises a single fully-connected layer.

17. The system of claim 16 , wherein the single fully-connected layer is configured to receive, as input, the single final audio feature vector and generate, as output, the score.

18. The system of claim 12 , wherein the shallow discriminator model comprises a logistic regression model.

19. The system of claim 12 , wherein the shallow discriminator model comprises a linear discriminant analysis model.

20. The system of claim 12 , wherein the shallow discriminator model comprises a random forest model.

21. The system of claim 12 , wherein the data processing hardware resides on the user device.

22. The system of claim 12 , wherein the shallow discriminator model comprises a single intelligent pooling layer and a single fully-connected layer.

Continuity (2)
Continuation 17110278 · Dec 2, 2020
Related Publication 20230386506A1 · Nov 30, 2023
References Cited (23)
US 11756572B2 · Shor · 2023 [cited by examiner]
US 20180254046A1 · Khoury et al. · 2018 [cited by applicant]
US 20200090644A1 · Klingler et al. · 2020 [cited by applicant]
US 20200279568A1 · Vaquero Avilés-Casco · 2020 [cited by examiner]
US 20200322377A1 · Lakhdhar et al. · 2020 [cited by applicant]
WO WO2018160943A1 · 2018 [cited by examiner]
WO 2019231624A2 · 2019 [cited by applicant]
WO WO2021075063A1 · 2021 [cited by examiner]
WO WO2021135454A1 · 2021 [cited by examiner]
Weicheng Cai, Haiwei Wu, Danwei Cai, Ming Li “The DKU Replay Detection System for the ASVspoof 2019 Challenge:On Data Augmentation, Feature Representation, Classification, and Fusion” arXiv:1907.02663v1 (Year: 2019). [cited by examiner]
H. Muckenhirn, P. Korshunov, M. Magimai-Doss and S. Marcel, “Long-Term Spectral Statistics for Voice Presentation Attack Detection,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 25, No. 11, p… [cited by examiner]
R. Reimao and V. Tzerpos, “FoR: A Dataset for Synthetic Speech Detection,” 2019 International Conference on Speech Technology and Human-Computer Dialogue (SpeD), Timisoara, Romania, 2019, pp. 1-10, doi: 10.1109/SPED.201… [cited by examiner]
Mari Ganesh Kumar, Suvidha Rupesh Kumar, Saranya M, B. Bharathi, Hema A. Murthy “Spoof detection using time-delay shallow neural network and feature switching” arXiv:1904.07453v2 (Year: 2020). [cited by examiner]
Synthetic Speech Detection using Fundamental Frequency Variation and Spectral Features, https://www.sciencedirect.com/science/a <https://www.sciencedirect.com/science/article/abs/pii/S0885230817301663> rticle/abs/pii/S0… [cited by applicant]
One Deep Music Representation to Rule Them All? A Comparative Analysis of Different Representation Learning Strategies, https://link.springer.com/article/10.1007/s <https://link.springer.com/article/10.1007/s00521-019-0… [cited by applicant]
Automatic Identification of Gender from Speech, http://www.cs.columbia.edu/˜sarah ita/papers/speech_prosody16.pdf <http://www.cs.columbia.edu/%7Esarahita/papers/speech_prosody16.pdf> Sarah Ita Levitan, Taniya Mishra, Sr… [cited by applicant]
Utterance-level Aggregation for Speaker Recognition in the Wild, http://www.robots.ox.ac.uk/˜vgg/research/speakerID/, Weidi Xie, Arsha Nagrani, Joon Son Chung, Andrew Zisserman, Aug. 26, 2020. [cited by applicant]
ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection <https://arxiv.org/pdf/1904.05441.pdf>). [cited by applicant]
Toward Learning a Universal Non-Semantic Representation of Speech, Cornel University, [Submitted on Feb. 25, 2020 (v1), last revised Aug. 6, 2020 (this version, v6)] <https://arxiv.org/abs/2002.12764>). [cited by applicant]
International Search Report for the related Application No. IDF-305093, Dated Aug. 8, 2020. [cited by applicant]
International Search Report and Written Opinion, related to Application No. PCT/US2021/059033, dated Feb. 16, 2022. [cited by applicant]
Haibin Wu et al., “Defense for Black-box Attacks on Anti-spoofing Models by Self-Supervised Learning”, Dec. 7, 2020 arXIv, pp. 3780-3784. DOI: 10.21437/Interspeech.2020-2026. [cited by applicant]
Office Action issued in related Japanese Patent Application No. 2023-533746, dated Jul. 16, 2024. [cited by applicant]