IP Library › Granted Patent US 12,562,184
Granted Patent B2
US 12,562,184 · App. 18/040,812 · Granted Feb 24, 2026

Synthetic speech detection

Inventors: Ke Wang (Beijing, CN); Lei He (Beijing, CN)
Assignee: Microsoft Technology Licensing, LLC.
G10L25/78G06F21/55
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,562,184
App. No.
18/040,812
Granted
Feb 24, 2026
Kind
B2
Abstract

Systems and methods for synthetic speech detection includes receiving an input sample comprising audio and extracting acoustic features corresponding to speech in the audio. The extracted acoustic features are processed using a plurality of neural networks to output abstracted features and generating a feature vector corresponding to the abstracted features using pooling. Training of an SSD task, a speaker classification task, and a channel classification task are performed at a same time, using the feature vector. Synthetic speech is detected using at least the trained SSD task.

Claims (39)

1 . A computerized method for synthetic speech detection (SSD), the computerized method comprising:

receiving an input sample comprising audio;

extracting a plurality of acoustic features corresponding to speech in the audio;

processing the plurality of extracted acoustic features in parallel using a plurality of neural networks to output abstracted feature vectors, wherein the plurality of neural networks includes at least deep neural networks (DNNs) having a shared output used to perform training of an SSD model, a speaker classification model, and a channel classification model;

generating, via the plurality of neural networks, a pooled feature vector from the output abstracted feature vectors using a pooling operation, wherein the pooling operation comprises multi-head attentive pooling that assigns to each of the output abstracted feature vectors, a respective attention weight, used to generate the pooled feature vector, and wherein the pooled feature vector is a single vector;

performing the training of an SSD task of the SSD model, a speaker classification task of the speaker classification model, and a channel classification task of the channel classification model, at substantially a same time, using the pooled feature vector; and

detecting synthetic speech using at least the SSD task of the SSD model.

2 . The computerized method of claim 1 , wherein the training is performed using a feed-forward layer comprising the SSD model, the speaker classification model and the channel classification model having shared information.

3 . The computerized method of claim 1 , further comprising identifying at least one of a physical attack (PA) and a logical attack (LA) using the detected synthetic speech.

4 . The computerized method of claim 1 , wherein the pooling operation further comprises an averaging operation using a plurality of weights corresponding to the plurality of extracted acoustic features.

5 . The computerized method of claim 1 , further comprising using a gradient reversal layer in combination with the pooling operation to generate the pooled feature vector.

6 . A system for synthetic speech detection (SSD), the system comprising: at least one processor; and

at least one memory comprising computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the at least one processor to:

receive an input sample comprising audio;

extract a plurality of acoustic features corresponding to speech in the audio;

process the plurality extracted acoustic features in parallel using a plurality of neural networks to output abstracted feature vectors wherein the plurality of neural networks includes at least deep neural networks (DNNs) having a shared output used to perform training of an SSD model, a speaker classification model, and a channel classification model;

generate, via the plurality of neural networks, a pooled feature vector from to the output abstracted feature vectors using a pooling operation, wherein the pooling operation comprises multi-head attentive pooling that assigns to each of the output abstracted feature vectors, a respective attention weight, used to generate the pooled feature vector, and wherein the pooled feature vector is a single vector;

perform the training of an SSD task of the SSD model, a speaker classification task of the speaker classification model, and a channel classification task of the channel classification model, at substantially a same time, using the pooled feature vector; and

detect synthetic speech using at least the SSD task of the SSD model.

7 . The system of claim 6 , wherein the training is performed using a feed-forward layer comprising the SSD model, the speaker classification model, and the channel classification model having shared information.

8 . The system of claim 6 , further comprising identifying at least one of a physical attack (PA) and a logical attack (LA) using the detected synthetic speech.

9 . The system of claim 6 , wherein the pooling operation further comprises an averaging operation using a plurality of weights corresponding to the plurality of extracted acoustic features.

10 . The system of claim 6 , further comprising using a gradient reversal layer in combination with the pooling operation to generate the feature vector.

11 . One or more computer storage media having computer-executable instructions for synthetic speech detection (SSD) that upon execution by a processor, cause the processor to:

receive an input sample comprising audio;

filter the audio to extract a plurality of acoustic features corresponding to speech in the audio of the input sample;

process the plurality of extracted acoustic features in parallel using a plurality of neural networks to output abstracted feature vectors, wherein the plurality of neural networks includes at least deep neural networks (DNNs) having a shared output used to perform training of an SSD model, a speaker classification model, and a channel classification model;

generate, via the plurality of neural networks, a pooled feature vector from the output abstracted feature vectors using a pooling operation, wherein the pooling operation comprises multi-head attentive pooling that assigns to each of the output abstracted feature vectors, a respective attention weight, used to generate the pooled feature vector, and wherein the pooled feature vector is a single vector;

perform the training of an SSD task of the SSD model, a speaker classification task of the speaker classification model, and a channel classification task of the channel classification model, at substantially a same time, using the pooled feature vector; and

detect synthetic speech using at least the SSD task of the SSD model.

12 . The one or more computer storage media of claim 11 , having further computer-executable instructions that cause the processor to:

decode the processed audio to generate an output, the output including an indication of a word or word sequence received as part of the input sample that includes the detected synthetic speech.

13 . The one or more computer storage media of claim 11 , having further computer-executable instructions that cause the processor to:

generate a log probability that one or more input segments of the audio in the input sample are the synthetic speech.

14 . The one or more computer storage media of claim 13 , wherein the log probability defines a score corresponding to a likelihood that the one or more input segments are the synthetic speech.

15 . The one or more computer storage media of claim 14 , having further computer-executable instructions that cause the processor to:

convert the score to user displayable information showing SSD results and speaker information corresponding to the score.

16 . The one or more computer storage media of claim 11 , having further computer-executable instructions that cause the processor to:

identify at least one of a physical attack (PA) or a logical attack (LA) using the detected synthetic speech.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2023
From: WANG, KE; HE, LEI
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 062607/0531 →
Continuity (1)
Related Publication 20240005947A1 · Jan 4, 2024
References Cited (28)
US 9142218B2 · Schroeter · 2015 [cited by applicant]
US 9466299B1 · Feltham et al. · 2016 [cited by applicant]
US 9484036B2 · Kons et al. · 2016 [cited by applicant]
US 9824692B1 · Khoury · 2017 [cited by examiner]
US 12015637B2 · Lakhdhar · 2024 [cited by examiner]
US 20110112833A1 · Frankel · 2011 [cited by examiner]
US 20160284347A1 · Sainath et al. · 2016 [cited by applicant]
US 20190325861A1 · Singh · 2019 [cited by examiner]
US 20200321009A1 · Khoury · 2020 [cited by applicant]
US 20210005067A1 · Salekin · 2021 [cited by examiner]
US 20220095061A1 · Diehl · 2022 [cited by examiner]
US 20220172739A1 · Shor · 2022 [cited by examiner]
CN 109655815A · 2019 [cited by applicant]
CN 110491391A · 2019 [cited by applicant]
CN 110648655A · 2020 [cited by applicant]
CN 111276131A · 2020 [cited by applicant]
CN 111324696A · 2023 [cited by applicant]
WO 2021002967A1 · 2021 [cited by applicant]
Choi, et al., “Neural MOS Prediction for Synthesized Speech using Multi-Task Learning with Spoofing Detection and Spoofing Type Classification”, In Repository of arXiv:2007.08267v2, Dec. 2, 2020, 8 pages. [cited by applicant]
He, et al., “Deep Residual Learning for Image Recognition”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 27, 2016, pp. 770-778. [cited by applicant]
Lavrentyeva, et al., “STC Antispoofing Systems for the ASVspoof2019 Challenge”, In Repository of arXiv:1904.05576v1, Apr. 11, 2019, 5 Pages. [cited by applicant]
Li, et al., “Anti-Spoofing Speaker Verification System with Multi-Feature Integration and Multi-Task Learning-”, In Proceedings of Annual Conference of the International Speech Communication Association, Sep. 15, 2019, … [cited by applicant]
Reimao, Ricardo, “Synthetic Speech Detection Using Deep Neural Networks”, In Thesis of York University, May 2019, 156 Pages. [cited by applicant]
Zhao, et al., “Multi-task Learning Based Spoofing-Robust Automatic Speaker Verification System”, In Journal of Latex Class Files, vol. 14, Issue 8, Aug. 2015, pp. 1-12. [cited by applicant]
Chen, et al., “Channel Invariant Speaker Embedding Learning with Joint Multi-Task and Adversarial Training”, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 6574-6578. [cited by applicant]
Extended European Search Report received for EP Application No. 21937300.8, mailed on Oct. 15, 2024, 10 pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/CN2021/088623”, Mailing Date: Jan. 19, 2022, 8 Pages. [cited by applicant]
First Office Action Received for Chinese Application No. 202180044082.8, mailed on Jul. 31, 2025, 17 pages. (English Translation Provided). [cited by applicant]