IP Library Granted Patent US 12,190,888
Granted Patent B2
US 12,190,888 · App. 17/840,795 · Granted Jan 7, 2025

System and method for secure training of speech processing systems

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,888
App. No.
17/840,795
Granted
Jan 7, 2025
Kind
B2
Abstract

A method, computer program product, and computing system for generating an obscured speech signal from an input speech signal and an obscured transcription from a transcription of the input speech signal. A speaker embedding may be extracted from the input speech signal. A speaker embedding delta may be generated based upon, at least in part, the extracted speaker embedding and a synthetic speaker embedding. A synthetic speech signal may be generated from the obscured speech signal using the synthetic speaker embedding. A residual signal may be generated based upon, at least in part, the obscured speech signal and the speaker embedding delta. A speech processing system may be trained using the obscured transcription, the synthetic speech signal, the speaker embedding delta, and the residual signal.

Claims (49)

1. A computer-implemented method, executed on a computing device, comprising:

generating an obscured speech signal from an input speech signal and an obscured transcription from a transcription of the input speech signal, wherein the obscured speech signal and the obscured transcription include obscured representations of sensitive content from the input speech signal and the transcription of the input speech signal;

extracting a speaker embedding from the input speech signal;

generating a speaker embedding delta based upon, at least in part, the extracted speaker embedding and a synthetic speaker embedding;

generating a synthetic speech signal from the obscured speech signal using the synthetic speaker embedding;

generating a residual signal based upon, at least in part, the obscured speech signal and the speaker embedding delta; and

training a speech processing system using the obscured transcription, the synthetic speech signal, the speaker embedding delta, and the residual signal.

2. The computer-implemented method of claim 1 , wherein the sensitive content from the input speech signal and the transcription of the input speech signal includes personally identifiable information (PII) and/or protected health information (PHI).

3. The computer-implemented method of claim 1 , wherein generating the residual signal includes:

generating a resynthesized speech signal from the synthetic speech signal using the speaker embedding delta.

4. The computer-implemented method of claim 3 , wherein generating the residual signal includes:

generating the residual signal as the difference between the obscured speech signal and the resynthesized speech signal.

5. The computer-implemented method of claim 1 , wherein training the speech processing system includes:

generating a reconstructed speech signal using the synthetic speech signal, the speaker embedding delta, and the residual signal.

6. The computer-implemented method of claim 5 , wherein training the speech processing system includes:

training the speech processing system with the reconstructed speech signal and the obscured transcription.

7. The computer-implemented method of claim 1 , wherein generating the obscured speech signal and the obscured transcription includes:

marking each obscured representation of sensitive content in the obscured transcription, thus defining a plurality of obscured portion markings.

8. The computer-implemented method of claim 7 , wherein training the speech processing system includes:

training the speech processing system with a reconstructed speech signal, the obscured transcription, and the plurality of obscured portion markings.

9. A computing system comprising:

a memory; and

a processor configured to generate an obscured speech signal from an input speech signal and an obscured transcription from a transcription of the input speech signal, wherein the obscured speech signal and the obscured transcription include obscured representations of sensitive content from the input speech signal and the transcription of the input speech signal, wherein the processor is further configured to extract a speaker embedding from the input speech signal, wherein the processor is further configured to generate a speaker embedding delta based upon, at least in part, the extracted speaker embedding and a synthetic speaker embedding, wherein the processor is further configured to generate a synthetic speech signal from the obscured speech signal using the synthetic speaker embedding, wherein the processor is further configured to generate a residual signal based upon, at least in part, the obscured speech signal and the speaker embedding delta, and wherein the processor is further configured to train a speech processing system using the obscured transcription, the synthetic speech signal, the speaker embedding delta, and the residual signal.

10. The computing system of claim 9 , wherein the sensitive content from the input speech signal and the transcription of the input speech signal includes personally identifiable information (PII) and/or protected health information (PHI).

11. The computing system of claim 9 , wherein generating the residual signal includes:

generating a resynthesized speech signal from the synthetic speech signal using the speaker embedding delta.

12. The computing system of claim 11 , wherein generating the residual signal includes:

generating the residual signal as the difference between the obscured speech signal and the resynthesized speech signal.

13. The computing system of claim 9 , wherein training the speech processing system includes:

generating a reconstructed speech signal using the synthetic speech signal, the speaker embedding delta, and the residual signal.

14. The computing system of claim 12 , wherein training the speech processing system includes:

training the speech processing system with the reconstructed speech signal and the obscured transcription.

15. The computing system of claim 9 , wherein generating the obscured speech signal and the obscured transcription includes:

marking each obscured representation of sensitive content in the obscured transcription, thus defining a plurality of obscured portion markings.

16. A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:

generating an obscured speech signal from an input speech signal and an obscured transcription from a transcription of the input speech signal, wherein the obscured speech signal and the obscured transcription include obscured representations of sensitive content from the input speech signal and the transcription of the input speech signal;

extracting a speaker embedding from the input speech signal;

generating a speaker embedding delta based upon, at least in part, the extracted speaker embedding and a synthetic speaker embedding;

generating a synthetic speech signal from the obscured speech signal using the synthetic speaker embedding;

generating a resynthesized speech signal from the synthetic speech signal using the speaker embedding delta;

generating a residual signal based upon, at least in part, the obscured speech signal and the resynthesized speech signal; and

training a speech processing system using the obscured transcription, the synthetic speech signal, the speaker embedding delta, and the residual signal.

17. The computer program product of claim 16 , wherein the sensitive content from the input speech signal and the transcription of the input speech signal includes personally identifiable information (PII) and/or protected health information (PHI).

18. The computer program product of claim 16 , wherein training the speech processing system includes:

generating a reconstructed speech signal using the synthetic speech signal, the speaker embedding delta, and the residual signal.

19. The computer program product of claim 18 , wherein training the speech processing system includes:

training the speech processing system with the reconstructed speech signal and the obscured transcription.

20. The computer program product of claim 16 , wherein generating the obscured speech signal and the obscured transcription includes:

marking each obscured representation of sensitive content in the obscured transcription, thus defining a plurality of obscured portion markings.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2024
From: NESPOLI, FRANCESCO
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 067048/0227 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2022
From: YIN, SHOU-CHUN; PARK, JUNHO; SHARMA, DUSHYANT; KIM, DOYEONG
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 060207/0888 →