IP Library Granted Patent US 12,531,051
Granted Patent B2
US 12,531,051 · App. 18/326,631 · Granted Jan 20, 2026

System and method for secure processing of speech signals using pseudo-speech representations

Inventors: Dushyant Sharma (Tracy, CA); William Francis Ganong, III (Brookline, MA); Daniel Paulino Almendro Barreda (Barcelona, ES); Patrick Aubrey Naylor (Reading, GB); Alvaro Martin Iturralde Zurita (Montrel, CA); Francesco Nespoli (London, GB)
Assignee: Microsoft Technology Licensing, LLC.
G10L13/033G10L13/08G10L15/22G10L15/26G10L19/018
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,531,051
App. No.
18/326,631
Granted
Jan 20, 2026
Kind
B2
Abstract

A method, computer program product, and computing system for processing a speech signal. A sensitive portion of the speech signal is identified. A pseudo-speech representation of the sensitive portion is generated using a voice converter system. Speech processing is performed on the speech signal and the pseudo-speech representation of the sensitive portion using a speech processing system.

Claims (30)

1 . A computer-implemented method comprising:

receiving, by a speech privacy processor, a speech signal recorded during a medical encounter between a medical professional and a patient, wherein a sensitive portion of the speech signal includes Personally Identifiable Information (PII) of the patient;

processing, by the speech privacy processor, the speech signal to obtain a synthetic speech signal at least by converting the PII within the sensitive portion of the speech signal into a pseudo-speech representation of the PII in the synthetic speech signal using a voice converter, wherein the voice converter reduces an acoustic quality of one or more linguistic units associated with the PII by filtering the sensitive portion of the speech signal in a modulation domain to remove modulations above a predefined threshold, thereby generating the pseudo-speech representation of the PII, wherein the modulation domain includes carrier signals and modulators; and

transmitting, by the speech privacy processor, the synthetic speech signal to an automatic speech recognition (ASR) processor.

2 . The computer-implemented method of claim 1 , wherein the PII within the sensitive portion of the speech signal is converted into the pseudo-speech representation based on a predefined acoustic signature.

3 . The computer-implemented method of claim 2 , wherein the predefined acoustic signature is based upon, at least in part, a content type of the sensitive portion of the speech signal.

4 . The computer-implemented method of claim 1 , wherein converting the PII within the sensitive portion of the speech signal into the pseudo-speech representation includes transforming an articulatory feature of the sensitive portion of the speech signal.

5 . The computer-implemented method of claim 1 , wherein converting the PII within the sensitive portion of the speech signal into the pseudo-speech representation includes performing a spectral modification of the sensitive portion of the speech signal.

6 . The computer-implemented method of claim 1 , further comprising:

marking the pseudo-speech representation with a watermark.

7 . A computing system comprising:

a processor; and

a memory storing programming instructions for execution by the processor, the programming instructions, upon execution by the processor, causing the computing system to perform the following steps:

receiving, by a speech privacy processor, a speech signal recorded during a medical encounter between a medical professional and a patient, wherein a sensitive portion of the speech signal includes Personally Identifiable Information (PII) of the patient;

processing, by the speech privacy processor, the speech signal to obtain a synthetic speech signal at least by converting the PII within the sensitive portion of the speech signal into a pseudo-speech representation of the PII in the synthetic speech signal using a voice converter, wherein the voice converter reduces an acoustic quality of one or more linguistic units associated with the PII by filtering the sensitive portion of the speech signal in a modulation domain to remove modulations above a predefined threshold, thereby generating the pseudo-speech representation of the PII, wherein the modulation domain includes carrier signals and modulators; and

transmitting, by the speech privacy processor, the synthetic speech signal to an automatic speech recognition (ASR) processor.

8 . The computing system of claim 7 , wherein the PII within the sensitive portion of the speech signal is converted into the pseudo-speech representation based on a predefined acoustic signature.

9 . The computing system of claim 8 , wherein the predefined acoustic signature is based upon, at least in part, a content type of the sensitive portion of the speech signal.

10 . The computing system of claim 7 , wherein converting the PII within the sensitive portion of the speech signal into the pseudo-speech representation includes transforming an articulatory feature of the sensitive portion of the speech signal.

11 . The computing system of claim 7 , wherein converting the PII within the sensitive portion of the speech signal into the pseudo-speech representation includes performing a spectral modification of the sensitive portion of the speech signal.

12 . The computer-implemented method of claim 1 , wherein the ASR processor is located remotely from the speech privacy processor, and wherein the synthetic speech signal is transmitted over a network that communicatively couples the speech privacy processor to the ASR processor.

13 . The computer-implemented method of claim 1 , wherein reducing the acoustic quality of the one or more linguistic units associated with the PII renders the pseudo-speech representation of the PII generally unintelligible during speech recognition by the ASR processor.

14 . The computer-implemented method of claim 1 , wherein the voice converter reduces the acoustic quality of the one or more linguistic units associated with the PII by shortening a duration of the one or more linguistic units associated with the PII.

15 . The computer-implemented method of claim 1 , wherein the voice converter reduces the acoustic quality of the one or more linguistic units associated with the PII by reducing a clarity of articulation of the one or more linguistic units associated with the PII.

16 . A computer program product residing on a non-transitory computer readable medium having programming instructions stored thereon which, when executed by a processor of a system, cause the system to perform the following operations:

receiving, by a speech privacy processor, a speech signal recorded during a medical encounter between a medical professional and a patient, wherein a sensitive portion of the speech signal includes Personally Identifiable Information (PII) of the patient;

processing, by the speech privacy processor, the to obtain a synthetic speech signal at least by converting the PII within the sensitive portion of the speech signal into a pseudo-speech representation of the PII in the synthetic speech signal using a voice converter, wherein the voice converter reduces an acoustic quality of one or more linguistic units associated with the PII by filtering the sensitive portion of the speech signal in a modulation domain to remove modulations above a predefined threshold, thereby generating the pseudo-speech representation of the PII, wherein the modulation domain includes carrier signals and modulators; and

transmitting, by the speech privacy processor, the synthetic speech signal to an automatic speech recognition (ASR) processor.

17 . The computer program product of claim 16 , wherein the voice converter reduces the acoustic quality of the one or more linguistic units associated with the PII by shortening a duration of the one or more linguistic units associated with the PII.

18 . The computer program product of claim 16 , wherein the voice converter reduces the acoustic quality of the one or more linguistic units associated with the PII by reducing a clarity of articulation of the one or more linguistic units associated with the PII.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2025
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 070761/0438 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2023
From: GANONG, WILLIAM FRANCIS, III; BARREDA, DANIEL PAULINO ALMENDRO
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 064609/0311 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 31, 2023
From: SHARMA, DUSHYANT; NAYLOR, PATRICK AUBREY; ITURRALDE ZURITA, ALVARO MARTIN; NESPOLI, FRANCESCO
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 063813/0507 →
Continuity (1)
Related Publication 20240404504A1 · Dec 5, 2024
References Cited (16)
US 20200258535A1 · Vatanparvar · 2020 [cited by examiner]
US 20210012026A1 · Taylor · 2021 [cited by examiner]
US 20220199093A1 · Ramadas · 2022 [cited by examiner]
WO 2014042715A1 · 2014 [cited by applicant]
WO 2021219554A1 · 2021 [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/030184, Sep. 16, 2024, 19 pages. [cited by applicant]
Invitation to pay additional fees received for PCT Application No. PCT/US2024/030184, Jul. 26, 2024, 9 pages. [cited by applicant]
Son, et al., “An Acoustic Description of Consonant Reduction,” Speech Communication, vol. 28, Issue 2, Jun. 1, 1999, pp. 125-140. [cited by applicant]
“Vowel reduction”, Retrieved From: https://en.wikipedia.org/wiki/Vowel_reduction, Jan. 23, 2022, 6 Pages. [cited by applicant]
Georges, et al., “Towards an articulatory-driven neural vocoder for speech synthesis”, 12th International Seminar on Speech Production, Dec. 2020, 2 Pages. [cited by applicant]
Koch, et al., “DigitalPhonetics / IMS-Toucan”, Retrieved From: https://github.com/DigitalPhonetics/IMS-Toucan, Mar. 23, 2023, 11 Pages. [cited by applicant]
Lux, et al., “Language-Agnostic Meta-Learning for Low-Resource Text-to-Speech with Articulatory Features”, In Repository of arXiv:2203.03191v1, Mar. 7, 2022, 12 Pages. [cited by applicant]
Murphy, James J. , “Demosthenes”, Retrieved From: https://www.britannica.com/biography/Demosthenes-Greek-statesman-and-orator, May 7, 2022, 10 Pages. [cited by applicant]
Sedivy, Julie, “Mumbling Isn't a Sign of Laziness—It's a Clever Data-Compression Trick”, Retrieved From: https://nautilus/mumbling-isnt-a-sign-of-lazinessits-a-clever-data_compression-trick-235301/, Feb. 18, 2015, 10 Pa… [cited by applicant]
Shor, et al., “Project Euphonia's Personalized Speech Recognition for Non-Standard Speech”, Retrieved From: https://ai.googleblog.com/2019/08/project-euphonias-personalized-speech.html, Aug. 13, 2019, 7 Pages. [cited by applicant]
Weirich, et al., “Mumbling: Macho or Morphology?”, In Journal of Speech, Language, and Hearing Research vol. 59, Dec. 2016, pp. 1587-1595. [cited by applicant]