IP Library Granted Patent US 12,676,164
Granted Patent B2
US 12,676,164 · App. 18/334,442 · Granted Jul 7, 2026

System and method for modulation domain-based audio signal encoding

Inventors: Dushyant Sharma (Tracy, CA); Patrick Aubrey Naylor (Reading, GB); William Francis Ganong, III (Brookline, MA)
Assignee: Microsoft Technology Licensing, LLC
G10L21/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,676,164
App. No.
18/334,442
Filed
Jun 14, 2023
Granted
Jul 7, 2026
Kind
B2
Art Unit
2658
USPC
704/500
Abstract

A method, computer program product, and computing system for processing an audio signal by converting the audio signal to the modulation domain. The modulation domain audio signal is encoded with a plurality of carrier signals and a plurality of modulator signals derived from the modulation domain audio signal. The encoded modulation domain audio signal is converted to the time domain.

Claims (32)

1 . A computer-implemented method comprising:

receiving, by an audio encoder, a voice audio signal carrying Personally Identifiable Information (PII) of a speaker;

encoding the voice audio signal to obtain an encoded voice audio signal for securely communicating to a distributed speech processing machine learning (ML) model via a telecommunications network, wherein encoding the voice audio signal includes converting the voice audio signal into the modulation domain to obtain a modulation domain representation of the voice audio signal, encoding the modulation domain representation of the voice audio signal with a plurality of carrier signals and a plurality of modulator signals derived from the modulation domain representation of the voice audio signal, and converting the encoded modulation domain representation of the voice audio signal to the time domain to obtain the encoded voice audio signal; and

transmitting, via the telecommunications network, the encoded voice audio signal to the distributed speech processing ML model.

2 . The computer-implemented method of claim 1 , wherein the distributed speech processing ML model is trained for speech recognition using training data that includes voice audio encoded in the modulation domain.

3 . The computer-implemented method of claim 1 , further comprising:

processing the encoded voice audio signal directly using the distributed speech processing ML model upon reception over the telecommunications network.

4 . The computer-implemented method of claim 3 , wherein encoding the modulation domain representation of the voice audio signal includes processing an encoding key defining an encoding process for the modulation domain representation of the voice audio signal.

5 . The computer-implemented method of claim 4 , wherein decoding the modulation domain representation of the encoded voice audio signal includes processing the encoding key to decode the encoded voice audio signal.

6 . The computer-implemented method of claim 1 , wherein encoding the modulation domain representation of the voice audio signal includes switching the plurality of modulator signals within a plurality of carrier-modulator signal pairs.

7 . The computer-implemented method of claim 6 , wherein switching the plurality of modulator signals within the plurality of carrier-modulator signal pairs includes switching frequency-adjacent modulator signals between the plurality of carrier-modulator signal pairs.

8 . The computer-implemented method of claim 6 , wherein encoding the modulation domain representation of the voice audio signal includes switching the plurality of modulator signals between the plurality of carrier-modulator signal pairs based upon, at least in part, pitch information associated with the voice audio signal.

9 . A computing system comprising:

a processor; and

a memory storing programming instructions for execution by the processor, the programming instructions, upon execution by the processor, causing the computing system to perform the following operations:

receiving, by an audio encoder, a voice audio signal carrying Personally Identifiable Information (PII) of a speaker;

encoding the voice audio signal to obtain an encoded voice audio signal for securely communicating to a distributed speech processing machine learning (ML) model via a telecommunications network, wherein encoding the voice audio signal includes converting the voice audio signal into the modulation domain to obtain a modulation domain representation of the voice audio signal, encoding the modulation domain representation of the voice audio signal with a plurality of carrier signals and a plurality of modulator signals derived from the modulation domain representation of the voice audio signal, and converting the encoded modulation domain representation of the voice audio signal to the time domain to obtain the encoded voice audio signal; and

transmitting, via the telecommunications network, the encoded voice audio signal to the distributed speech processing ML model.

10 . The computing system of claim 9 , wherein the distributed speech processing ML model is trained for speech recognition using training data that includes voice audio encoded in the modulation domain.

11 . The computing system of claim 9 , wherein encoding the modulation domain representation of the voice audio signal includes switching the plurality of modulator signals within a plurality of carrier-modulator signal pairs.

12 . The computing system of claim 11 , wherein switching the plurality of modulator signals within the plurality of carrier-modulator signal pairs includes switching frequency-adjacent modulator signals between the plurality of carrier-modulator signal pairs.

13 . The computing system of claim 12 , wherein encoding the modulation domain representation of the voice audio signal includes switching the plurality of modulator signals between the plurality of carrier-modulator signal pairs based upon, at least in part, pitch information associated with the voice audio signal.

14 . A computer program product residing on a non-transitory computer readable medium having programming instructions stored thereon which, when executed by a processor of a system, cause the system to perform the following operations:

receiving, by an audio encoder, a voice audio signal carrying Personally Identifiable Information (PII) of a speaker;

encoding the voice audio signal to obtain an encoded voice audio signal for securely communicating to a distributed speech processing machine learning (ML) model via a telecommunications network, wherein encoding the voice audio signal includes converting the voice audio signal into the modulation domain to obtain a modulation domain representation of the voice audio signal, encoding the modulation domain representation of the voice audio signal with a plurality of carrier signals and a plurality of modulator signals derived from the modulation domain representation of the voice audio signal to obtain the encoded voice audio signal, and converting the encoded modulation domain representation of the voice audio signal to the time domain to obtain the encoded voice audio signal; and

transmitting, via the telecommunications network, the encoded voice audio signal to the distributed speech processing ML model.

15 . The computer program product of claim 14 , wherein the distributed speech processing ML model is trained for speech recognition using training data that includes voice audio encoded in the modulation domain.

16 . The computer program product of claim 14 , wherein encoding the modulation domain representation of the voice audio signal includes processing an encoding key defining an encoding process for the modulation domain representation of the voice audio signal.

17 . The computer program product of claim 16 , wherein decoding the modulation domain representation of the voice audio signal includes processing the encoding key to decode the modulation domain representation of the voice audio signal.

18 . The computer program product of claim 14 , wherein decoding the modulation domain representation of the voice audio signal includes switching the plurality of modulator signals within a plurality of carrier-modulator signal pairs.

19 . The computer program product of claim 18 , wherein switching the plurality of modulator signals within the plurality of carrier-modulator signal pairs includes switching frequency-adjacent modulator signals between the plurality of carrier-modulator signal pairs.

20 . The computer program product of claim 18 , wherein encoding the modulation domain representation of the voice audio signal includes switching the plurality of modulator signals between the plurality of carrier-modulator signal pairs based upon, at least in part, pitch information associated with the voice audio signal.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2025
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 070747/0464 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2023
From: SHARMA, DUSHYANT; NAYLOR, PATRICK AUBREY; GANONG, WILLIAM FRANCIS, III
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 063944/0123 →
Continuity (1)
Related Publication 20240420722A1 · Dec 19, 2024
References Cited (33)
US 5995539A · Miller · 1999 [cited by examiner]
US 8223872B1 · Zhang · 2012 [cited by examiner]
US 11799646B2 · Ramabadran · 2023 [cited by examiner]
US 20030081705A1 · Miller · 2003 [cited by examiner]
US 20070100610A1 · Disch · 2007 [cited by examiner]
US 20170245173A1 · Qian · 2017 [cited by examiner]
US 20200022143A1 · Abdoli · 2020 [cited by examiner]
US 20200258535A1 · Vatanparvar · 2020 [cited by applicant]
US 20210012026A1 · Taylor · 2021 [cited by applicant]
US 20220182271A1 · Shin · 2022 [cited by examiner]
US 20220199093A1 · Ramadas · 2022 [cited by applicant]
US 20230163902A1 · Saggar · 2023 [cited by examiner]
US 20240404504A1 · Sharma · 2024 [cited by examiner]
WO 2007107046A1 · 2007 [cited by applicant]
WO 2014042715A1 · 2014 [cited by applicant]
WO 2021219554A1 · 2021 [cited by applicant]
Ozkan, et al., “Secure Voice Communication via GSM Network”, In Proceedings of 7th International Conference on Electrical and Electronics Engineering, Dec. 1, 2011, pp. 267-271. [cited by applicant]
Qi, et al., “A Speech Privacy Protection Method Based on Sound Masking and Speech Corpus”, In Journal of Procedia computer science, vol. 131, Jan. 1, 2018, pp. 1269-1274. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/032459, Aug. 20, 2024, 13 pages. [cited by applicant]
“Vowel reduction”, Retrieved From. https://en.wikipedia.org/wiki/Vowei_reduction, Jan. 23, 2022, 6 Pages. [cited by applicant]
Georges, et al., “Towards an articulatory-driven neural vocoder for speech synthesis”, 12th International Seminar on Speech Production, Dec. 2020, 2 Pages. [cited by applicant]
International Preliminary Report on Patentability (Chapter I) received for PCT Application No. PCT/US2024/030184, mailed on Dec. 11, 2025, 11 pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/030184, Sep. 16, 2024, 19 pages. [cited by applicant]
Invitation to pay additional fees received for PCT Application No. PCT/US2024/030184, Jul. 26, 2024, 9 pages. [cited by applicant]
Koch, et al., “DigitalPhonetics / IMS-Toucan”, Retrieved From: httpsligithub.comi'DigitalPhoneticsiliVIS-Toucan, Mar. 23, 2023, 11 Pages. [cited by applicant]
Lux, et al., “Language-Agnostic Meta-Learning for Low-Resource Text-to-Speech with Articulatory Features”, In Repository of arXiv:2203.03191v1, Mar. 7, 2022, 12 Pages. [cited by applicant]
Murphy, James J. , “Demosthenes”, Retrieved From: https://www.britannica.com/biographyiDernosthenes-Greekstatesman-and-orator, May 7, 2022, 10 Pages. [cited by applicant]
Non-Final Office Action mailed on Jun. 4, 2025, in U.S. Appl. No. 18/326,631, 12 pages. [cited by applicant]
Notice of Allowance mailed on Sep. 30, 2025, in U.S. Appl. No. 18/326,631, 16 Pages. [cited by applicant]
Sedivy, Jul1e, “Mumbling Isn't a Sign of Laziness—It's a Clever Data-Compression Trick”, Retrieved From: https://nautil.usimumbling-isnt-a-sign-of-lazinessits-a-clever-data_compression-trick-235301,/, Feb. 18, 2015, 10 … [cited by applicant]
Shor, et al., “Project Euphonia's Personalized Speech Recognition for Non-Standard Speech”, Retrieved From: https://ai.googleblog.com/2019/08/project-euphonias-personalized-speech.html, Aug. 13, 2019, 7 Pages. [cited by applicant]
Son, et al., “An Acoustic Description of Consonant Reduction,” Speech Communication, vol. 28, Issue 2, Jun. 1, 1999, pp. 125-140. [cited by applicant]
Weirich, et al., “Mumbling: Macho or Morphology?”, In Journal of Speech, Language, and Hearing Research, vol. 59, Dec. 2016, pp. 1587-1595. [cited by applicant]