IP Library Granted Patent US 12,375,624
Granted Patent B2
US 12,375,624 · App. 18/102,916 · Granted Jul 29, 2025

Accent conversion for virtual conferences

Inventors: Tuan Nam Nguyen (Karlsruhe, DE); Alexander Waibel (Sammamish, WA)
Assignee: Zoom Communications, Inc.
H04N7/157G10L15/063G10L21/007H04N7/147H04N7/152G10L2021/0135
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,375,624
App. No.
18/102,916
Granted
Jul 29, 2025
Kind
B2
Abstract

One example method includes receiving, during a virtual conference hosted by a virtual conference provider, a first audio stream comprising speech having first speech patterns according to a first accent, the first audio stream received from a first client device associated with a first participant in the virtual conference; generating, by a first trained machine learning (“ML”) model, a second audio stream comprising the speech having second speech patterns according to a second accent; and outputting the second audio stream.

Claims (53)

1. A method comprising:

receiving, during a virtual conference hosted by a virtual conference provider, a first audio stream comprising speech according to a first voice having first speech patterns according to a first accent, the first audio stream received from a first client device associated with a first participant in the virtual conference;

generating, by a first trained machine learning (“ML”) model, a second audio stream comprising the speech having second speech patterns according to a second accent and in the first voice, wherein the first trained ML model was trained according to pairs of training audio streams, each pair of training audio streams comprising:

a respective first training audio stream comprising first speech uttered by a first speaker having a first training voice having first identity characteristics, the respective first training audio stream having first training speech patterns according to a first training accent, the first speech comprising a first set of words; and

a respective third training audio stream generated by a second trained ML model based on a respective second training audio stream, the respective second training audio stream comprising second speech uttered by a second speaker having a second training voice having second identity characteristics, the second speech having second training speech patterns according to a second training accent, the respective third training audio stream comprising third speech according to the first identity characteristics and having the second training speech patterns according to the second training accent, the second speech comprising a second set of words; and

outputting the second audio stream.

2. The method of claim 1 , wherein the first speech patterns comprise pronunciation patterns, cadence, and prosody.

3. The method of claim 1 , wherein the first identity characteristics comprise a first timbre and first pitch and the second identity characteristics comprise a second timbre and second pitch, the first identity characteristics different from the second identity characteristics.

4. The method of claim 1 , wherein the first audio stream is provided by a client device associated with a participant, further comprising:

joining the client device to a virtual conference hosted by a virtual conference provider; and

receiving a request to convert the first audio stream to the second accent.

5. The method of claim 4 , wherein the request is received from the participant.

6. The method of claim 4 , wherein the request is received from another participant in the virtual conference.

7. The method of claim 6 , wherein the participant is not informed of the request.

8. The method of claim 4 , wherein generating the second audio stream is performed by the virtual conference provider, and the second audio stream is output to participants of the virtual conference other than the participant.

9. The method of claim 1 , further comprising:

receiving a third audio stream comprising second speech having third speech patterns according to a third accent;

generating, by the first trained ML model, a fourth audio stream comprising the second speech having fourth speech patterns according to a fourth accent; and

outputting the fourth audio stream.

10. The method of claim 1 , wherein the first audio stream is provided by a client device associated with a participant, further comprising:

joining the client device to a virtual conference hosted by a virtual conference provider;

receiving a third audio stream from a second client device associated with a second participant, the third audio stream comprising second speech having the second speech patterns according to the second accent;

determining that the first and second accents are different based on detecting the first accent in the first audio stream and detecting the second accent in the second audio stream; and

providing a suggestion to each of the client device and the second client device to perform accent conversion based on the first and second accents being different.

11. A system comprising:

a communications interface;

a non-transitory computer-readable medium; and

one or more processors communicatively coupled to the communications interface and the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:

receive, during a virtual conference hosted by a virtual conference provider, a first audio stream comprising speech according to a first voice having first speech patterns according to a first accent, the first audio stream received from a first client device associated with a first participant in the virtual conference;

generate, by a first trained machine learning (“ML”) model, a second audio stream comprising the speech having second speech patterns according to a second accent in the first voice, wherein the first trained ML model was trained according to pairs of training audio streams, each pair of training audio streams comprising:

a respective first training audio stream comprising first speech uttered by a first speaker having a first training voice having first identity characteristics, the respective first training audio stream having first training speech patterns according to a first training accent, the first speech comprising a first set of words; and

a respective third training audio stream generated by a second trained ML model based on a respective second training audio stream, the respective second training audio stream comprising second speech uttered by a second speaker having a second training voice having second identity characteristics, the second speech having second training speech patterns according to a second training accent, the respective third training audio stream comprising third speech according to the first identity characteristics and having the second training speech patterns according to the second training accent, the second speech comprising a second set of words; and

output the second audio stream.

12. The system of claim 11 , wherein the first speech patterns comprise pronunciation patterns, cadence, and prosody.

13. The system of claim 11 , wherein the first identity characteristics comprise a first timbre and first pitch and the second identity characteristics comprise a second timbre and second pitch, the first identity characteristics different from the second identity characteristics.

14. The system of claim 11 , wherein the first audio stream is provided by a client device associated with a participant, and wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

join the client device to a virtual conference hosted by a virtual conference provider; and

receive a request to convert the first audio stream to the second accent.

15. The system of claim 11 , wherein the first audio stream is provided by a client device associated with a participant, and wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

join the client device to a virtual conference hosted by a virtual conference provider;

receive a third audio stream from a second client device associated with a second participant, the third audio stream comprising second speech having the second speech patterns according to the second accent;

determine that the first and second accents are different based on detecting the first accent in the first audio stream and detecting the second accent in the second audio stream; and

provide a suggestion to each of the client device and the second client device to perform accent conversion based on the first and second accents being different.

16. A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to:

receive, during a virtual conference hosted by a virtual conference provider, a first audio stream comprising speech according to a first voice having first speech patterns according to a first accent, the first audio stream received from a first client device associated with a first participant in the virtual conference;

generate, by a first trained machine learning (“ML”) model, a second audio stream comprising the speech having second speech patterns according to a second accent in the first voice, wherein the first trained ML model was trained according to pairs of training audio streams, each pair of training audio streams comprising:

a respective first training audio stream comprising first speech uttered by a first speaker having a first training voice having first identity characteristics, the respective first training audio stream having first training speech patterns according to a first training accent, the first speech comprising a first set of words; and

a respective third training audio stream generated by a second trained ML model based on a respective second training audio stream, the respective second training audio stream comprising second speech uttered by a second speaker having a second training voice having second identity characteristics, the second speech having second training speech patterns according to a second training accent, the respective third training audio stream comprising third speech according to the first identity characteristics and having the second training speech patterns according to the second training accent, the second speech comprising a second set of words; and

output the second audio stream.

17. The non-transitory computer-readable medium of claim 16 , wherein the first audio stream is provided by a client device associated with a participant, and further comprising processor-executable instructions configured to cause one or more processors to:

join the client device to a virtual conference hosted by a virtual conference provider; and

receive a request to convert the first audio stream to the second accent.

18. The non-transitory computer-readable medium of claim 17 , wherein generating the second audio stream is performed by the virtual conference provider, and further comprising processor-executable instructions configured to cause one or more processors to output the second audio stream to participants of the virtual conference other than the participant.

Assignments (2)
CHANGE OF NAME Recorded Jul 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 071831/0502 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2023
From: NGUYEN, TUAM NAM; WAIBEL, ALEXANDER
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 062533/0327 →
Continuity (2)
Provisional Application 63406942 · Sep 15, 2022
Related Publication 20240098218A1 · Mar 21, 2024
References Cited (27)
US 11134217B1 · Goel · 2021 [cited by examiner]
US 20110282669A1 · Michaelis · 2011 [cited by examiner]
US 20180174595A1 · Dirac · 2018 [cited by applicant]
US 20220122579A1 · Biadsy · 2022 [cited by examiner]
US 20220358903A1 · Serebryakov · 2022 [cited by examiner]
US 20220415340A1 · Barhate · 2022 [cited by examiner]
US 20230223011A1 · Golman · 2023 [cited by applicant]
US 20240355346A1 · Lubin · 2024 [cited by examiner]
“Github v0.2.1”, Available Online at: https://github.com/coqui-ai/TTS/releases/tag/v0.2.1, Aug. 31, 2021, pp. 1-2. [cited by applicant]
Casanova, et al., “YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for Everyone”, Proceeding of the 39th International Conference on Machine Learning, Apr. 30, 2023, 9 pages. [cited by applicant]
Kim, et al., “Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech”, Proceedings of the 38th International Conference on Machine Learning, Jun. 11, 2021, 15 pages. [cited by applicant]
Kim, et al., “Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search”, Available Online at: https://arxiv.org/abs/2005.11129, Oct. 23, 2020, pp. 1-14. [cited by applicant]
Kolluru, et al., “Generating Multiple-accent Pronunciations for TTS using Joint Sequence Model Interpolation”, Interspeech, Sep. 14-18, 2014, pp. 1273-1277. [cited by applicant]
Kong, et al., “HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis”, 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Oct. 23, 2020, pp. 1-14. [cited by applicant]
Nguyen, et al., “Accent Conversion using Pre-trained Model and Synthesized Data from Voice Conversion”, Interspeech, Sep. 18-22, 2022, pp. 2583-2587. [cited by applicant]
Nguyen, et al., “KIT's IWSLT 2021 Offline Speech Translation System”, Proceedings of the 18th International Conference on Spoken Language Translation, Aug. 5-6, 2021, pp. 125-130. [cited by applicant]
Pham, et al., “Adaptive Multillingual Speech Recognition with Pretrained Models”, Interspeech, Available Online at: https://arxiv.org/abs/2205.12304, May 24, 2022, 5 pages. [cited by applicant]
Pham, et al., “Efficient Weight Factorization for Multilingual Speech Recognition”, Available Online at: https://arxiv.org/abs/2105.03010, May 7, 2021, 5 pages. [cited by applicant]
Pham, et al., “KIT's IWSLT 2020 SLT Translation System”, Proceedings of the 17th International Conference on Spoken Language Translation, Jul. 9-10, 2020, pp. 55-61. [cited by applicant]
Ren, et al., “FastSpeech: Fast, Robust and Controllable Text to Speech”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Nov. 20, 2019, pp. 1-13. [cited by applicant]
Vaswani, et al., “Attention is All you Need”, 31st Conference on Neural Information Processing Systems (NIPS 2017), Dec. 6, 2017, pp. 1-15. [cited by applicant]
Zhao, et al., “L2-ARCTIC: A Non-Native English Speech Corpus”, Interspeech, Sep. 2018, 7 pages. [cited by applicant]
Kumar, et al., “Learning Robust Latent Representations for Controllable Speech Synthesis”, Cornell University Library, May 10, 2021, 10 pages. [cited by applicant]
Li, et al., “Neural Speech Synthesis with Transformer Network”, Proceedings of the AAAI Conference on Artificial Intelligence, Sep. 1, 2018, 8 pages. [cited by applicant]
Melechovsky, et al., “Learning Accent Representation with Multi-Level VAE Towards Controllable Speech Synthesis”, Institute of Electrical and Electronics Engineers Spoken Language Technology Workshop, Jan. 9, 2023, pp. … [cited by applicant]
Nguyen, et al., “Syntacc : Synthesizing Multi-Accent Speech by Weight Factorization”, Institute of Electrical and Electronics Engineers International Conference on Acoustics, Jun. 4, 2023, 5 pages. [cited by applicant]
PCT/US2024/028469, “International Search Report and Written Opinion”, Aug. 26, 2024, 13 pages. [cited by applicant]