IP Library › Granted Patent US 12,417,756
Granted Patent B2
US 12,417,756 · App. 19/027,799 · Granted Sep 16, 2025

Systems and methods for real-time accent mimicking

Inventors: Ankita Jha (Bangalore, IN); Lukas Pfeifenberger (Salzburg, AT); Piotr Dura (Warsaw, PL); David Braude (Edinburgh, GB); Alvaro Escudero (San Sebastian de los Reyes, ES); Shawn Zhang (Palo Alto, CA); Maxim Serebryakov (Palo Alto, CA); Sharath Kashava Narayana (Palo Alto, CA)
Assignee: Sanas.ai Inc.
G10L13/027G10L13/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,756
App. No.
19/027,799
Granted
Sep 16, 2025
Kind
B2
Abstract

The disclosed technology relates to methods, speech processing systems, and non-transitory computer readable media for real-time accent mimicking. In some examples, trained machine learning model(s) are applied to first input audio data to extract accent features of first input speech associated with a first accent of a first user. Obtained second input data associated with second input speech associated with a second accent of a second user is analyzed to generate characteristics specific to a natural voice of the second user. A modified version of the second input speech is synthesized based on the generated characteristics and the extracted accent features. The modified version of the second input speech advantageously preserves aspects of the natural voice of the second user and mimics the first accent. Output audio data generated based on the modified version of the second input speech is provided for output via an audio output device.

Claims (32)

1. A speech processing system, comprising an audio interface coupled to a microphone and an audio output device, memory having instructions stored thereon, and one or more processors coupled to the memory and the audio interface and configured to execute the instructions to:

apply one or more trained machine learning models to first input audio data obtained via the microphone and the audio interface to extract accent features of first input speech associated with a first accent of a first user;

analyze obtained second input audio data associated with second input speech associated with a second accent of a second user to generate characteristics specific to a natural voice of the second user, wherein the generated characteristics correspond to vocal traits that are distinct to the second user and comprise one or more of a voice quality or one or more phonetic patterns, prosodic features, articulation styles, or intonation patterns;

synthesize a modified version of the second input speech by modifying the obtained second input audio data based on the generated characteristics and the extracted accent features, wherein the modified version of the second input speech preserves aspects of the natural voice of the second user and mimics the first accent; and

provide to the audio interface output audio data for output via the audio output device, wherein the output audio data is generated based on the modified version of the second input speech.

2. The speech processing system of claim 1 , wherein the one or more processors are further configured to execute the instructions to extract from the first input audio data one or more prosodic features, linguistic features, or global speaker characteristics.

3. The speech processing system of claim 1 , wherein the accent features comprise one or more pitch contours, other intonation patterns, or phoneme pronunciations and the pitch contours comprise variations in pitch throughout the first input speech, the other intonation patterns comprise the rise and fall of pitch at the ends of phrases or sentences, or the phoneme pronunciations comprise a unique production of phonemes in the first accent.

4. The speech processing system of claim 1 , wherein the one or more processors are further configured to execute the instructions to apply a mel frequency cepstral coefficient (MFCC) analysis to extract a unique fingerprint of the voice of the second user, wherein the generated characteristics comprise the unique fingerprint.

5. The speech processing system of claim 1 , wherein the one or more processors are further configured to execute the instructions to apply a speaker identity encoding technique to encode speaker-specific voice characteristics, wherein the generated characteristics comprise the speaker-specific voice characteristics.

6. The speech processing system of claim 1 , wherein the one or more processors are further configured to execute the instructions to receive the second input audio data via one or more communication networks and from a user computing device that is remote from the speech processing system, wherein the second input audio data is captured at the user computing device.

7. A method for real-time accent mimicking, the method implemented by a speech processing system and comprising:

applying one or more trained machine learning models to first input audio data to extract accent features of first input speech associated with a first accent of a first user;

analyzing obtained second input audio data associated with second input speech associated with a second accent of a second user to generate characteristics specific to a natural voice of the second user, wherein the generated characteristics comprise a unique fingerprint of the voice of the second user;

synthesizing a modified version of the second input speech by modifying the obtained second input audio data based on the generated characteristics and the extracted accent features; and

providing output audio data generated based on the modified version of the second input speech.

8. The method of claim 7 , wherein the modified version of the second input speech preserves aspects of the natural voice of the second user and mimics the first accent.

9. The method of claim 7 , further comprising extracting from the first input audio data one or more prosodic features, linguistic features, or global speaker characteristics.

10. The method of claim 7 , wherein the accent features comprise one or more pitch contours, intonation patterns, or phoneme pronunciations and the pitch contours comprise variations in pitch throughout the first input speech, the intonation patterns comprise the rise and fall of pitch at the ends of phrases or sentences, or the phoneme pronunciations comprise a unique production of phonemes in the first accent.

11. The method of claim 7 , further comprising applying a mel frequency cepstral coefficient (MFCC) analysis to extract the unique fingerprint of the voice of the second user.

12. The method of claim 7 , further comprising applying a speaker identity encoding technique to encode speaker-specific voice characteristics, wherein the generated characteristics comprise the speaker-specific voice characteristics.

13. The method of claim 7 , further comprising receiving the second input audio data via one or more communication networks and from a user computing device that is remote from the speech processing system, wherein the second input audio data is captured at the user computing device.

14. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to:

apply one or more trained machine learning models to first input audio data obtained via a microphone to extract accent features of first input speech associated with a first accent of a first user;

analyze obtained second input audio data associated with second input speech associated with a second accent of a second user to generate characteristics specific to a natural voice of the second user, wherein the generated characteristics comprise speaker-specific voice characteristics;

synthesize a modified version of the second input speech by modifying the obtained second input audio data based on the generated characteristics and the extracted accent features; and

provide for output via an audio output device output audio data generated based on the modified version of the second input speech.

15. The non-transitory computer-readable medium of claim 14 , wherein the modified version of the second input speech preserves aspects of the natural voice of the second user and mimics the first accent.

16. The non-transitory computer-readable medium of claim 14 , wherein the instructions, when executed by the at least one processor further causes the at least one processor to extract from the first input audio data one or more prosodic features, linguistic features, or global speaker characteristics.

17. The non-transitory computer-readable medium of claim 14 , wherein the accent features comprise one or more pitch contours, intonation patterns, or phoneme pronunciations and the pitch contours comprise variations in pitch throughout the first input speech, the intonation patterns comprise the rise and fall of pitch at the ends of phrases or sentences, or the phoneme pronunciations comprise a unique production of phonemes in the first accent.

18. The non-transitory computer-readable medium of claim 14 , wherein the instructions, when executed by the at least one processor further causes the at least one processor to apply a mel frequency cepstral coefficient (MFCC) analysis to extract a unique fingerprint of the voice of the second user, wherein the generated characteristics comprise the unique fingerprint.

19. The non-transitory computer-readable medium of claim 14 , wherein the instructions, when executed by the at least one processor further causes the at least one processor to apply a speaker identity encoding technique to encode the speaker-specific voice characteristics.

20. The non-transitory computer-readable medium of claim 14 , wherein the instructions, when executed by the at least one processor further causes the at least one processor to receive the second input audio data via one or more communication networks and from a user computing device that is remote from the speech processing system, wherein the second input audio data is captured at the user computing device.

Assignments (4)
SECURITY INTEREST Recorded Apr 17, 2026
From: SANAS.AI INC.
To: CANADIAN IMPERIAL BANK OF COMMERCE
Reel/Frame 074405/0229 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2026
From: BRAUDE, DAVID
To: SANAS.AI INC.
Reel/Frame 074253/0731 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2025
From: NARAYANA, SHARATH KASHAVA
To: SANAS.AI INC.
Reel/Frame 071924/0929 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2025
From: JHA, ANKITA; PFEIFENBERGER, LUKAS; DURA, PIOTR; ESCUDERO, ALVARO; ZHANG, SHAWN; SEREBRYAKOV, MAXIM
To: SANAS.AI INC.
Reel/Frame 069963/0080 →
Continuity (2)
Provisional Application 63678180 · Aug 1, 2024
Related Publication 20250166603A1 · May 22, 2025
References Cited (31)
US 9129602B1 · Shepard · 2015 [cited by examiner]
US 9715873B2 · Graham · 2017 [cited by examiner]
US 11120219B2 · Alloh · 2021 [cited by examiner]
US 11134217B1 · Goel · 2021 [cited by examiner]
US 11600284B2 · Pearson · 2023 [cited by examiner]
US 11741965B1 · Opp · 2023 [cited by examiner]
US 12039975B2 · Krishnan · 2024 [cited by examiner]
US 12087270B1 · Cygert · 2024 [cited by examiner]
US 12243511B1 · Joly · 2025 [cited by examiner]
US 20040225499A1 · Wang · 2004 [cited by examiner]
US 20050234727A1 · Chiu · 2005 [cited by examiner]
US 20080205629A1 · Basson · 2008 [cited by examiner]
US 20080269958A1 · Filev · 2008 [cited by examiner]
US 20140187210A1 · Chang · 2014 [cited by examiner]
US 20160140952A1 · Graham · 2016 [cited by examiner]
US 20160275952A1 · Kashtan · 2016 [cited by examiner]
US 20170169814A1 · Pashine · 2017 [cited by examiner]
US 20180146370A1 · Krishnaswamy · 2018 [cited by examiner]
US 20180203847A1 · Akkiraju · 2018 [cited by examiner]
US 20200004820A1 · Chhaya · 2020 [cited by examiner]
US 20200193971A1 · Feinauer · 2020 [cited by examiner]
US 20210098013A1 · Day · 2021 [cited by examiner]
US 20210217431A1 · Pearson · 2021 [cited by examiner]
US 20220093094A1 · Krishnan · 2022 [cited by examiner]
US 20220131973A1 · Talib · 2022 [cited by examiner]
US 20230267941A1 · Nagpal · 2023 [cited by examiner]
US 20230335123A1 · Liu · 2023 [cited by examiner]
US 20240146560A1 · Swerdlow · 2024 [cited by examiner]
US 20240161764A1 · Maikhuri · 2024 [cited by examiner]
US 20240221719A1 · Kothari · 2024 [cited by examiner]
US 20250118286A1 · Badlani · 2025 [cited by examiner]