IP Library › Granted Patent US 12,482,476
Granted Patent B2
US 12,482,476 · App. 18/053,886 · Granted Nov 25, 2025

Accent personalization for speakers and listeners

Inventors: Ajay Maikhuri (Bangalore, IN); Dhilip Kumar (Bangalore, IN)
Assignee: Dell Products L.P.
G10L21/007G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,482,476
App. No.
18/053,886
Filed
Nov 9, 2022
Granted
Nov 25, 2025
Kind
B2
Art Unit
2658
USPC
704/278
Abstract

In one aspect, an example methodology implementing the disclosed techniques includes, by a computing device, receiving audio data corresponding to a spoken utterance by a first user and determining an accent of the audio data. The method also includes, by the computing device, neutralizing the accent of the audio data to a preconfigured accent and transmitting a modified audio data in the preconfigured accent to another computing device. The modified audio data includes the spoken utterance by the first user.

Claims (48)

1 . A method comprising:

providing a service on a computing device and an application on another computing device;

receiving, by a computing device, audio data corresponding to a spoken utterance by a first user;

determining, by the computing device, an accent of the audio data using an accent classification module, the accent classification module comprising a convolution layer, a maximum pooling layer, and a softmax layer;

selecting a neutralization model from a plurality of models, each of the plurality of models trained to neutralize a different accent;

neutralizing, by the computing device, the accent of the audio data to generate a modified audio data in a preconfigured accent using the selected neutralization model;

transmitting, by the computing device, the modified audio data in the preconfigured accent to the another computing device being used by a second user;

personalizing, by the application, the modified audio data to a first accent selected by the second user;

personalizing, by the application, the voice of the spoken utterance to a second voice selected by the second user; and

outputting, by the application, the cloned voice to an audio playback device.

2 . The method of claim 1 , wherein the preconfigured accent is a native English accent.

3 . The method of claim 1 , wherein the accent of the audio data is a non-native English accent.

4 . The method of claim 1 , wherein the accent of the audio data is a native English accent.

5 . The method of claim 1 , wherein the accent of the second user is a native English accent.

6 . The method of claim 1 , wherein the accent of the second user is a non-native English accent.

7 . The method of claim 1 , further comprising:

wherein the second voice is selected to be one of a voice of the first user and a voice of the second user.

8 . The method of claim 7 , wherein the second voice is a cloning of the voice of the second user.

9 . A computing device comprising:

one or more non-transitory machine-readable mediums configured to store instructions; and

one or more processors configured to execute the instructions stored on the one or more non-transitory machine-readable mediums, wherein execution of the instructions causes the one or more processors to carry out a process comprising:

providing a service on a computing device and an application on another computing device;

receiving audio data corresponding to a spoken utterance by a first user;

determining, by the computing device, an accent of the audio data using an accent classification module, the accent classification module comprising a convolution layer, a maximum pooling layer, and a softmax layer;

selecting a neutralization model from a plurality of models, each of the plurality of models trained to neutralize a different accent;

neutralizing, by the computing device, the accent of the audio data to generate a modified audio data in a preconfigured accent using the selected neutralization model;

transmitting, by the computing device, the modified audio data in the preconfigured accent to the another computing device being used by a second user

personalizing, by the application, the modified audio data to a first accent selected by the second user;

personalizing, by the application, the voice of the spoken utterance to a second voice selected by the second user; and

outputting, by the application, the cloned voice to an audio playback device.

10 . The computing device of claim 9 , wherein the preconfigured accent is a native English accent.

11 . The computing device of claim 9 , wherein the accent of the audio data is a non-native English accent.

12 . The computing device of claim 9 , wherein the accent of the audio data is a native English accent.

13 . The computing device of claim 9 , wherein the accent of the second user is a native English accent.

14 . The computing device of claim 9 , wherein the accent of the second user is a non-native English accent.

15 . The computing device of claim 9 , wherein the second voice is selected to be one of a voice of the first user and a voice of the second user.

16 . The computing device of claim 15 , wherein the second voice is a cloning of the voice of the second user.

17 . The computing device of claim 15 , wherein the another voice is a voice specified by the second user.

18 . A non-transitory machine-readable medium encoding instructions that when executed by one or more processors cause a process to be carried out, the process including:

providing a service on a computing device and an application on another computing device;

receiving audio data corresponding to a spoken utterance by a first user;

determining an accent of the audio data using an accent classification module, the accent classification module comprising a convolution layer, a maximum pooling layer, and a softmax layer;

selecting a neutralization model from a plurality of models, each of the plurality of models trained to neutralize a different accent;

neutralizing the accent of the audio data to generate a modified audio data in a preconfigured accent using the selected neutralization model;

transmitting the modified audio data in the preconfigured accent to the another computing device being used by a second user;

personalizing the modified audio data to a first accent selected by the second user;

personalizing the voice of the spoken utterance to a second voice selected by the second user; and

outputting the cloned voice to an audio playback device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 29, 2022
From: MAIKHURI, AJAY; KUMAR, DHILIP
To: DELL PRODUCTS L.P.
Reel/Frame 061901/0005 →
Continuity (1)
Related Publication 20240161764A1 · May 16, 2024
References Cited (9)
US 20180174595A1 · Dirac · 2018 [cited by examiner]
US 20220044688A1 · Perret · 2022 [cited by examiner]
US 20220358903A1 · Serebryakov · 2022 [cited by examiner]
US 20220415340A1 · Barhate · 2022 [cited by examiner]
US 20230223006A1 · Fan · 2023 [cited by examiner]
US 20230352001A1 · Carmiel · 2023 [cited by examiner]
US 20240098218A1 · Nguyen · 2024 [cited by examiner]
Ding, S. (2021). Few-Shot Voice and Foreign Accent Conversion and Its Applications in Pronunciation Training (Doctoral dissertation). [cited by examiner]
Grigoriadis, S. (2019). Convolutional neural networks for accent classification. [cited by examiner]