IP Library Granted Patent US 12,505,830
Granted Patent B2
US 12,505,830 · App. 18/046,137 · Granted Dec 23, 2025

Automatic speech recognition with voice personalization and generalization

Inventor: Keyvan Mohajer (Los Gatos, CA)
Assignee: SOUNDHOUND AI IP, LLC
G10L15/18G10L15/063G10L2021/0135
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,830
App. No.
18/046,137
Granted
Dec 23, 2025
Kind
B2
Abstract

A voice morphing model can transform diverse voices to one or a small number of target voices. Speech recognition on diverse voices can be performed by morphing it to a target voice and then performing recognition on audio with the target voice. A source of requests for speech recognition can pass audio and a voiceprint with requests. Speech recognition can run with improved accuracy by biasing an acoustic model for the voice in the audio using the voiceprint. The audio can be used to calculate a new voiceprint, which can be used to update the voiceprint included with the audio. The updated voiceprint can be sent back to the source and then used with future speech recognition requests.

Claims (33)

1 . A computer-implemented method of training an acoustic model, the method comprising:

obtaining a voiceprint calculator that calculates a score for the distance between a voice in speech audio and a target voice;

training a voice morphing model to morph speech audio to the target voice, the training using speech audio of multiple distinct voices with a loss function dependent on the score;

training an acoustic model on transcribed speech in the target voice; and

tuning the voice morphing model and acoustic model by backpropagation of error reduction based on a measurement of the error rate of phoneme inference,

wherein the acoustic model can infer phonemes from audio morphed by the voice morphing model.

2 . The method of claim 1 wherein the transcribed speech in the target voice is from a single speaker without morphing.

3 . The method of claim 1 wherein the transcribed speech in the target voice is generated by morphing speech audio of multiple distinct voices.

4 . The method of claim 3 further comprising:

finetuning the voice morphing model with a loss function dependent on an error rate of the acoustic model when run on the morphed audio of transcribed speech.

5 . The method of claim 1 further comprising tuning the voice morphing model while keeping the acoustic model fixed.

6 . The method of claim 1 further comprising tuning the acoustic model while keeping the voice morphing model fixed.

7 . The method of claim 1 further comprising measuring the amount of noise in the morphed speech audio, wherein the loss function further depends on the amount of noise.

8 . A computer implemented method of phoneme inference, the method comprising:

calculating a plurality of scores for the distances between a voice in speech audio from multiple distinct voices and a target voice;

training a voice morphing model to morph speech audio to the target voice, the training using speech audio of the multiple distinct voices with a loss function dependent on the scores;

morphing audio of sampled speech to a target voice using the voice morphing model to generate morphed audio; and

inferring a sequence of phonemes from the morphed audio using an acoustic model,

wherein the acoustic model has an accuracy bias in favor of the target voice.

9 . The method of claim 8 wherein the acoustic model is conditioned by a choice of the target voice from among a plurality of target voices.

10 . A computer-implemented method of training an acoustic model, the method comprising:

obtaining a voiceprint calculator that calculates a plurality of scores for the distances between each voice of multiple distinct voices in speech audio and a target voice;

training a voice morphing model to morph speech audio to the target voice, the training using speech audio of the multiple distinct voices with a loss function dependent on the scores; and

training an acoustic model on transcribed speech in the target voice,

wherein the acoustic model can infer phonemes from audio morphed by the voice morphing model.

11 . The method of claim 1 wherein the transcribed speech in the target voice is from a single speaker without morphing.

12 . The method of claim 1 wherein the transcribed speech in the target voice is generated by morphing speech audio of multiple distinct voices.

13 . The method of claim 12 further comprising:

finetuning the voice morphing model with a loss function dependent on an error rate of the acoustic model when run on the morphed audio of transcribed speech.

14 . The method of claim 1 further comprising tuning the voice morphing model while keeping the acoustic model fixed.

15 . The method of claim 1 further comprising tuning the acoustic model while keeping the voice morphing model fixed.

16 . The method of claim 1 further comprising tuning the voice morphing model and acoustic model by backpropagation of error reduction based on a measurement of the error rate of phoneme inference.

17 . The method of claim 1 further comprising measuring the amount of noise in the morphed speech audio, wherein the loss function further depends on the amount of noise.

Assignments (5)
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2022
From: MOHAJER, KEYVAN
To: SOUNDHOUND, INC.
Reel/Frame 061407/0566 →
Continuity (1)
Related Publication 20240127803A1 · Apr 18, 2024
References Cited (13)
US 10068565B2 · Yassa · 2018 [cited by examiner]
US 10431236B2 · Gloge · 2019 [cited by examiner]
US 10726828B2 · Fukuda · 2020 [cited by examiner]
US 11270721B2 · Lu · 2022 [cited by examiner]
US 20180247640A1 · Yassa · 2018 [cited by examiner]
US 20190180759A1 · Fontaine · 2019 [cited by examiner]
US 20210089626A1 · Ross · 2021 [cited by examiner]
US 20210304769A1 · Ye · 2021 [cited by examiner]
US 20240105203A1 · Bittner · 2024 [cited by examiner]
WO WO2013000868A1 · 2013 [cited by examiner]
Serizel, Romain, and Diego Giuliani. “Deep-neural network approaches for speech recognition with heterogeneous groups of speakers including children.” Natural Language Engineering 23.3 (2017): 325-350. (Year: 2017). [cited by examiner]
Serizel, Romain, and Diego Giuliani. “Vocal tract length normalisation approaches to DNN-based children's and adults' speech recognition.” 2014 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2014. (Year: 2014). [cited by examiner]
Sekii, Yusuke, et al. “Fast many-to-one voice conversion using autoencoders.” International Conference on Agents and Artificial Intelligence. vol. 2. SCITEPRESS, 2017. (Year: 2017). [cited by examiner]