IP Library Granted Patent US 11,600,284
Granted Patent B2
US 11,600,284 · App. 16/740,440 · Granted Mar 7, 2023

Voice morphing apparatus having adjustable parameters

Inventor: Steve Pearson (Felton, CA)
Assignee: SOUNDHOUND, INC.
G10L21/013G06N3/08G06N20/00G10L21/0208G10L2021/0135
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,600,284
App. No.
16/740,440
Granted
Mar 7, 2023
Kind
B2
Abstract

A voice morphing apparatus having adjustable parameters is described. The disclosed system and method include a voice morphing apparatus that morphs input audio to mask a speaker's identity. Parameter adjustment uses evaluation of an objective function that is based on the input audio and output of the voice morphing apparatus. The voice morphing apparatus includes objectives that are based adversarially on speaker identification and positively on audio fidelity. Thus, the voice morphing apparatus is adjusted to reduce identifiability of speakers while maintaining fidelity of the morphed audio. The voice morphing apparatus may be used as part of an automatic speech recognition system.

Claims (24)

1. A voice morphing apparatus comprising:

a neural network architecture to map input audio data to output audio data, the input audio data comprising a representation of speech from a speaker, the neural network architecture including a set of parameters, the set of parameters being trained to maximize a speaker identification distance from the input audio data to a set of speaker identification vectors and to optimize a speaker intelligibility score for the output audio data.

2. The voice morphing apparatus of claim 1 further comprising a noise filter to pre-process the input audio data.

3. The voice morphing apparatus of claim 2 , wherein the noise filter removes a noise component from the input audio data and the voice morphing apparatus adds the noise component to the set of speaker identification vectors from the neural network architecture.

4. The voice morphing apparatus of claim 1 , wherein the neural network architecture comprises one or more recurrent connections.

5. The voice morphing apparatus of claim 1 , wherein the voice morphing apparatus is configured to output time-series audio waveform data based on the set of speaker identification vectors from the neural network architecture.

6. A non-transitory computer-readable storage medium for storing instructions that, when executed by at least one processor, cause the at least one processor to:

load input audio data from a data source;

input the input audio data to a voice morphing apparatus, the voice morphing apparatus including a set of trainable parameters;

process the input audio data using the voice morphing apparatus to generate morphed audio data;

apply a speaker identification system to at least the morphed audio data to output a measure of speaker identification;

apply an audio fidelity system to the morphed audio data and the input audio data to output a measure of audio fidelity;

evaluate an objective function based on the measure of speaker identification and the measure of audio fidelity; and

adjust the set of trainable parameters for the voice morphing apparatus based on a gradient of the objective function,

wherein the objective function is configured to adjust the set of trainable parameters to optimize the measure of audio fidelity between the morphed audio data and the input audio data and to reduce the measure of speaker identification while maintaining speech intelligibility.

7. A method for optimizing training parameters, the method comprising:

loading input audio data from a data source;

inputting the input audio data to a voice morphing apparatus, the voice morphing apparatus including a set of trainable parameters;

processing the input audio data using the voice morphing apparatus to generate morphed audio data;

applying a speaker identification system to at least the morphed audio data to output a measure of speaker identification;

applying an audio fidelity system to the morphed audio data and the input audio data to output a measure of audio fidelity;

evaluating an objective function based on the measure of speaker identification and the measure of audio fidelity; and

adjusting the set of trainable parameters for the voice morphing apparatus based on a gradient of the objective function,

wherein the objective function is configured to adjust the set of trainable parameters to optimize the measure of audio fidelity between the morphed audio data and the input audio data and to reduce the measure of speaker identification while maintaining speech intelligibility.

Assignments (7)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2020
From: PEARSON, STEVE
To: SOUNDHOUND, INC.
Reel/Frame 052266/0463 →