IP Library Granted Patent US 11,308,938
Granted Patent B2
US 11,308,938 · App. 16/704,216 · Granted Apr 19, 2022

Synthesizing speech recognition training data

Inventors: Maisy Wieman (Boulder, CO); Jonah Probell (Alviso, CA); Sudharsan Krishnaswamy (San Jose, CA)
Assignee: SoundHound, Inc.
G10L15/063G10L13/02G10L15/16G10L15/187G10L15/1815G10L15/197G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,308,938
App. No.
16/704,216
Granted
Apr 19, 2022
Kind
B2
Abstract

To train a speech recognizer, such as for recognizing variables in a neural speech-to-meaning system, compute, within an embedding space, a range of vectors of features of natural speech. Generate parameter sets for speech synthesis and synthesis speech according to the parameters. Analyze the synthesized speech to compute vectors in the embedding space. Using a cost function that favors an even spread (minimal clustering) generates a multiplicity of speech synthesis parameter sets. Using the multiplicity of parameter sets, generate a multiplicity of speech of known words that can be used as training data for speech recognition.

Claims (25)

1. A method of training a speech-to-meaning model, the method comprising:

determining a multiplicity of words that may be values of a variable in a phrasing of an intent;

determining a multiplicity of parameter sets representative of the voices of users of a virtual assistant, the parameter sets being chosen with an approximately even distribution within a range in an embedding space defined by vectors computed from recorded speech;

synthesizing a multiplicity of speech audio segments for the multiplicity of words, the segments being synthesized according to the multiplicity of parameter sets; and

training, using the synthesized speech audio segments, a variable recognizer that is able to compute a probability of the presence of any of the multiplicity of words in speech audio.

2. The method of claim 1 further comprising:

training an intent recognizer on segments of speech audio of a phrasing, wherein an input to the intent recognizer is a probability output from the variable recognizer.

3. The method of claim 2 further comprising:

determining a context of the variable within the phrasing with respect to adjacent phonetic information,

wherein the context is a further parameter of the speech synthesis.

4. The method of claim 2 further comprising:

determining a context of the variable within the phrasing with respect to emphasis,

wherein the context is a further parameter of the speech synthesis.

5. A method of computing a plurality of speech synthesis parameter sets representative of a diversity of voices, the method comprising:

procuring a multiplicity of speech audio recordings of natural people representative of the diversity of voices;

analyzing the recordings to compute recorded speech vectors within an embedding space of voice features;

computing a region representing a range of recorded speech vectors within the embedding space; and

learning the plurality of speech synthesis parameter sets by gradient descent according to a loss function computed by:

synthesizing speech segments according to parameter sets in the plurality of parameter sets;

analyzing the synthesized speech segments to compute synthesized speech vectors in the space; and

computing a loss in proportion to a clustering of the synthesized speech vectors within the space.

6. The method of claim 5 wherein the embedding space is learned.

7. The method of claim 5 further comprising:

synthesizing segments of speech of an enumerated word according to the speech synthesis parameter sets; and

training a variable recognizer, wherein the training data includes the synthesized segments of speech of the enumerated word.

Assignments (7)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2019
From: WIEMAN, MAISY; PROBELL, JONAH; KRISHNASWAMY, SUDHARSAN
To: SOUNDHOUND, INC.
Reel/Frame 051195/0662 →
Cited By (1)
US 12,265,576