IP Library › Granted Patent US 11,887,582
Granted Patent B2
US 11,887,582 · App. 17/173,659 · Granted Jan 30, 2024

Training and testing utterance-based frameworks

Inventor: Daniel Bromand (Stockholm, SE)
Assignee: Spotify AB
G10L15/063G06F7/582G10L13/02G10L15/07G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,887,582
App. No.
17/173,659
Granted
Jan 30, 2024
Kind
B2
Abstract

Systems, methods, and devices for training and testing utterance based frameworks are disclosed. The training and testing can be conducting using synthetic utterance samples in addition to natural utterance samples. The synthetic utterance samples can be generated based on a vector space representation of natural utterances. In one method, a synthetic weight vector associated with a vector space is generated. An average representation of the vector space is added to the synthetic weight vector to form a synthetic feature vector. The synthetic feature vector is used to generate a synthetic voice sample. The synthetic voice sample is provided to the utterance-based framework as at least one of a testing or training sample.

Claims (65)

1. A method for training an audio voice framework, the method comprising:

generating a set of natural voice samples as voice training samples from one or more audio clips of the speech of one or more individuals;

generating a set of synthetic voice samples by playing one or more audio voice samples through a speaker to an appliance;

inputting the set of natural voice samples and the set of synthetic voice samples to an audio voice framework to produce test voice samples;

comparing the produced test voice samples against an expected test voice samples result;

if the produced test voice samples do not match the expected test samples result:

determining from the produced test voice samples a plurality of the input natural voice samples and synthetic voice samples that caused the audio voice framework to produce incorrect test voice samples output;

generating new synthetic voice samples by modifying one or more of the plurality of input natural voice samples and synthetic voice samples that caused the incorrect output; and

adding the generated new synthetic voice samples to the set of synthetic voice samples to be input to the audio voice framework to retrain the audio voice framework.

2. The method according to claim 1 , wherein the plurality set of natural voice samples include a first kind of voice and wherein the plurality of input natural voice samples that cause the audio voice framework to generate incorrect output are the second kind of voice.

3. The method according to claim 1 ,

wherein the plurality of natural voice samples include a first kind of voice and a second kind of voice, the first kind of voice having a first pitch characteristic and the second kind of voice having a second pitch characteristic;

wherein the second kind of voice has a pitch that is relatively lower than the first kind of voice; and

generating the new synthetic voice samples based on the plurality of natural voice samples having the second kind of voice.

4. The method according to claim 1 ,

wherein the plurality of natural voice samples include a first kind of voice and a second kind of voice, the first kind of voice having a first pitch characteristic and the second kind of voice having a second pitch characteristic;

wherein the second kind of voice has a pitch that is relatively higher than the first kind of voice; and

further comprising the step of generating the new synthetic voice samples using the plurality of natural voice samples having the second kind of voice.

5. The method according to claim 1 , wherein modifying one or more of the plurality of input natural voice samples and synthetic voice samples comprises:

modifying at least one characteristic of the plurality of input natural voice samples and synthetic voice samples.

6. The method according to claim 5 , wherein the at least one characteristic is: a duration characteristic, a pitch characteristic, a timbre characteristic, or a volume characteristic.

7. A system for training an audio voice framework, comprising:

one or more processors; and

a computer-readable storage medium coupled to the one or more processors and comprising instructions thereon that, when executed by the one or more processors, cause the one or more processors to:

generate a set of natural voice samples as voice training samples from one or more audio clips of the speech of one or more individuals;

generate a set of synthetic voice samples by playing one or more audio voice samples through a speaker to an appliance;

input the set of natural voice samples and the set of synthetic voice samples to an audio voice framework to produce test voice samples;

compare the produced test samples against an expected test voice samples result;

if the produced test voice samples do not match the expected test voice samples result:

determine from the produced test voice samples a plurality of the input natural voice samples and synthetic voice samples that caused the audio voice framework to produce incorrect test voice samples output;

generate new synthetic voice samples by modifying one or more of the plurality of input natural voice samples and synthetic voice samples that caused the incorrect output; and

add the generated new synthetic voice samples to the set of synthetic voice samples to be input to the audio voice framework to retrain the audio voice framework.

8. The system according to claim 7 , wherein the set of natural voice samples include a first kind of voice and wherein the plurality of input natural voice samples that cause the audio voice framework to generate incorrect output are the second kind of voice.

9. The system according to claim 7 ,

wherein the plurality of natural voice samples include a first kind of voice and a second kind of voice, the first kind of voice having a first pitch characteristic and the second kind of voice having a second pitch characteristic;

wherein the second kind of voice has a pitch that is relatively lower than the first kind of voice; and

wherein the computer-readable storage medium coupled to the one or more processors further comprises instructions thereon that, when executed by the one or more processors, cause the one or more processors to generate the new synthetic voice samples based on the plurality of natural voice samples having the second kind of voice.

10. The system according to claim 7 ,

wherein the plurality of natural voice samples include a first kind of voice and a second kind of voice, the first kind of voice having a first pitch characteristic and the second kind of voice having a second pitch characteristic;

wherein the second kind of voice has a pitch that is relatively higher than the first kind of voice; and

wherein the computer-readable storage medium coupled to the one or more processors further comprises instructions thereon that, when executed by the one or more processors, cause the one or more processors to generate the new synthetic voice samples using the plurality of natural voice samples having the second kind of voice.

11. The system according to claim 7 , wherein modifying one or more of the plurality of input natural voice samples and synthetic voice samples comprises:

modify at least one characteristic of the plurality of input natural voice samples and synthetic voice samples.

12. The system according to claim 11 , wherein the at least one characteristic is: a duration characteristic, a pitch characteristic, a timbre characteristic, or a volume characteristic.

13. A non-transitory computer readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to:

generate a set of natural voice samples as voice training samples from one or more audio clips of the speech of one or more individuals;

generate a set of synthetic voice samples by playing one or more audio voice samples through a speaker to an appliance;

input the set of natural voice samples and the set of synthetic voice samples to an audio voice framework to produce test voice samples;

compare the produced test samples against an expected test voice samples result;

if the produced test voice samples do not match the expected test voice samples result:

determine from the produced test voice samples a plurality of the input natural voice samples and synthetic voice samples that caused the audio voice framework to produce incorrect test voice samples output;

generate new synthetic voice samples by modifying one or more of the plurality of input natural voice samples and synthetic voice samples that caused the incorrect output; and

add the generated new synthetic voice samples to the set of synthetic voice samples to be input to the audio voice framework to retrain the audio voice framework.

14. The non-transitory computer-readable medium of claim 13 , wherein the set of natural voice samples include a first kind of voice and wherein the plurality of input natural voice samples that cause the audio voice framework to generate incorrect output are the second kind of voice.

15. The non-transitory computer-readable medium of claim 13 ,

wherein the plurality of natural voice samples include a first kind of voice and a second kind of voice, the first kind of voice having a first pitch characteristic and the second kind of voice having a second pitch characteristic;

wherein the second kind of voice has a pitch that is relatively lower than the first kind of voice; and

further comprising instructions thereon that, when executed by the one or more processors, cause the one or more processors to generate the new synthetic voice samples based on the plurality of natural voice samples having the second kind of voice.

16. The non-transitory computer-readable medium of claim 13 ,

wherein the plurality of natural voice samples include a first kind of voice and a second kind of voice, the first kind of voice having a first pitch characteristic and the second kind of voice having a second pitch characteristic;

wherein the second kind of voice has a pitch that is relatively higher than the first kind of voice; and

further comprising instructions thereon that, when executed by the one or more processors, cause the one or more processors to generate the new synthetic voice samples using the plurality of natural voice samples having the second kind of voice.

17. The non-transitory computer-readable medium of claim 13 , wherein modifying one or more of the plurality of input natural voice samples and synthetic voice samples comprises:

modify at least one characteristic of the plurality of input natural voice samples and synthetic voice samples.

18. The non-transitory computer-readable medium of claim 17 , wherein the at least one characteristic is: a duration characteristic, a pitch characteristic, a timbre characteristic, or a volume characteristic.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2021
From: BROMAND, DANIEL
To: SPOTIFY AB
Reel/Frame 055233/0498 →
Priority Claims (1)
EP 18167001 · Apr 12, 2018 · regional
Continuity (2)
Continuation 16168478 · Oct 23, 2018
Related Publication 20210174785A1 · Jun 10, 2021
Cited By (1)
US 12,555,567