IP Library › Granted Patent US 12,555,567
Granted Patent B2
US 12,555,567 · App. 18/530,702 · Granted Feb 17, 2026

Training and testing audio voice frameworks

Inventor: Daniel Bromand (Stockholm, SE)
Assignee: Spotify AB
G10L15/063G06F7/582G10L13/02G10L15/07G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,567
App. No.
18/530,702
Granted
Feb 17, 2026
Kind
B2
Abstract

Systems, methods, and devices for training and testing utterance based frameworks are disclosed. The training and testing can be conducting using synthetic utterance samples in addition to natural utterance samples. The synthetic utterance samples can be generated based on a vector space representation of natural utterances. In one method, a synthetic weight vector associated with a vector space is generated. An average representation of the vector space is added to the synthetic weight vector to form a synthetic feature vector. The synthetic feature vector is used to generate a synthetic voice sample. The synthetic voice sample is provided to the utterance-based framework as at least one of a testing or training sample.

Claims (56)

1 . A method for training an audio voice framework, the method comprising:

obtaining a first natural weight vector of a vector space and a second natural weight vector of the vector space, each respectively associated with natural voice samples, wherein the second natural weight vector is associated with a higher confidence than the first natural weight vector;

generating a synthetic weight vector by way of applying a genetic algorithm to the first natural weight vector and the second natural weight vector, wherein the genetic algorithm selects the synthetic weight vector based on a fitness function applied to the first natural weight vector and the second natural weight vector;

forming a synthetic feature vector based on a product of multiplying the synthetic weight vector and eigenvoices associated with the vector space;

generating a synthetic voice sample based on the synthetic feature vector; and

providing the synthetic voice sample to the audio voice framework.

2 . The method of claim 1 , wherein generating the synthetic weight vector includes generating the synthetic weight vector based on a training sample that the audio voice framework processed incorrectly.

3 . The method of claim 1 , wherein providing the synthetic voice sample to the audio voice framework includes:

playing the synthetic voice sample through a speaker to an appliance having the audio voice framework.

4 . The method of claim 1 , further comprising:

providing a natural sample to the audio voice framework as at least one of a testing or training sample.

5 . The method of claim 1 , further comprising:

providing the synthetic feature vector to the audio voice framework for training the audio voice framework.

6 . The method of claim 1 , further comprising:

creating audio files by inputting the synthetic feature vector into a speech synthesizer.

7 . The method of claim 1 , wherein providing the synthetic voice sample to the audio voice framework includes:

providing the synthetic voice sample to an activation trigger engine, the activation trigger engine being configured to detect an activation trigger and, in response thereto, transition a speech analysis engine from an inactive state to an active state,

wherein the synthetic voice sample includes a representation of a synthetic voice uttering the activation trigger.

8 . The method of claim 7 , further comprising:

adjusting weights of the activation trigger engine based on comparing an expected output of the activation trigger engine and an actual output of the activation trigger engine responsive to the synthetic voice sample being provided to the activation trigger engine.

9 . The method of claim 7 , further comprising:

determining, by the speech analysis engine, an intent associated with the synthetic voice sample.

10 . A system for training an audio voice framework, comprising:

one or more processors; and

a computer-readable storage medium coupled to the one or more processors and comprising instructions thereon that, when executed by the one or more processors, cause the one or more processors to:

obtain a first natural weight vector of a vector space and a second natural weight vector of the vector space, each respectively associated with natural voice samples wherein the second natural weight vector is associated with a higher confidence than the first natural weight vector;

generate a synthetic weight vector by way of applying a genetic algorithm to the first natural weight vector and the second natural weight vector, wherein the genetic algorithm selects the synthetic weight vector based on a fitness function applied to the first natural weight vector and the second natural weight vector;

form a synthetic feature vector based on a product of multiplying the synthetic weight vector and eigenvoices associated with the vector space;

generate a synthetic voice sample based on the synthetic feature vector; and

provide the synthetic voice sample to the audio voice framework.

11 . The system of claim 10 , wherein the instructions further cause the one or more processors to:

determine, by a speech analysis engine, an intent associated with the synthetic voice sample.

12 . The system of claim 10 , further comprising:

providing the synthetic feature vector to the audio voice framework for training the audio voice framework.

13 . The system of claim 10 , wherein the instructions further cause the one or more processors to:

provide the synthetic voice sample to an activation trigger detection framework, wherein the synthetic voice sample includes a representation of a synthetic voice uttering an activation trigger; and

determine a fitness of the activation trigger detection framework based on comparing an expected output of the activation trigger detection framework and an actual output of the activation trigger detection framework responsive to the synthetic voice sample being provided to the activation trigger detection framework.

14 . The system of claim 13 , wherein the instructions further cause the one or more processors to:

adjust weights of the activation trigger detection framework based on comparing the expected output of the activation trigger detection framework and the actual output of the activation trigger detection framework responsive to the synthetic voice sample being provided to the activation trigger detection framework.

15 . A method for training an audio voice framework, the method comprising:

generating a set of audio files from a plurality of audio clips of speech of one or more individuals;

representing each audio file of the plurality of audio clips as a feature vector to form a plurality of feature vectors;

generating an average representation vector from the plurality of feature vectors;

subtracting the average representation vector from the plurality of feature vectors to obtain a mean-centered result;

performing singular value decomposition based on the mean-centered result to obtain eigenvoices;

defining a vector space of the plurality of feature vectors based on the eigenvoices;

generating synthetic voice samples within the vector space based on the plurality of feature vectors; and

providing the synthetic voice samples to the audio voice framework as audio samples for training the audio voice framework.

16 . The method of claim 15 , further comprising selecting a subset of the eigenvoices to define the vector space.

17 . The method of claim 15 , further comprising:

determining, by a speech analysis engine, an intent associated with the synthetic voice samples.

18 . The method of claim 15 , wherein providing the synthetic voice samples to the audio voice framework includes:

playing the synthetic voice samples through a speaker to an appliance having the audio voice framework.

19 . The method of claim 15 , further comprising:

providing a natural sample to the audio voice framework as at least one of a testing or training sample.

20 . The method of claim 15 , wherein generating the audio files comprises inputting the plurality of feature vectors into a speech synthesizer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2024
From: BROMAND, DANIEL
To: SPOTIFY AB
Reel/Frame 066678/0625 →
Priority Claims (1)
EP 18167001 · Apr 12, 2018 · regional
Continuity (3)
Continuation 17173659 · Feb 11, 2021
Continuation 16168478 · Oct 23, 2018
Related Publication 20240203401A1 · Jun 20, 2024
References Cited (52)
US 5796916A · Meredith · 1998 [cited by applicant]
US 6141644A · Kuhn et al. · 2000 [cited by applicant]
US 6625587B1 · Erten · 2003 [cited by examiner]
US 7096183B2 · Junqua · 2006 [cited by applicant]
US 7567896B2 · Coorman et al. · 2009 [cited by applicant]
US 7574359B2 · Huang · 2009 [cited by applicant]
US 9098467B1 · Blanksteen et al. · 2015 [cited by applicant]
US 9613620B2 · Agiomyrgiannakis · 2017 [cited by applicant]
US 9852729B2 · Hoffmeister · 2017 [cited by applicant]
US 10943581B2 · Bromand · 2021 [cited by examiner]
US 11887582B2 · Bromand · 2024 [cited by examiner]
US 20010042082A1 · Ueguri · 2001 [cited by applicant]
US 20030023442A1 · Akabane · 2003 [cited by applicant]
US 20030113002A1 · Philomin et al. · 2003 [cited by applicant]
US 20030171930A1 · Junqua · 2003 [cited by examiner]
US 20060085187A1 · Barquilla · 2006 [cited by examiner]
US 20060161437A1 · Akabane · 2006 [cited by applicant]
US 20080112542A1 · Sharma · 2008 [cited by applicant]
US 20080115112A1 · Sharma · 2008 [cited by applicant]
US 20100169093A1 · Washio · 2010 [cited by examiner]
US 20100211392A1 · Tokuda · 2010 [cited by examiner]
US 20110293076A1 · Sharma · 2011 [cited by applicant]
US 20120035917A1 · Kim · 2012 [cited by examiner]
US 20140222415A1 · Legat · 2014 [cited by applicant]
US 20150058019A1 · Chen · 2015 [cited by examiner]
US 20150301796A1 · Visser et al. · 2015 [cited by applicant]
US 20160027430A1 · Dachiraju et al. · 2016 [cited by applicant]
US 20160072945A1 · Kulkarni · 2016 [cited by applicant]
US 20160358600A1 · Nallasamy · 2016 [cited by applicant]
US 20160379622A1 · Patel · 2016 [cited by examiner]
US 20170162186A1 · Tamura · 2017 [cited by applicant]
US 20170169811A1 · Sabbavarapu · 2017 [cited by examiner]
US 20170270907A1 · Mori · 2017 [cited by applicant]
US 20170352353A1 · Dachiraju et al. · 2017 [cited by applicant]
US 20180005628A1 · Xue · 2018 [cited by applicant]
US 20180012593A1 · Prasad et al. · 2018 [cited by applicant]
US 20190066656A1 · Mori · 2019 [cited by applicant]
US 20190318722A1 · Bromand · 2019 [cited by applicant]
EP 1079615 · 2002 [cited by applicant]
Bucila et al., “Model Compression”, KDD'06, Aug. 20-23, 2006. Available at: https://www.cs.cornell.edu/.about.caruana/compression.kdd06.pdf. [cited by applicant]
Dosovitskiy et al., “Learning to Generate Chairs with Convolutional Neural Networks”, IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1538-1546 (Nov. 21, 2014). [cited by applicant]
European Communication in Application 18167001.9, mailed Jan. 3, 2019, 4 pages. [cited by applicant]
European Communication pursuant to Article 94(3) EPC in Application 20165140.3, mailed Oct. 18, 2022, 5 pages. [cited by applicant]
European Extended Search Report from European Application No. 18167001.9, dated Oct. 17, 2018. [cited by applicant]
European Extended Search Report in Application 20165140.3, mailed Jul. 8, 2020, 9 pages. [cited by applicant]
Grauman et al., “Face detection and recognition” (2009). Available at: http://www.cs.unc.edu/.about.lazebnik/spring09/lec22_eigenfaces.pdf. [cited by applicant]
Luise Valentin Rygaard: “Using Synthesized Speech to Improve Speech Recognition for Low-Resource Languages”, pp. 1-6 (Dec. 31, 2015). [cited by applicant]
Prahallad, K., “Speech Technology: A Practical Introduction. Topic: Spectrogram, Cepstrum and Mel-Frequency Analysis” (2008), Carnegie Mellon University & Int'l Inst. of Info. Tech. Hyderabad. Available at: http://www.s… [cited by applicant]
Shichiri et al., “Eigenvoices for HMM-Based Speech Synthesis”, ICSLP 2002: 7th Int'l Conf. on Spoken Language Processing, Denver, CO., pp. 1269 (Sep. 16-20, 2002). [cited by applicant]
Van Den Oord, A., “Wavenet: A Generative Model for Raw Audio”, SSW (2016), arXiv: 1609.03499v2 (Sep. 19, 2016). [cited by applicant]
Van Dyk et al., “The Art of Data Augmentation”, American Statistical Association Institute of Mathematical Statistics, and Interface Foundation of North America Journal of Computational and Graphical Statistics, vol. 10… [cited by applicant]
Zhang et al., “Advanced Data Exploitation in Speech Analysis: An overview”, IEEE Signal Processing Magazine, IEEE Service Center, Piscataway, NJ, US, vol. 34, No. 4, pp. 107-129 (Jul. 1, 2017). [cited by applicant]