IP Library Granted Patent US 10,311,855
Granted Patent B2
US 10,311,855 · App. 15/473,103 · Granted Jun 4, 2019

Method and apparatus for designating a soundalike voice to a target voice from a database of voices

Inventors: Fathy Yassa (Soquel, CA); Benjamin Reaves (Menlo Park, CA); Sandeep Mohan (San Jose, CA)
Assignee: SPEECH MORPHING SYSTEMS, INC.
G10L13/033G10L13/00G10L13/047G10L13/10G10L17/00G10L17/12G10L25/27G10L25/51G10L25/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,311,855
App. No.
15/473,103
Granted
Jun 4, 2019
Kind
B2
Abstract

A soundalike system to improve speech synthesis by training a text to speech engine on a voice like the target speakers voice.

Claims (15)

1. A speech synthesis system comprising:

a text-to-speech (TTS) system;

a database that stores a collection of voices; and

a processor configured to:

perform training on each voice among the collection of voices stored in the database by building a mathematical model of each voice among the collection of voices stored in the database,

cluster each voice among the collection of voices stored in the database into a voice cluster among a plurality of voice clusters based on similarities between voice characteristics of each voice among the collection of voices stored in the database,

calculate an i-vector for each voice cluster among the plurality of voice clusters, the i-vector for each voice cluster among the plurality of voice clusters representing a mathematical model of each voice cluster among the plurality of voice clusters;

calculate a target i-vector for a target voice,

identify a matching voice cluster among the plurality of voice clusters having an i-vector among the i-vectors for each voice cluster among the plurality of voice clusters that most closely matches the target i-vector for the target voice,

calculate i-vectors of each voice within the matching voice cluster,

identify a soundalike voice that most closely matches the target voice among each voice in the matching voice cluster having an i-vector among the i-vectors for each voice in the matching voice cluster that most closely matches the target i-vector for the target voice, and

build a TTS voice from the soundalike voice for training the TTS.

2. The speech synthesis system of claim 1 , wherein the i-vector of the matching voice cluster has a lowest Euclidean distance from the target i-vector for the target voice among the i-vectors for each voice cluster among the plurality of voice clusters, and

wherein the i-vector of the soundalike voice has a lowest Euclidean distance from the target i-vector for the target voice among the i-vectors for each voice in the matching voice cluster.

3. The speech synthesis system of claim 2 , wherein voice characteristics comprise at least one of speech pitch and speech speed.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2017
From: YASSA, FATHY; REAVES, BENJAMIN; MOHAN, SANDEEP
To: SPEECH MORPHING SYSTEMS, INC.
Reel/Frame 042932/0396 →
Continuity (2)
Provisional Application 62314759 · Mar 29, 2016
Related Publication 20170301340A1 · Oct 19, 2017