IP Library Granted Patent US 11,854,563
Granted Patent B2
US 11,854,563 · App. 17/307,397 · Granted Dec 26, 2023

System and method for creating timbres

Inventors: William Carter Huffman (Cambridge, MA); Michael Pappas (Cambridge, MA)
Assignee: Modulate, Inc.
G10L21/013G10L15/02G10L15/063G10L15/22G10L19/018G10L25/30G10L2015/025G10L2021/0135
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,854,563
App. No.
17/307,397
Granted
Dec 26, 2023
Kind
B2
Abstract

A method of building a new voice having a new timbre using a timbre vector space includes receiving timbre data filtered using a temporal receptive field. The timbre data is mapped in the timbre vector space. The timbre data is related to a plurality of different voices. Each of the plurality of different voices has respective timbre data in the timbre vector space. The method builds the new timbre using the timbre data of the plurality of different voices using a machine learning system.

Claims (42)

1. A method of training a speech conversion system, the method comprising:

receiving source speech data that is a function of a first speech segment of a source voice;

receiving target timbre data relating to a target voice, the target timbre data being within a timbre space;

using a generative machine learning system to produce, as a function of the source speech data and the target timbre data, first candidate data that is a function of a first candidate speech segment in a first candidate voice;

receiving inconsistency data relating to a difference between the first candidate data and data relating to the target voice, the inconsistency data being a function of a plurality of voices;

feeding back the inconsistency data to the generative machine learning system;

refining the target timbre data in the timbre space as a result of said feeding back to produce refined target timbre data.

2. The method as defined by claim 1 , wherein the source speech data is from an audio input of the source voice.

3. The method as defined by claim 1 , further comprising:

using a generative machine learning system to produce second candidate data in a second candidate voice as a function of the source speech data and the refined target timbre data;

receiving second inconsistency data, the second inconsistency data being a function of a plurality of voices, the second inconsistency data having information relating to a difference between the second candidate data and data relating to the target voice.

4. The method as defined by claim 1 , further comprising transforming the source speech data to into the target timbre.

5. The method as defined by claim 1 , wherein the target timbre data is obtained from an audio input in the target voice.

6. The method as defined by claim 1 , wherein the machine learning system is a neural network.

7. The method as defined by claim 1 , further comprising:

mapping a representation of the plurality of voices and the first candidate voice in a vector space as a function of a frequency distribution in the speech segment provided by each voice.

8. The method as defined by claim 7 , further comprising:

adjusting a representation of the first candidate voice relative to representations of the plurality of voices in the vector space to reflect the second candidate voice as a function of the inconsistency message.

9. The method as defined by claim 1 , wherein the inconsistency message is produced when the discriminative neural network has less than a 95 percent confidence interval that the first candidate voice is the target voice.

10. A system for training a speech conversion system, the system comprising:

source speech data that represents a first speech segment of a source voice;

target timbre data that relates to a target voice;

a generative machine learning system configured to produce first candidate data that represents a first candidate voice as a function of the source speech data and the target timbre data;

an inconsistency message having information relating to a distinction between the first candidate data and data relating to the target voice, the inconsistency message being a function of a plurality of voices, wherein the inconsistency message is used to refine the target timbre data in the timbre space to produce refined target timbre data.

11. The system as defined by claim 10 , wherein the source speech data is from an audio input of the source voice.

12. The system as defined by claim 10 , wherein the generative machine learning system is configured to produce second candidate data in a second candidate voice as a function of the source speech data and the refined target timbre data.

13. The system as defined by claim 12 , further comprising second inconsistency data, the second inconsistency data being a function of a plurality of voices, the second inconsistency data having information relating to a difference between the second candidate data and data relating to the target voice.

14. The system as defined by claim 10 , wherein the target timbre data is obtained from an audio input in the target voice.

15. The system as defined by claim 10 , wherein the machine learning system is a neural network.

16. A method of building a speech conversion system using target voice information from a target voice, and speech data that represents a speech segment of a source voice, the method comprising:

receiving source speech data that is a function of a first speech segment of a source voice;

receiving target timbre data relating to the target voice, the target timbre data being within a timbre space;

using a generative machine learning system to produce first candidate data that is a function of a first candidate speech segment in a first candidate voice as a function of the source speech data and the target timbre data;

receiving inconsistency data, the inconsistency data being a function of a plurality of voices, the inconsistency data having information relating to a difference between the first candidate data and data relating to the target voice;

feeding back the inconsistency data to the generative machine learning system; and

refining the generative machine learning system as a result of said feeding back.

17. The method as defined by claim 16 , wherein the inconsistency data is a function of a plurality of timbre data.

18. The method as defined by claim 16 , further comprising:

using the generative machine learning system to produce second candidate data in a second candidate voice as a function of the source speech data and the feeding back;

receiving second inconsistency data, the second inconsistency data being a function of a plurality of voices, the second inconsistency data having information relating to a difference between the second candidate data and data relating to the target voice.

19. The method as defined by claim 16 , further comprising:

using the generative machine learning system to produce sequential candidate data in a sequential candidate voice as a function of the source speech data and the feeding back until the inconsistency data indicates no difference between the sequential candidate data and the data relating to the target voice.

Assignments (2)
CERTIFICATE OF CONVERSION Recorded May 13, 2021
From: MODULATE, LLC
To: MODULATE, INC.
Reel/Frame 056237/0633 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2021
From: HUFFMAN, WILLIAM CARTER; PAPPAS, MICHAEL
To: MODULATE, LLC
Reel/Frame 056237/0641 →
Continuity (4)
Continuation 16846460 · Apr 13, 2020
Continuation 15989072 · May 24, 2018
Provisional Application 62510443 · May 24, 2017
Related Publication 20210256985A1 · Aug 19, 2021
Cited By (2)
US 12,341,619 US 12,412,588