IP Library Granted Patent US 10,347,238
Granted Patent B2
US 10,347,238 · App. 15/796,292 · Granted Jul 9, 2019

Text-based insertion and replacement in audio narration

Inventors: Zeyu Jin (Princeton, NJ); Gautham J. Mysore (San Francisco, CA); Stephen DiVerdi (Oakland, CA); Jingwan Lu (Santa Clara, CA); Adam Finkelstein (Princeton, NJ)
Assignees: Adobe Inc.; The Trustees of Princeton University
G10L13/08G10L13/04G10L13/07G10L15/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,347,238
App. No.
15/796,292
Granted
Jul 9, 2019
Kind
B2
Abstract

Systems and techniques are disclosed for synthesizing a new word or short phrase such that it blends seamlessly in the context of insertion or replacement in an existing narration. In one such embodiment, a text-to-speech synthesizer is utilized to say the word or phrase in a generic voice. Voice conversion is then performed on the generic voice to convert it into a voice that matches the narration. An editor and interface are described that support fully automatic synthesis, selection among a candidate set of alternative pronunciations, fine control over edit placements and pitch profiles, and guidance by the editors own voice.

Claims (44)

1. A method for performing text-based insertion in a target voice waveform, the method comprising:

receiving a query text that indicates an insertion in a voice transcript associated with said target voice waveform; and

for each phoneme associated with said query text, generating an audio frame range that comprises a portion of said target voice waveform, wherein generating the audio frame range comprises

generating a query waveform from said query text,

processing the query waveform to generate a query exemplar,

performing a range selection process to generate the audio frame range, wherein the range selection process uses the query exemplar and a first segment that corresponds to a phoneme in said target voice waveform, and wherein the audio frame range comprises a range of consecutive frames that encompasses at least a portion of the first segment, and

generating an edited waveform by modifying said target voice waveform using said audio frame range.

2. The method according to claim 1 , wherein said first segment is generated by processing said query waveform that is generated from the query text.

3. The method according to claim 1 , further comprising processing said query waveform to generate the first segment by:

processing said query waveform to generate a second segment; and,

processing said second segment to generate said first segment.

4. The method according to claim 3 , wherein said second segment is generated by performing a triphone pre-selection process and said first segment is generated by performing a dynamic triphone pre-selection process.

5. The method according to claim 1 , wherein processing said query waveform to generate the query exemplar further comprises:

extracting a feature associated with said query waveform to generate query feature data; and,

processing said query feature data to generate said query exemplar.

6. The method according to claim 5 , wherein said query exemplar is generated by concatenating a plurality of features associated with the query feature data.

7. The method according to claim 1 , wherein said query waveform is generated by applying a text-to-speech (“TTS”) voice to said query text.

8. A system for performing text-based insertion in a target voice waveform comprising:

a corpus processing engine that comprises

a text-to-speech (“TTS”) selection module that generates a source TTS voice, and

a voice conversion module generator that generates a voice conversion module;

an interactive voice editing module that comprises a query input module, the source TTS voice, and the voice conversion module; and

an exemplar-to-edited waveform block that comprises an exemplar to waveform translator module and a concatenative synthesis module, wherein the exemplar to waveform translator module receives an exemplar and a target voice waveform and generates an audio snippet and a context waveform.

9. The system according to claim 8 , wherein said voice conversion module further comprises:

an exemplar extraction module, wherein said exemplar extraction module generates a query exemplar based upon a query waveform;

a segment selection module, wherein said segment selection module generates a segment based upon said query waveform and segment selection data; and,

a range selection module, wherein said range selection module generates a range of exemplars based upon said query exemplar and said segment.

10. The system according to claim 8 , wherein said concatenative synthesis module generates an edited waveform based upon said audio snippet and said context waveform.

11. A computer program product including one or more non-transitory machine-readable mediums encoded with instructions that when executed by one or more processors cause a process to be carried out for performing text-based insertion and replacement a target voice waveform, said process comprising:

receiving a query text;

generating a query waveform from said query text;

processing said query waveform to generate a first segment, wherein said first segment corresponds to a phoneme in said target voice waveform;

processing said query waveform to generate a query exemplar;

performing a range selection process utilizing said query exemplar and said first segment to generate a proposed range, wherein the proposed range comprises a range of consecutive frames that encompasses at least a portion of the first segment; and,

generating an edited waveform by modifying said target voice waveform using said proposed range.

12. The computer program product according to claim 11 , wherein processing said query waveform to generate the first segment further comprises:

processing said query waveform to generate a second segment; and,

processing said second segment to generate said first segment.

13. The computer program product according to claim 12 , wherein said second segment is generated by performing a triphone pre-selection process and said first segment is generated by performing a dynamic triphone pre-selection process.

14. The computer program product according to claim 11 , wherein processing said query waveform to generate the query exemplar further comprises:

extracting a feature associated with said query waveform to generate query feature data; and,

processing said query feature data to generate said query exemplar.

15. The computer program product according to claim 14 , wherein said query exemplar is generated by concatenating a plurality of features associated with the query feature data.

16. The computer program product according to claim 11 , wherein said query waveform is generated by applying a text-to-speech (“TTS”) voice to said query text.

Assignments (3)
CHANGE OF NAME Recorded Nov 30, 2018
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 047688/0530 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2017
From: FINKELSTEIN, ADAM; JIN, ZEYU
To: THE TRUSTEES OF PRINCETON UNIVERSITY
Reel/Frame 044303/0156 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2017
From: MYSORE, GAUTHAM J.; DIVERDI, STEPHEN; LU, JINGWAN
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 044005/0245 →
Continuity (1)
Related Publication 20190130894A1 · May 2, 2019
Cited By (2)
US 12,288,547 US 12,573,374