IP Library › Granted Patent US 11,710,474
Granted Patent B2
US 11,710,474 · App. 17/146,804 · Granted Jul 25, 2023

Text-to-speech from media content item snippets

Inventors: Rohit Kumar (Austin, TX); Henrik Lindström (Stockholm, SE); Henriette Cramer (San Francisco, CA); Sarah Mennicken (San Francisco, CA); Sravana Reddy (Cambridge, MA); Jennifer Thom-Santelli (Boston, MA)
Assignee: Spotify AB
G10L13/00G06F16/685G10L13/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,710,474
App. No.
17/146,804
Granted
Jul 25, 2023
Kind
B2
Abstract

A text-to-speech engine creates audio output that includes synthesized speech and one or more media content item snippets. The input text is obtained and partitioned into text sets. A track having lyrics that match a part of one of the text sets is identified. The location of the track's audio that contains the lyric is extracted based on forced alignment data. The extracted audio is combined with synthesized speech corresponding to the remainder of the input text to form audio output.

Claims (48)

1. A system providing text-to-speech functionality, comprising:

a text-to-speech engine configured to:

portion an input text into a first text set and a second text set using key-value pairs that match text data with a corresponding text set, wherein the input text is not pre-portioned;

identify from a plurality of tracks, at least one audio snippet from a first track, the first track having lyrics corresponding to at least one word of the first text set;

use a speech synthesizer to create a synthesized utterance of at least one word of the second text set, thereby generating a synthesized utterance;

combine the at least one audio snippet and the synthesized utterance to form combined audio; and

provide an audio output that includes the combined audio.

2. The system according to claim 1 , further comprising:

a forced alignment data store having stored thereon forced alignment data for the plurality of tracks.

3. The system of claim 2 , wherein the forced alignment data describes alignment between audio data of the tracks and lyrics data of the tracks.

4. The system of claim 2 , wherein the forced alignment data, for each of the tracks, describes lyrics data and time data of where lyrics occur in the audio data.

5. The system of claim 2 , further comprising:

a forced alignment engine is configured to:

receive a track of the plurality of tracks having track metadata;

receive track lyrics associated with the track;

select an acoustic model from an acoustic model data store based on the track metadata;

using the acoustic model, generate track forced alignment data that aligns audio data of the track and the track lyrics; and

add the track forced alignment data to the forced alignment data store.

6. The system of claim 1 , the text-to-speech engine further operable to:

determine a suitability of the at least one audio snippet, including by determining how clearly the lyrics of the at least one audio snippet can be heard.

7. The system of claim 1 ,

wherein the text-to-speech engine is further configured to:

portion the input text into a third text set; and

identify from a plurality of tracks, at least one audio snippet from a second track having lyrics corresponding to at least one word of the third text set;

wherein the combined audio further includes the at least one audio snippet from the second track.

8. The system of claim 7 , wherein the identifying the at least one audio snippet from the second track is further based on audio characteristics of the first track and the second track.

9. The system of claim 8 , wherein identifying the at least one audio snippet from the second track includes:

selecting the second track based on similarities in musical style between the first track and the second track.

10. The system of claim 9 , wherein the similarities in musical style between the first track and the second track is determined based on a distance between the first track and the second track in a vector space.

11. The system of claim 1 , further comprising a forced alignment engine for aligning input track audio and input track lyrics.

12. The system of claim 11 , wherein the forced alignment engine is configured to find a Viterbi path through the input track lyrics and the input track audio under an acoustic model.

13. The system of claim 11 , wherein the forced alignment engine is configured to align the input track lyrics and input track audio on a line-by-line basis.

14. The system of claim 1 , wherein the combining engine is further configured to:

determine a quality described in the input text,

wherein identifying at least one audio snippet from the first track is based at least in part on the quality.

15. The system of claim 14 , wherein the quality is a music genre or music artist.

16. A method for providing text-to-speech functionality, comprising:

portioning an input text into a first text set and a second text set using key-value pairs that match data with a corresponding text set, wherein the input text is not pre-portioned;

identifying from a plurality of tracks, at least one audio snippet from a first track, the first track having lyrics corresponding to at least one word of the first text set;

creating a synthesized utterance of at least one word of the second text set, thereby generating a synthesized utterance;

combining the at least one audio snippet and the synthesized utterance to form combined audio; and

providing an audio output that includes the combined audio.

17. The method according to claim 16 , further comprising:

storing, in a forced alignment data store, forced alignment data for the plurality of tracks.

18. The method of claim 17 , wherein the forced alignment data describes alignment between audio data of the tracks and lyrics data of the tracks.

19. The method of claim 17 , wherein the forced alignment data, for each of the tracks, describes lyrics data and time data of where lyrics occur in the audio data.

20. The method of claim 16 , further comprising:

determining a suitability of the at least one audio snippet, including by determining how clearly lyrics of the at least one audio snippet can be heard.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2023
From: SPOTIFY USA INC.
To: SPOTIFY AB
Reel/Frame 063105/0815 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE'S NAME PREVIOUSLY RECORDED AT REEL: 056062 FRAME: 0215. ASSIGNOR(S) HEREBY CONFIRMS THE EMPLOYMENT AGREEMENT. Recorded Mar 13, 2023
From: KUMAR, ROHIT
To: SPOTIFY USA INC.
Reel/Frame 063068/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2021
From: LINDSTROM, HENRIK; CRAMER, HENRIETTE; MENNICKEN, SARAH; REDDY, SRAVANA; THOM-SANTELLI, JENNIFER
To: SPOTIFY AB
Reel/Frame 056052/0627 →
EMPLOYMENT AGREEMENT Recorded Apr 27, 2021
From: KRAMER, ROHIT
To: SPOTIFY AB
Reel/Frame 056062/0215 →
Continuity (2)
Continuation 16235776 · Dec 28, 2018
Related Publication 20210241753A1 · Aug 5, 2021