IP Library › Granted Patent US 12,437,744
Granted Patent B2
US 12,437,744 · App. 18/310,136 · Granted Oct 7, 2025

Text-to-speech from media content item snippets

Inventors: Rohit Kumar (Austin, TX); Henrik Lindström (Stockholm, SE); Henriette Cramer (San Francisco, CA); Sarah Mennicken (San Francisco, CA); Sravana Reddy (Cambridge, MA); Jennifer Thom-Santelli (Boston, MA)
Assignee: Spotify AB
G10L13/00G06F16/685G10L13/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,744
App. No.
18/310,136
Granted
Oct 7, 2025
Kind
B2
Abstract

A text-to-speech engine creates audio output that includes synthesized speech and one or more media content item snippets. The input text is obtained and partitioned into text sets. A track having lyrics that match a part of one of the text sets is identified. The location of the track's audio that contains the lyric is extracted based on forced alignment data. The extracted audio is combined with synthesized speech corresponding to the remainder of the input text to form audio output.

Claims (44)

1. A method for providing text-to-speech functionality, comprising:

splitting an input text into a first text set and a second text set, wherein the first text set is selected for the splitting based on the first text set being associated with suitable lyrics or a suitable lyrical characteristic matched by a key value pair in a dictionary structure, and wherein the second text set is selected for the splitting based on the second text set being suitable for speech synthesis;

extracting an audio snippet from a track corresponding to the first text set, wherein the audio snippet is a portion of the track that begins at a start time within the track and ends at an end time within the track;

creating a synthesized utterance of at least one word of the second text set;

concatenating the audio snippet and the synthesized utterance to form combined audio; and

providing an audio output that includes the combined audio.

2. The method of claim 1 , wherein the track is selected randomly.

3. The method of claim 1 , wherein the track is selected by presenting, via a user device interface, a prompt to a user to select the track from a plurality of tracks.

4. The method of claim 1 , wherein the track is selected from a plurality of tracks by:

sorting the plurality of tracks based on a quality associated with the first text set into a sorted list; and

selecting the track based on the sorted list.

5. The method of claim 1 , wherein the track is selected from a plurality of tracks by:

sorting the plurality of tracks based on a taste profile of a user into a sorted list; and

selecting the track based on the sorted list.

6. The method of claim 1 , wherein the track is selected based on a play list associated with the input text.

7. The method of claim 1 , wherein the track is selected from a plurality of tracks by:

determining a suitability of each track in the plurality of tracks, including by determining how clearly a lyric can be heard; and

selecting the track based on the suitability.

8. The method of claim 1 , wherein the audio snippet contains one or more words of the first text set and, and wherein the combined audio contains the one or more words of the first text set and the at least one word of the second text set.

9. The method of claim 1 , wherein splitting the input text into the first text set and the second text set comprises splitting the input text into the first text set, the second text set, and a third text set, the method further comprising:

extracting a second audio snippet from a second track based on the third text set, wherein the combined audio further includes the second audio snippet.

10. A system providing text-to-speech functionality, comprising a text-to-speech engine configured to:

split an input text into a first text set and a second text set, wherein the first text set is selected for the splitting based on the first text set being associated with suitable lyrics or a suitable lyrical characteristic matched by a key value pair in a dictionary structure, and wherein the second text set is selected for the splitting based on the second text set being suitable for speech synthesis;

extract an audio snippet from a track corresponding to the first text set, wherein the audio snippet is a portion of the track that begins at a start time within the track and ends at an end time within the track;

create a synthesized utterance of at least one word of the second text set;

concatenate the audio snippet and the synthesized utterance to form combined audio; and

provide an audio output that includes the combined audio.

11. The system of claim 10 , wherein the track is selected randomly.

12. The system of claim 10 , wherein the track is selected by presenting, via a user device interface, a prompt to a user to select the track from a plurality of tracks.

13. The system of claim 10 , wherein the track is selected from a plurality of tracks by:

sorting the plurality of tracks based on a quality into a sorted list; and

selecting the track based on the sorted list.

14. The system of claim 10 , wherein the track is selected from a plurality of tracks by:

sorting the plurality of tracks based on a taste profile of a user into a sorted list; and

selecting the track based on the sorted list.

15. The system of claim 10 , wherein the track is selected based on a play list associated with the input text.

16. The system of claim 10 , wherein the track is selected from a plurality of tracks by:

determining a suitability of each track in the plurality of tracks, including by determining how clearly a lyric can be heard; and

selecting the track based on the suitability.

17. The system of claim 10 , wherein the audio snippet contains one or more words of the first text set and, and wherein the combined audio contains the one or more words of the first text set and the at least one word of the second text set.

18. The system of claim 10 , wherein splitting the input text into the first text set and the second text set comprises splitting the input text into the first text set, the second text set, and a third text set, and wherein the text-to-speech engine is further configured to:

extract a second audio snippet from a second track based on the third text set, wherein the combined audio further includes the second audio snippet.

19. The method of claim 1 , wherein the first text set being associated with the suitable lyrics or the suitable lyrical characteristic matched by the key value pair in the dictionary structure comprises the first text set being in a title of the track.

20. The system of claim 10 , wherein the first text set being associated with the suitable lyrics or the suitable lyrical characteristic matched by the key value pair in the dictionary structure comprises the first text set being in a title of the track.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2025
From: THOM-SANTELLI, JENNIFER
To: SPOTIFY AB
Reel/Frame 072030/0870 →
CORRECTIVE ASSIGNMENT TO CORRECT THE EXECUTION DATE PREVIOUSLY RECORDED AT REEL: 63495 FRAME: 318. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 15, 2025
From: LINDSTRÖM, HENRIK; CRAMER, HENRIETTE; MENNICKEN, SARAH; REDDY, SRAVANA
To: SPOTIFY AB
Reel/Frame 072469/0170 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2023
From: LINDSTRÖM, HENRIK; CRAMER, HENRIETTE; MENNICKEN, SARAH; REDDY, SRAVANA
To: SPOTIFY AB
Reel/Frame 063495/0318 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2023
From: SPOTIFY USA INC.
To: SPOTIFY AB
Reel/Frame 063496/0015 →
EMPLOYMENT AGREEMENT Recorded May 1, 2023
From: KUMAR, ROHIT
To: SPOTIFY USA INC.
Reel/Frame 063502/0421 →
Continuity (3)
Continuation 17146804 · Jan 12, 2021
Continuation 16235776 · Dec 28, 2018
Related Publication 20230267912A1 · Aug 24, 2023
References Cited (42)
US 5860064A · Henton · 1999 [cited by applicant]
US 5913193A · Huang · 1999 [cited by examiner]
US 6173263B1 · Conkie · 2001 [cited by examiner]
US 7096183B2 · Junqua · 2006 [cited by applicant]
US 7124082B2 · Freedman · 2006 [cited by applicant]
US 7567896B2 · Coorman et al. · 2009 [cited by applicant]
US 8073854B2 · Whitman et al. · 2011 [cited by applicant]
US 8433431B1 · Master et al. · 2013 [cited by applicant]
US 8666749B1 · Subramanya et al. · 2014 [cited by applicant]
US 10140973B1 · Dalmia · 2018 [cited by applicant]
US 11114085B2 · Kumar et al. · 2021 [cited by applicant]
US 20030074196A1 · Kamanaka · 2003 [cited by applicant]
US 20030200858A1 · Xie · 2003 [cited by applicant]
US 20050131559A1 · Kahn et al. · 2005 [cited by applicant]
US 20060028951A1 · Tozun · 2006 [cited by applicant]
US 20070055527A1 · Jeong et al. · 2007 [cited by applicant]
US 20080071529A1 · Silverman · 2008 [cited by applicant]
US 20080091571A1 · Sater · 2008 [cited by applicant]
US 20090076819A1 · Wouters · 2009 [cited by applicant]
US 20090292535A1 · Seo · 2009 [cited by applicant]
US 20100145705A1 · Kirkeby · 2010 [cited by applicant]
US 20110066438A1 · Lindahl · 2011 [cited by applicant]
US 20110246186A1 · Takeda · 2011 [cited by applicant]
US 20110288862A1 · Todic · 2011 [cited by applicant]
US 20140076125A1 · Kellett · 2014 [cited by examiner]
US 20140200894A1 · Osowski et al. · 2014 [cited by applicant]
US 20170300527A1 · Colangelo · 2017 [cited by examiner]
US 20180041462A1 · Halt · 2018 [cited by examiner]
US 20180103000A1 · Guthery et al. · 2018 [cited by applicant]
US 20210241753A1 · Kumar · 2021 [cited by applicant]
EP 1835488 · 2007 [cited by applicant]
JP 20138357 · 2013 [cited by applicant]
“TextGrid file formats” Phonetic Sci., Amsterdam (Aug. 21, 2018). Available at: http://www.fon.hum.uva.nl/praat/manual/TextGrid_file_formats.html. [cited by applicant]
Diemo Schwarz, “Current Research in Concatenative Sound Synthesis”, Icmc, Ircam - Centre Pompidou, 4 pages (5-9 Sep. 2005). [cited by applicant]
Dzhambazov et al. “Automatic Lyrics-To-Audio Alignment in Classical Turkish Music”, Univ. Pompeu Fabra, p. 61-64 (2014). [cited by applicant]
European Search Report for EP Appln. No. 21167170.6 mailed Jul. 30, 2021 (9 pages). [cited by applicant]
European Extended Search Report from European Appl'n No. 19214458.2, dated May 28, 2020. [cited by applicant]
Fujihara et al. “LyricSyncrhonizer: Automatic Synchronization System Between Musical Audio Signals and Lyrics.” IEEE J. of Selected Topics in Signal Processing, vol. 5, No. 6, Oct. 2011, pp. 1251-1261. [cited by applicant]
Mustapha et al., “Text-to-Speech Synthesis Using Concatenative Approach”, IJTRD, vol. 3(5), pp. 459-462 (Sep.-Oct. 2016). [cited by applicant]
Paul Kehrer “Frinkiac—The Simpsons Screenshot Search Engine”, Langui.sh (Feb. 2, 2016). Available at: https://langui.sh/2016/02/02/frinkiac-the-simpsons-screenshot-search-engine/. [cited by applicant]
Shen and Lee, “Digital Storytelling Book Generator with MIDI-to-Singing”, Applied Mechanics and Materials, vol. 145, Trans Tech Publications, Ltd., pp. 441-445 (Dec. 2011). [cited by applicant]
Van den Oord et al., “Wavenet: A Generative Model for Raw Audio”, CoRR abs/1609.03499 (2016). [cited by applicant]