IP Library Granted Patent US 12,711,980
Granted Patent B2
US 12,711,980 · App. 18/436,143 · Granted Aug 18, 2026

Identifying shifts in audio content via machine learning

Inventors: Peter Shoebridge (Boulder, CO); Jeffrey Thramann (Boulder, CO); Pablo Calderon Rodriguez (Boulder, CO)
Assignee: Auddia Inc.
G10L25/51G10L25/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,711,980
App. No.
18/436,143
Filed
Feb 8, 2024
Granted
Aug 18, 2026
Kind
B2
Art Unit
2692
USPC
700/94
Abstract

A method and system for identifying the beginning and ending of songs via a machine learning analysis. A machine learning model analyzes streaming audio (such as a radio broadcast) in overlapping, 3-second samples. Each sample is labeled into groups such as “song,” “talk,” “commercial” and “transition.” Based on the location of the transition samples, an exact second a given song begins and ends in the audio stream is derivable. The model further identifies when two songs shift between one another.

Claims (44)

1 . A method for classifying segments of an audio stream of a radio program comprising:

labeling a plurality of consecutive audio samples of the audio stream with a trained machine learning model via successive inspection, the trained machine learning model configured to output a label corresponding to each audio sample indicating whether each respective audio sample is a song portion, a talk portion, or a commercial portion of the audio stream resulting in a sequence of labels;

executing a first probabilistic correction on the sequence of labels based on patterns represented within the sequence of labels and resulting in a corrected sequence of labels;

identifying a set of consecutive audio samples as having song portion labels; and

determining, via the trained machine learning model, whether the set of consecutive audio samples are a matching song.

2 . The method of claim 1 , further comprising:

in response to identification that the set of consecutive audio samples belong to different songs, determining, via the trained machine learning model, a transition time between two different songs through use of consecutive overlapping audio samples.

3 . The method of claim 2 , wherein determining the transition time between the two different songs includes executing a second probabilistic correction on a sequence of comparisons of contiguous audio samples.

4 . The method of claim 2 , wherein the transition time between the two different songs further comprises:

comparing a set of contiguous audio samples of the plurality of audio samples.

5 . The method of claim 1 , wherein the plurality of consecutive audio samples are overlapping.

6 . The method of claim 2 , wherein said determining the transition time further includes:

inserting a marker at an end of a song where the consecutive audio samples transition between songs.

7 . The method of claim 1 , wherein the successive inspection of consecutive audio samples further comprises:

advancing a frame of inspection by a temporal period that is shorter than a temporal length of each audio sample.

8 . The method of claim 6 , wherein the successive inspections overlap by 1 second.

9 . A computing device for classifying segments of an audio stream of a radio program comprising:

a processor; and

a non-transitory computer-readable medium having stored thereon instructions that, when executed by the processor, cause the processor to perform operations including:

label a plurality of consecutive audio samples of the audio stream with a trained machine learning model via successive inspection, the trained machine learning model configured to output a label corresponding to each audio sample indicating whether each respective audio sample is a song portion, a talk portion, or a commercial portion of the audio stream resulting in a sequence of labels;

execute a first probabilistic correction on the sequence of labels based on patterns represented within the sequence of labels and resulting in a corrected sequence of labels;

identify a set of consecutive audio samples as having song portion labels; and

determine, via the trained machine learning model, whether the set of consecutive audio samples are a matching song.

10 . The computing device of claim 9 , the instructions further comprising:

in response to identification that the set of consecutive audio samples belong to different songs, determining, via the trained machine learning model, a transition time between two different songs through use of consecutive overlapping audio samples.

11 . The computing device of claim 10 , wherein determining the transition time between the two different songs includes executing a second probabilistic correction on a sequence of comparisons of contiguous audio samples.

12 . The computing device of claim 10 , wherein the transition time between the two different songs further comprises:

comparing a set of contiguous audio samples of the plurality of audio samples.

13 . The computing device of claim 9 , wherein the plurality of consecutive audio samples are overlapping.

14 . The computing device of claim 10 , wherein said determining the transition time further includes:

inserting a marker at an end of a song where the consecutive audio samples transition between songs.

15 . A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processor to perform operations for classifying segments of an audio stream of a radio program comprising:

labeling a plurality of consecutive audio samples of the audio stream with a trained machine learning model via successive inspection, the trained machine learning model configured to output a label corresponding to each audio sample indicating whether each respective audio sample is a song portion, a talk portion, or a commercial portion of the audio stream resulting in a sequence of labels;

executing a first probabilistic correction on the sequence of labels based on patterns represented within the sequence of labels and resulting in a corrected sequence of labels;

identifying a set of consecutive audio samples as having song portion labels; and

determining, via the trained machine learning model, whether the set of consecutive audio samples are a matching song.

16 . The non-transitory computer-readable medium of claim 15 , the instructions further comprising:

in response to identification that the set of consecutive audio samples belong to different songs, determining, via the trained machine learning model, a transition time between two different songs through use of consecutive overlapping audio samples.

17 . The non-transitory computer-readable medium of claim 16 , wherein determining the transition time between the two different songs includes executing a second probabilistic correction on a sequence of comparisons of contiguous audio samples.

18 . The non-transitory computer-readable medium of claim 16 , wherein the transition time between the two different songs further comprises:

comparing a set of contiguous audio samples of the plurality of audio samples.

19 . The non-transitory computer-readable medium of claim 15 , wherein the plurality of consecutive audio samples are overlapping.

20 . The non-transitory computer-readable medium of claim 16 , wherein said determining the transition time further includes:

inserting a marker at an end of a song where the consecutive audio samples transition between songs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2026
From: SHOEBRIDGE, PETER; RODRIGUEZ, PABLO CALDERON; THRAMANN, JEFFREY
To: AUDDIA INC.
Reel/Frame 074022/0900 →
Continuity (4)
Continuation In Part 17123761 · Dec 16, 2020
Provisional Application 63444449 · Feb 9, 2023
Provisional Application 62949228 · Dec 17, 2019
Related Publication 20240185878A1 · Jun 6, 2024
References Cited (75)
US 7174293B2 · Kenyon et al. · 2007 [cited by applicant]
US 8170701B1 · Lu · 2012 [cited by examiner]
US 8209713B1 · Lai et al. · 2012 [cited by applicant]
US 8706272B2 · Lindahl · 2014 [cited by examiner]
US 8798776B2 · Schildbach · 2014 [cited by examiner]
US 8892450B2 · Schildbach · 2014 [cited by examiner]
US 10652298B2 · Thomas · 2020 [cited by applicant]
US 11662972B2 · Raikar · 2023 [cited by examiner]
US 11755273B2 · Graham · 2023 [cited by examiner]
US 11785387B2 · Carrigan · 2023 [cited by examiner]
US 11935520B1 · Shoebridge · 2024 [cited by examiner]
US 20040260682A1 · Herley · 2004 [cited by examiner]
US 20050126369A1 · Kirkeby · 2005 [cited by examiner]
US 20070016918A1 · Alcorn · 2007 [cited by applicant]
US 20080082510A1 · Wang · 2008 [cited by examiner]
US 20090053991A1 · Mantel · 2009 [cited by examiner]
US 20100042412A1 · Aronowitz et al. · 2010 [cited by applicant]
US 20100198380A1 · Peiffer · 2010 [cited by examiner]
US 20100293072A1 · Murrant · 2010 [cited by examiner]
US 20110054647A1 · Chipchase · 2011 [cited by examiner]
US 20110075851A1 · LeBoeuf · 2011 [cited by examiner]
US 20110196517A1 · Lindahl · 2011 [cited by examiner]
US 20110208521A1 · McClain · 2011 [cited by examiner]
US 20120053710A1 · Lindahl · 2012 [cited by examiner]
US 20120071995A1 · Topchy · 2012 [cited by examiner]
US 20120136466A1 · Weiss · 2012 [cited by examiner]
US 20120221131A1 · Wang · 2012 [cited by examiner]
US 20130024016A1 · Bhat · 2013 [cited by examiner]
US 20130211567A1 · Oganesyan · 2013 [cited by examiner]
US 20130261781A1 · Topchy · 2013 [cited by examiner]
US 20130317635A1 · Bates · 2013 [cited by examiner]
US 20130318087A1 · Yu · 2013 [cited by examiner]
US 20130318114A1 · Emerson, III · 2013 [cited by examiner]
US 20130331972A1 · Sagne · 2013 [cited by examiner]
US 20140031960A1 · Hill · 2014 [cited by examiner]
US 20140052770A1 · Gran et al. · 2014 [cited by applicant]
US 20140195028A1 · Emerson, III · 2014 [cited by examiner]
US 20140214190A1 · Wang · 2014 [cited by examiner]
US 20140277652A1 · Watts · 2014 [cited by examiner]
US 20140277653A1 · Watts · 2014 [cited by examiner]
US 20140288686A1 · Sant · 2014 [cited by examiner]
US 20140336797A1 · Emerson, III · 2014 [cited by examiner]
US 20140336798A1 · Emerson, III · 2014 [cited by examiner]
US 20150025663A1 · Cameron · 2015 [cited by examiner]
US 20150120336A1 · Grokop · 2015 [cited by examiner]
US 20150148928A1 · Malsbary · 2015 [cited by examiner]
US 20150199968A1 · Singhal · 2015 [cited by examiner]
US 20150271598A1 · Thompson · 2015 [cited by examiner]
US 20150301791A1 · Harwood · 2015 [cited by examiner]
US 20150341410A1 · Schrempp et al. · 2015 [cited by applicant]
US 20160019876A1 · Jeffrey · 2016 [cited by examiner]
US 20160072599A1 · Kariyappa · 2016 [cited by examiner]
US 20160092926A1 · Herberger · 2016 [cited by examiner]
US 20160125892A1 · Bowen · 2016 [cited by examiner]
US 20160140224A1 · Jin · 2016 [cited by examiner]
US 20170301340A1 · Yassa · 2017 [cited by examiner]
US 20180121159A1 · Thompson et al. · 2018 [cited by applicant]
US 20180166066A1 · Dimitriadis · 2018 [cited by examiner]
US 20180314979A1 · Talwar et al. · 2018 [cited by applicant]
US 20190102458A1 · Roblek · 2019 [cited by examiner]
US 20190392852A1 · Hijazi · 2019 [cited by examiner]
US 20200021375A1 · Stavrowski et al. · 2020 [cited by applicant]
US 20200251115A1 · Farinelli · 2020 [cited by examiner]
US 20200412864A1 · Al Majid · 2020 [cited by examiner]
US 20240185878A1 · Shoebridge · 2024 [cited by examiner]
Baluja et al., “Waveprint: Efficient wavelet-based audio fingerprinting,” Pattern Recognition 41:3467-3480 (2008). [cited by applicant]
Cano et al., “Robust Sound Modeling for Song Detection in Broadcast Audio,” Proc. Audio Engineering Society Convention Paper, Presented at the 112th Covention, Munich, Germany pp. 1-7 (May 10-13, 2002). [cited by applicant]
Herley, “ARGOS: Automatically Extracting Repeating Objects From Multimedia Streams.” IEEE Transactions on Multimedia, 8(1):115-129 (publication date: Feb. 2006). [cited by applicant]
Jung et al., “A Probabilistic Ranking Model for Audio Stream Retrieval,” Proceedings of the 1st International Workshop on Multimedia Analysis and Retrieval for Multimodal Interaction pp. 33-38 2016. [cited by applicant]
Koolagudi et aI. “Advertisement Detection in Commercial Radio Channels,” 2015 IEEE 10th International Conference on Industrial and Information Systems, ICIIS, 272-277, (publication date: Dec. 18-20, 2015). [cited by applicant]
Senevirathna et al., “Audio Music Monitoring: Analyzing Current Techniques for Song Recognition and Identification,” GSTF Journal on Computing (JoC), 4(3):23-34 (publication date: Oct. 2015). [cited by applicant]
Senevirathna et al., “Radio Broadcast Monitoring to Ensure Copyright Ownership,” International Journal on Advances in ICT for Emerging Regions (publication date: Jul. 2018). [cited by applicant]
Shah et al., “Efficient Broadcast Monitoring using Audio Change Detection,” Proceedings of the Fifth Indian International Conference on Artificial Intelligence, Tumkur, India (2011). [cited by applicant]
Muller-Cajar, Robin, Univ. of Canterbury, student thesis entitled Detecting Advertising in Radio using Machine Learning, 2007, pp. 1-34. (Year: 2007). [cited by applicant]
SHOUTcast XML Metadata Specification, available on archive.org as of Jan. 17, 2016, https://web.archive.org/web/2016011717 4300/http://wiki .shoutcast.com/index.php?title=SHOUTcast_XML_Metadata_Specification&oldid=74966… [cited by applicant]