IP Library Granted Patent US 12,236,956
Granted Patent B2
US 12,236,956 · App. 17/524,679 · Granted Feb 25, 2025

Audio content processing systems and methods

Inventors: Mari Joller (New York City, NY); Ottokar Tilk (Tallinn, EE); Aleksandr Tkatšenko (Tallinn, EE); Johnathan Joseph Groat (Taylor, MI); Mark Fišel (Tartu, EE); Kaur Karus (Tartu, EE)
Assignee: A9.com, Inc.
G10L17/00G06F16/635G10L25/63G10L25/81G10L25/90
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,956
App. No.
17/524,679
Granted
Feb 25, 2025
Kind
B2
Abstract

This disclosure relates to systems and methods for processing content and, particularly, but not exclusively, systems and methods for processing audio content. Systems and methods are described that provide techniques for processing, analyzing, and/or structuring of longer-form content to, among other things, make the content searchable, identify relevant and/or interesting segments within the content, provide for and/or otherwise generate search results and/or coherent shorter-form summaries and/or highlights, enable new shorter-form audio listening experiences, and/or the like. Various aspects of the disclosed systems and methods may further enable relatively efficient transcription and/or indexing of content libraries at scale, while also generating effective formats for users interacting with such libraries to engage with search results.

Claims (49)

1. A method of processing content performed by a content processing system comprising a processor and a non-transitory computer-readable storage medium storing instructions that, when executed, cause the content processing system to perform the method, the method comprising:

receiving a first content file;

generating, based on the first content file, a first text file comprising transcribed text corresponding to the first content file;

extracting one or more words from the first text file;

identifying a plurality of segments in the first text file based on the extracted one or more words, wherein identifying the plurality of segments includes iteratively splitting relatively larger segments into smaller segments at split points, and wherein determining a first split point of the split points in a first segment is based on an assessment of a first cohesion of the first segment, a second cohesion of a second segment resulting from the split of the first segment, and a third cohesion of a third segment resulting from the split of the first segment;

identifying based on the extracted one or more words, one or more topics, wherein segments of the plurality of segments are associated with at least one topic of the identified one or more topics, wherein the one or more topics includes a first topic;

identifying, in accordance with one or more parameters, one or more highlight segments from the plurality of segments based on a determined relevance of the one or more highlight segments relative to the first topic; and

generating a second content file, the second content file comprising at least an indication of one more content highlight segments of the content file, wherein a first content highlight segment of the one or more content highlight segments corresponds to a first highlight segment of the one or more highlight segments.

2. The method of claim 1 , wherein the one or more parameters comprise one or more user specified parameters comprise at least one of a target size of the one or more content highlight segments, a specified number of content highlight segments, and relative scoring threshold information associated with the one or more highlight segments.

3. The method of claim 1 , wherein the method further comprises: scoring one or more segments of the plurality of segments based on the determined relevance of the one or more segments relative to at least the first topic,

wherein identifying the one or more highlight segments is based on a comparison between one or more scores associated with the one or more highlight segments and at least one threshold.

4. The method of claim 1 , wherein the method further comprises:

analyzing the first content file to identify one or more speakers; and

labeling portions of the first text file based on the identified one or more speakers,

wherein identifying the plurality of the segments in the first text file is further based on the labeled portions of the first text file.

5. The method of claim 1 , wherein identifying the plurality of segments in the first text file is further based on one or more of a lexical feature, and a grammatical feature of the first text file.

6. The method of claim 1 , wherein the method further comprises adding punctuation to the first text file prior to identifying the plurality of segments in the first text file.

7. The method of claim 1 , wherein identifying the plurality of segments in the first text file comprises,

identifying at least one filtered segment; and

excluding the at least one filtered segment from the plurality of segments.

8. The method of claim 7 , wherein the at least one filtered segment comprises at least one of an introduction segment, an advertisement segment, a conclusion segment, and a music segment.

9. The method of claim 1 , wherein the one or more topics includes a second topic, and wherein identifying one or more highlight segments from the plurality of segments is further based on a determined relevance of the one or more highlight segments relative to the second topic.

10. The method of claim 1 , wherein identifying, based on the extracted one or more words, one or more topics comprises generating a content graph linking at least a portion of the extracted one or more words associated with a first segment to the first topic, and at least a portion of the extracted one or more word associated with a second segment to a second topic.

11. A system comprising:

one or more processors; and

memory storing executable instructions that, when executed by the one or more processors,

cause the system to:

receive a first content file;

generate, based on the first content file, a first text file comprising transcribed text corresponding to the first content file;

extract one or more words from the first text file;

identify a plurality of segments in the first text file based on the extracted one or more words, wherein identifying the plurality of segments includes splitting a first segment into a second segment and a third segment, and assessing a first cohesion of the first segment, a second cohesion of a second segment, and a third cohesion of a third segment resulting;

identify, based on the extracted one or more words, one or more topics, wherein segments of the plurality of segments are associated with at least one topic of the identified one or more topics, wherein the one or more topics includes a first topic,

identify, in accordance with one or more parameters, one or more highlight segments from the plurality of segments based on a determined relevance of the one or more highlight segments relative to the first topic; and

generate a second content file, the second content file comprising at least an indication of one more content highlight segments of the content file, wherein a first content highlight segment of the one or more content highlight segments corresponds to a first highlight segment of the one or more highlight segments.

12. The system of claim 11 , wherein the one or more parameters comprise one or more user specified parameters comprise at least one of a target size of the one or more content highlight segments, a specified number of content highlight segments, and relative scoring threshold information associated with the one or more highlight segments.

13. The system of claim 11 , wherein the executable instructions include further instructions that, when executed by the one or more processors, further cause the system to:

score one or more segments of the plurality of segments based on the determined relevance of the one or more segments relative to at least the first topic,

wherein identifying the one or more highlight segments is based on a comparison between one or more scores associated with the one or more highlight segments and at least one threshold.

14. The system of claim 11 , wherein the executable instructions include further instructions that, when executed by the one or more processors, further cause the system to:

analyze the first content file to identify one or more speakers; and

label portions of the first text file based on the identified one or more speakers, wherein identifying the plurality of the segments in the first text file is further based on the labeled portions of the first text file.

15. The system of claim 11 , wherein identifying the plurality of segments in the first text file is further based on one or more of a lexical feature, and a grammatical feature of the first text file.

16. The system of claim 11 , wherein the executable instructions include further instructions that, when executed by the one or more processors, further cause the system to:

add punctuation to the first text file prior to identifying the plurality of segments in the first text file.

17. The system of claim 11 , wherein the executable instructions include further instructions that, when executed by the one or more processors to identify the plurality of segments in the text file, further cause the system to,

identify at least one filtered segment; and

exclude the at least one filtered segment from the plurality of segments, wherein the at least one filtered segment comprises at least one of an introduction segment, an advertisement segment, a conclusion segment, and a music segment.

18. The system of claim 11 , wherein the one or more topics includes a second topic, and wherein identifying one or more highlight segments from the plurality of segments is further based on a determined relevance of the one or more highlight segments relative to the second topic.

19. The system of claim 11 , wherein identifying, based on the extracted one or more words, one or more topics comprises generating a content graph linking at least a portion of the extracted one or more words associated with a first segment to the first topic, and at least a portion of the extracted one or more word associated with a second segment to a second topic.

Assignments (2)
CHANGE OF NAME Recorded Oct 18, 2023
From: SNACKABLE LLC
To: A9.COM
Reel/Frame 065270/0227 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2022
From: JOLLER, MARI; TILK, OTTOKAR; TKAT¿ENKO, ALEKSANDR; GROAT, JONATHAN JOSEPH; FI¿EL, MARK; KARUS, KAUR
To: SNACKABLE INC.
Reel/Frame 060102/0783 →
Continuity (3)
Continuation 16584893 · Sep 26, 2019
Provisional Application 62737672 · Sep 27, 2018
Related Publication 20220139398A1 · May 5, 2022
References Cited (34)
US 8600745B2 · Hirschberg et al. · 2013 [cited by applicant]
US 8612211B1 · Shires et al. · 2013 [cited by applicant]
US 20040006748A1 · Srivastava et al. · 2004 [cited by applicant]
US 20040062367A1 · Fellenstein et al. · 2004 [cited by applicant]
US 20080177536A1 · Sherwani et al. · 2008 [cited by applicant]
US 20090204399A1 · Akamine · 2009 [cited by examiner]
US 20100318537A1 · Surendran et al. · 2010 [cited by applicant]
US 20130185303A1 · Ceusters · 2013 [cited by examiner]
US 20130226930A1 · Arngren et al. · 2013 [cited by applicant]
US 20140074471A1 · Sankar · 2014 [cited by examiner]
US 20150279390A1 · Mani · 2015 [cited by applicant]
US 20160027442A1 · Burton et al. · 2016 [cited by applicant]
US 20160162476A1 · Munro · 2016 [cited by examiner]
US 20160283185A1 · McLaren et al. · 2016 [cited by applicant]
US 20160284354A1 · Chen et al. · 2016 [cited by applicant]
US 20170091154A1 · Eppolito · 2017 [cited by examiner]
US 20170147544A1 · Modani et al. · 2017 [cited by applicant]
US 20170169816A1 · Blandin et al. · 2017 [cited by applicant]
US 20170177715A1 · Chang · 2017 [cited by examiner]
US 20170199934A1 · Nongpiur · 2017 [cited by examiner]
US 20170277784A1 · Hay · 2017 [cited by examiner]
US 20170323643A1 · Arslan · 2017 [cited by examiner]
US 20180060289A1 · Grueneberg · 2018 [cited by examiner]
US 20180308519A1 · Chik et al. · 2018 [cited by applicant]
US 20190121851A1 · Shires · 2019 [cited by examiner]
US 20190180175A1 · Meteer et al. · 2019 [cited by applicant]
Maskey, Sameer, and Julia Hirschberg. “Comparing lexical, acoustic/prosodic, structural and discourse features for speech summarization.” Ninth European conference on speech communication and technology. 2005. (Year: 20… [cited by examiner]
Hori, Chiori, and Sadaoki Furui. “A new approach to automatic speech summarization.” IEEE Transactions on Multimedia 5.3 (2003): 368-378. (Year: 2003). [cited by examiner]
Furui, Sadaoki, et al. “Speech-to-text and speech-to-speech summarization of spontaneous speech.” IEEE Transactions on Speech and Audio Processing 12.4 (2004): 401-408. (Year: 2004). [cited by examiner]
Tiun, Sabrina, Rosni Abdullah, and Tang Enya Kong. “Automatic topic identification using ontology hierarchy.” Computational Linguistics and Intelligent Text Processing: Second International Conference, published in 2001… [cited by examiner]
“SCANMail: Audio Navigation in the Voicemail Domain.” Proceedings of the First International Conference of Human Language Technology Research. Bacchiani, et al. 2001 (3 pgs). [cited by applicant]
“Extractive Speech Summarization Using Rhetorical Structure Modeling.” IEEE Transactions of Audio Speech and Language Processing. Zhang, et al. Aug. 6, 2010. (11 pgs). [cited by applicant]
Non-Final Office Action Issued in U.S. Appl. No. 16/584,893. Mar. 15, 2021. (45 pgs). [cited by applicant]
Notice of Allowance Issued in U.S. Appl. No. 16/584,893. Jul. 22, 2021. (7 pgs). [cited by applicant]