IP Library › Granted Patent US 11,749,241
Granted Patent B2
US 11,749,241 · App. 17/172,201 · Granted Sep 5, 2023

Systems and methods for transforming digitial audio content into visual topic-based segments

Inventors: Michael Kakoyiannis (Garden City, NY); Sherry Mills (New York, NY); Christoforos Lambrou (Nicosia, CY); Vladimir Canic (Belgrade, RS); Srdjan Jovanovic (Belgrade, RS)
Assignee: TREE GOAT MEDIA, INC.
G10H1/0008G06F16/685G06F16/686G10H2220/106
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,749,241
App. No.
17/172,201
Filed
Feb 10, 2021
Granted
Sep 5, 2023
Kind
B2
Art Unit
2837
USPC
84/609
Abstract

A system for platform-independent visualization of audio content, in particular audio tracks utilizing a central computer system in communication with user devices via a computer network. The central system utilizes various algorithms to identify spoken content from audio tracks and selects visual assets associated with the identified content. Thereafter, a visualized audio track is available for users to listen and view. Audio tracks, for example Podcasts, may be segmented into topical audio segments based upon themes or topics, with segments from disparate podcasts combined into a single listening experience, based upon certain criteria, e.g., topics, themes, keywords, and the like.

Claims (32)

1. A method for packaging audio content by an audio content system to facilitate in viewing, searching and/or sharing of said audio content comprising:

(a) providing an audio track and storing the audio track in a data storage;

(b) converting content contained within the audio to an audio text;

(c) using AI to identify interest segments of said audio text;

(d) displaying said interest segments on an app or computer;

(e) enabling a user to select one of said interest segments being displayed on said app or computer to enable said interest segment to be displayed in association with a portion of said audio text that exists immediately prior to and immediately after said interest segment selected by said user;

(f) enabling said user to hear a portion of audio track that includes said interest segment selected by said user; and,

(g) enabling said user to hear said audio track from beginning to end.

2. The method as defined in claim 1 , further including the steps of:

pairing at least one visual asset to one or more of said interest segments by:

(i) determining a proposed set of visual assets by performing an automated analysis of said one or more interest segments; and,

(ii) using AI to associate a particular visual asset to said one or more interest segments.

3. The method as defined in claim 2 , wherein the at least one visual asset is one of an image, photograph, video, cinemograph, video loop, and/or collage.

4. The method as defined in claim 1 , wherein said audio track is a podcast.

5. The method as defined in claim 1 , wherein a voice recognition module is configured to convert content contained within said audio track to an audio text and a segmentation module is used to divide said audio signal into the at least one audio segment based on keywords derived from said audio text.

6. The method as defined in claim 1 , wherein a voice recognition module is configured said interest segments directly from said content contained within said audio signal and/or audio text, and wherein automated analysis of said audio signal and/or audio text includes analyzing said audio signal and/or audio text based on extracted keywords.

7. The method as defined in claim 1 , further one or more interest segments in an associated database.

8. A content system for platform-independent visualization of audio content, comprising:

(a) a central computer system comprising: a processor, a memory in communication with the processor, the memory storing instructions which are executed by the processor;

(b) an audio segmenting subsystem including an audio resource containing at least one audio track, the audio segmenting subsystem configured to divide the at least one audio track into at least one audio segment and generate an indexed audio segment by associating the at least one audio segment with at least one audio textual element, wherein the at least one audio textual element relates to a spoken content captured within the audio track; and,

(c) a visual subsystem including a video resource storing at least one visual asset, the visual subsystem configured to generate an indexed visual asset by associating at least one visual textual element to the at least one visual asset, wherein the content system is configured to generate a packaged audio segment by associating the indexed audio segment with the indexed visual asset;

wherein the at least one audio track and the at least one visual asset are not relationally associated, and wherein the processor is configured to:

cause a plurality of indexed visual assets, including the indexed visual asset, to display on a user device;

(ii) cause the indexed audio segment to play on the user device and cause the indexed visual asset to display on the user device;

(iii) while playing the indexed audio segment, cause a second plurality of indexed visual assets to display on the user device, based on a textual association between the second plurality of indexed visual assets and the indexed visual asset.

9. The content system as defined in claim 8 , further comprising, prior to generating the indexed audio segment and associating the at least one visual asset to the indexed audio segment, receiving a request from a user device to listen to the audio track, and in response to the request:

(a) causing the user device to play the audio track; and

(b) causing the user device to display the visual asset during play of the corresponding at least one audio segment of the audio track.

10. The content system as defined in claim 8 , wherein the plurality of indexed visual assets includes a plurality of background images and a plurality of foreground images, further comprising, when associating the at least one visual asset to the indexed audio segment:

(a) selecting a background image from the plurality of background images based on the indexed audio segment;

(b) selecting a foreground image from the plurality of foreground images based on the indexed audio segment; and

(c) overlaying the foreground image on the background image to produce the at least one visual asset.

Assignments (2)
ENTITY CONVERSION Recorded Jun 21, 2021
From: TREE GOAT MEDIA, LLC
To: TREE GOAT MEDIA, INC.
Reel/Frame 056635/0689 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2021
From: KAKOYIANNIS, MICHAEL; MILLS, SHERRY; LAMBROU, CHRISTOFOROS; CANIC, VLADIMIR; JOVANOVIC, SRDJAN
To: TREE GOAT MEDIA, LLC
Reel/Frame 056138/0093 →
Continuity (5)
Continuation 16506231 · Jul 9, 2019
Provisional Application 62814018 · Mar 5, 2019
Provisional Application 62695439 · Jul 9, 2018
Related Publication 20210166666A1 · Jun 3, 2021
Related Publication 20230230564A9 · Jul 20, 2023