IP Library Granted Patent US 9,542,393
Granted Patent B2
US 9,542,393 · App. 14/843,912 · Granted Jan 10, 2017

Method and system for indexing and searching timed media information based upon relevance intervals

Inventors: Michael Scott Morton (Washington, DC); Sibley Verbeck Simon (Wahington, DC); Noam Carl Unger (Somerville, MA); Robert Rubinoff (Potomac, MD); Anthony Ruiz Davis (Takoma Park, MD); Kyle Aveni-Deforge (Columbia, SC)
Assignee: Streamsage, Inc.
G06F17/3002G06F17/30029G06F17/30064G06F17/30424G06F17/30613G06F17/30796G06F17/274G06F17/2705G06F17/275G06F17/30684Y10S707/913Y10S707/99931Y10S707/99935
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,542,393
App. No.
14/843,912
Granted
Jan 10, 2017
Kind
B2
Abstract

A method and system for indexing, searching, and retrieving information from timed media files based upon relevance intervals. The method and system for indexing, searching, and retrieving this information is based upon relevance intervals so that a portion of a timed media file is returned, which is selected specifically to be relevant to the given information representations, thereby eliminating the need for a manual determination of the relevance and avoiding missing relevant portions. The timed media includes streaming audio, streaming video, timed HTML, animations such as vector-based graphics, slide shows, other timed media, and combinations thereof.

Claims (57)

1. A method comprising:

receiving, by a computing device, a file comprising audio;

recognizing speech in the file comprising the audio;

parsing the speech to determine a grammatical structure of the speech;

based on the grammatical structure of the speech, dividing the speech into sentences; and

determining a topic of one of the sentences.

2. The method of claim 1 , wherein dividing the speech into sentences comprises using a set of rules to determine a most likely sentence boundary for the speech.

3. The method of claim 2 , wherein the set of rules to determine the most likely sentence boundary for the speech comprises rules for determining sentence boundaries based upon one or more of word sequences, pauses, parts of speech, grammatical data, and prosodic cues.

4. The method of claim 2 , comprising:

determining a sentence number of a particular sentence in the speech by counting a number of sentences in the speech before the particular sentence.

5. The method of claim 1 , comprising:

determining a pronoun in the speech; and

determining a reference to which the pronoun in the speech refers.

6. The method of claim 5 , wherein the determining the reference to which the pronoun in the speech refers comprises:

determining a group of potential antecedents for the pronoun in the speech; and

filtering the group of potential antecedents for the pronoun based on whether each of the potential antecedents can represent one or more of a human, a group of humans, or a gendered non-human.

7. The method of claim 6 , wherein filtering the group of potential antecedents for the pronoun is further based on one or more semantic constraints on the potential antecedents.

8. A method comprising:

receiving, by a computing device, a file comprising a video;

identifying speech in the video;

determining a grammatical structure of the speech in the video;

based on the grammatical structure of the speech in the video, determining one or more sentences in the speech; and

determining a concept of one of the one or more sentences.

9. The method of claim 8 , comprising:

reducing words of the speech in the video to canonical form; and

adding the canonical form of the words of the speech to an index.

10. The method of claim 8 , comprising:

determining a grammatical structure of the one or more sentences in the speech; and

based on the grammatical structure of the one or more sentences in the speech, determining centrality of the concept of the one of the one or more sentences.

11. The method of claim 10 , wherein determining the grammatical structure of the one or more sentences in the speech comprises:

determining a part of speech of each word in the one or more sentences.

12. The method of claim 11 , comprising:

assigning a centrality weight to each word in the one or more sentences, based on the part of speech of each word in the one or more sentences.

13. The method of claim 8 , comprising:

determining a semantic meaning of each word in the one or more sentences in the speech; and

filtering words in the one or more sentences based on the semantic meaning of each word.

14. The method of claim 13 , wherein filtering the words in the one or more sentences based on the semantic meaning of each word comprises:

determining whether a word of the words in the one or more sentences is a meaningful word; and

excluding the word if the word is not a meaningful word.

15. The method of claim 14 , wherein the meaningful word is not a conjunction, an article, or a preposition.

16. A method comprising:

receiving, by a computing device, a media file;

receiving a text transcript associated with the media file;

parsing the text transcript to determine a grammatical structure of words in the text transcript;

based on the grammatical structure of the words in the text transcript, dividing the words in the text transcript into sentences; and

determining a topic of one of the sentences.

17. The method of claim 16 , comprising:

determining a corpus based on centrality of a concept in a sentence of the sentences; and

creating a search index using the corpus.

18. The method of claim 17 , comprising:

determining a centrality score for each word in the sentence based on a structure of the sentence; and

determining the centrality of the concept in the sentence based on the centrality score for each word in the sentence.

19. The method of claim 16 , comprising:

determining, for each of the sentences, whether the sentence is a topic boundary.

20. The method of claim 16 , comprising:

determining a plurality of named entities in the text transcript associated with the media file; and

determining whether the plurality of named entities in the text transcript are referring to a same entity.

Assignments (3)
MERGER Recorded Apr 17, 2017
From: STREAMSAGE, INC.
To: COMCAST CABLE COMMUNICATIONS MANAGEMENT, LLC
Reel/Frame 042028/0670 →
MERGER Recorded Apr 17, 2017
From: STREAMSAGE, INC.
To: COMCAST CABLE COMMUNICATIONS MANAGEMENT, LLC
Reel/Frame 042028/0766 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2016
From: MORTON, MICHAEL SCOTT; SIMON, SIBLEY VERBECK; UNGER, NOAM CARL; RUBINOFF, ROBERT; DAVIS, ANTHONY RUIZ; AVENI-DEFORGE, KYLE
To: STREAMSAGE, INC.
Reel/Frame 040393/0446 →
Continuity (8)
Continuation 14176367 · Feb 10, 2014
Continuation 13955582 · Jul 31, 2013
Continuation 13347914 · Jan 11, 2012
Continuation 12349934 · Jan 7, 2009
Continuation 10364408 · Feb 12, 2003
Continuation In Part 09611316 · Jul 6, 2000
Provisional Application 60356632 · Feb 12, 2002
Related Publication 20160098396A1 · Apr 7, 2016