Time ordered indexing of an information stream
View Patent ↗Methods and apparatuses in which two or more types of attributes from an information stream are identified. Each of the identified attributes from the information stream is encoded. A time ordered indication is assigned with each of the identified attributes. Each of the identified attributes shares a common time reference measurement. A time ordered index of the identified attributes is generated.
1. A method, comprising:
converting, by a computer, spoken words in an information stream to written text, the information stream containing at least audio information; and
generating, by the computer, a separate encoded file for every spoken word, wherein each encoded file shares a common time reference.
2. The method of claim 1 , wherein the information stream contains audio-visual information and the each encoded file shares a common time reference to a video frame.
3. The method of claim 2 , further comprising:
detecting shot change information;
creating thumbnail images for each shot change;
generating an encoded file for every shot change, each encoded file for the respective shot change containing a time ordered indication referencing to one or more corresponding spoken words.
4. The method of claim 2 , further comprising:
generating a link to the video frame based upon a user selecting one or more of the spoken words.
5. The method of claim 2 , further comprising:
generating a link to respective material based upon the spoken words and synchronizing a display of the link in less than five seconds from analyzing the information stream.
6. A non-transitory machine-readable medium storing instructions, which when executed by a machine, cause the machine to
convert spoken words in an information stream to written text, the information stream containing audio-visual information; and
generate a separate encoded file for every spoken word, each encoded file containing a time ordered indication reference to a respective video frame.
7. The machine-readable medium of claim 6 , storing further instructions which when executed cause the machine to:
detect shot change information;
create thumbnail images for each shot change; and
generate an encoded file for every shot change, each encoded file for the respective shot change containing a time ordered indication referencing to one or more corresponding spoken words.
8. The machine-readable medium of claim 6 , storing further instructions which when executed cause the machine to:
generate a link to respective material based upon the spoken words and synchronizing a display of the link in less than five seconds from analyzing the information stream.
9. The machine-readable medium of claim 6 , storing further instructions which when executed cause the machine to:
individually synchronize each word in a transcript from the information stream to be frame accurate to corresponding video data based upon both sharing a common time reference.
10. An apparatus, comprising:
a non-transitory computer readable medium storing instructions; and
at least one processor, the instructions executable on the at least one processor to:
convert spoken words in an information stream to written text, the information stream containing at least audio information; and
generate a separate encoded file for every spoken word, wherein each encoded file shares a common time reference.
11. The apparatus of claim 10 , wherein the information stream contains audio-visual information and the each encoded file shares a common time reference to a video frame.
12. The apparatus of claim 10 , wherein the instructions are executable on the at least one processor to:
detect shot change information;
create thumbnail images for each shot change;
generate an encoded file for every shot change, each encoded file for the respective shot change containing a time ordered indication referencing to one or more corresponding spoken words.
13. The apparatus of claim 11 , wherein the instructions are executable on the at least one processor to:
generate a link to the video frame based upon a user selecting one or more of the spoken words.