Automatically enhancing streaming media using content transformation
A method includes receiving media content comprising audio data for distribution through content distribution platform that requires the media content to include video content, transforming the audio data into textual content, determining, based on a search of a searchable database, that the textual content of the audio data matches characteristics of visual data in the searchable database, integrating the visual data having the matched characteristics with the media content to create an augmented content stream in response to the determination that the textual content of the audio data matches the characteristics of the visual data, and distributing the augmented content stream through the content distribution platform that requires the media content to include video content.
1 . A method, comprising:
receiving audio data for distribution through content distribution platform, wherein the audio data includes first speech of a first person and second speech of a second person;
differentiating the first speech of the first person in the audio data from the second speech of the second person in the audio data;
transforming the audio data into textual content;
flagging a first set of the textual content as being the first speech of the first person in the audio data;
flagging a second set of the textual content as being the second speech of the second person in the audio data;
selecting, based on a search of a searchable database using the first set of textual content instead of the second set of textual content, visual data in the searchable database based on the first set of the textual content being flagged as being the first speech of the first person in the audio data;
integrating the visual data selected based on the first set of textual content with the audio data to create an augmented content stream that includes the visual data, audio of the first speech of the first person, and audio of the second speech of the second person; and
distributing, to a user that has requested the audio data, the augmented content stream through the content distribution platform.
2 . The method of claim 1 , further comprising:
detecting an annotation located at a particular temporal location within the audio data; and
wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying, based on the annotation, the visual data with the audio data at the particular temporal location within the media content.
3 . The method of claim 2 , wherein the annotation specifies one or more visual data characteristics.
4 . The method of claim 2 , wherein integrating the visual data with the audio data to create the augmented content stream further comprises editing the visual data based on one or more visual data characteristics.
5 . The method of claim 2 , wherein the annotation specifies that visual data cannot be overlaid with the audio data at the particular temporal location.
6 . The method of claim 1 , further comprising:
determining a first context of the audio data based on the textual content of the audio data;
determining a second context of the visual data based on characteristics of the visual data in the searchable database; and
determining that the textual content of the audio data matches the characteristics of the visual data in the searchable database based on the first context matching the second context.
7 . The method of claim 6 , further comprising:
identifying a particular temporal location within the audio data based on the first context of the audio data; and
wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying the visual data with the audio data at the particular temporal location within the audio data.
8 . A system comprising:
one or more processors; and
one or more memory elements including instructions that, when executed, cause the one or more processors to perform operations including:
receiving audio data for distribution through content distribution platform, wherein the audio data includes first speech of a first person and second speech of a second person;
differentiating the first speech of the first person in the audio data from the second speech of the second person in the audio data;
transforming the audio data into textual content;
flagging a first set of the textual content as being the first speech of the first person in the audio data;
flagging a second set of the textual content as being the second speech of the second person in the audio data;
selecting, based on a search of a searchable database using the first set of textual content instead of the second set of textual content, visual data in the searchable database based on the first set of the textual content being flagged as being the first speech of the first person in the audio data;
integrating the visual data selected based on the first set of textual content with the audio data to create an augmented content stream that includes the visual data, audio of the first speech of the first person, and audio of the second speech of the second person; and
distributing, to a user that has requested the audio data, the augmented content stream through the content distribution platform.
9 . The system of claim 8 , the operations further comprising:
detecting an annotation located at a particular temporal location within the audio data; and
wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying, based on the annotation, the visual data with the audio data at the particular temporal location within the audio data.
10 . The system of claim 9 , wherein the annotation specifies one or more visual data characteristics.
11 . The system of claim 9 , wherein integrating the visual data with the audio data to create the augmented content stream further comprises editing the visual data based on one or more visual data characteristics.
12 . The system of claim 9 , wherein the annotation specifies that visual data cannot be overlaid with the audio data at the particular temporal location.
13 . The system of claim 8 , the operations further comprising:
determining a first context of the audio data based on the textual content of the audio data;
determining a second context of the visual data based on characteristics of the visual data in the searchable database; and
determining that the textual content of the audio data matches the characteristics of the visual data in the searchable database based on the first context matching the second context.
14 . The system of claim 13 , the operations further comprising:
identifying a particular temporal location within the audio data based on the first context of the audio data; and
wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying the visual data with the audio data at the particular temporal location within the audio data.
15 . A non-transitory computer-storage medium encoded with instructions that when executed by a distributed computing system cause the distributed computing system to perform operations comprising:
receiving audio data for distribution through content distribution platform, wherein the audio data includes first speech of a first person and second speech of a second person;
differentiating the first speech of the first person in the audio data from the second speech of the second person in the audio data;
transforming the audio data into textual content;
flagging a first set of the textual content as being the first speech of the first person in the audio data;
flagging a second set of the textual content as being the second speech of the second person in the audio data;
selecting, based on a search of a searchable database using the first set of textual content instead of the second set of textual content, visual data in the searchable database based on the first set of the textual content being flagged as being the first speech of the first person in the audio data;
integrating the visual data selected based on the first set of textual content with the audio data to create an augmented content stream that includes the visual data, audio of the first speech of the first person, and audio of the second speech of the second person; and
distributing, to a user that has requested the audio data, the augmented content stream through the content distribution platform.
16 . The non-transitory computer-storage medium of claim 15 , the operations further comprising:
detecting an annotation located at a particular temporal location within the audio data; and
wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying, based on the annotation, the visual data with the audio data at the particular temporal location within the audio data.
17 . The non-transitory computer-storage medium of claim 16 , wherein the annotation specifies one or more visual data characteristics.
18 . The non-transitory computer-storage medium of claim 16 , wherein integrating the visual data with the audio data to create the augmented content stream further comprises editing the visual data based on one or more visual data characteristics.
19 . The non-transitory computer-storage medium of claim 16 , wherein the annotation specifies that visual data cannot be overlaid with the audio data at the particular temporal location.
20 . The non-transitory computer-storage medium of claim 15 , the operations further comprising:
determining a first context of the audio data based on the textual content of the audio data;
determining a second context of the visual data based on characteristics of the visual data in the searchable database; and
determining that the textual content of the audio data matches the characteristics of the visual data in the searchable database based on the first context matching the second context.
21 . The method of claim 1 , further comprising:
emphasizing portions of the first set of textual content based on audio characteristics other than identification of the first person, wherein selecting, based on a search of a searchable database using the first set of textual content instead of the second set of textual content, visual data in the searchable database based on the first set of the textual content being flagged as being the first speech of the first person in the audio data comprises selecting the visual data based on the emphasized portions of the first set of textual content.