Content system with event identification and audio-based editing feature
In one aspect, an example method includes (i) obtaining, by a computing system, video data representing video content; (ii) analyzing, by the computing system, the video data to identify an event that is a subject of the video content; (iii) using, by the computing system, the identified event as a basis to select audio content; and (iv) performing, by the computing system, an operation that facilitates editing the video content to include the selected audio content.
1 . A method comprising:
obtaining, by a computing system, video data representing video content;
providing the video data and closed-captioning data associated with the video content to a trained model, wherein the trained model is configured to use at least video data and closed-captioning data as runtime input data to generate event identification data as runtime output data;
responsive to providing the video data and the closed-captioning data associated with the video data to the trained model, receiving from the trained model, corresponding generated event identification data;
using, by the computing system, the generated event identification data as a basis to select audio content; and
performing, by the computing system, an operation that facilitates editing the video content to include the selected audio content.
2 . The method of claim 1 , wherein the received event identification data further includes event position data that specifies a temporal position of the event within the video content.
3 . The method of claim 1 , wherein the received event identification data further includes event duration data that specifies a duration of the event within the video content.
4 . The method of claim 1 , wherein the model was trained by providing to the model as training data, multiple instances of training video data and for each instance of training video data, corresponding training event identification data.
5 . The method of claim 1 , wherein using the generated event identification data as the basis to select the audio content comprises using mapping data to map the generated event identification data to the selected audio content.
6 . The method of claim 1 , wherein performing the operation facilitates editing the video content to include the selected audio content comprises:
editing the video content by adding the selected audio content to the video content.
7 . The method of claim 1 , wherein performing the operation facilitates editing the video content to include the selected audio content comprises:
editing the video content by replacing an existing audio content portion of the video content with the selected audio content.
8 . The method of claim 1 , wherein performing the operation that facilitates editing the video content to include the selected audio content comprises:
prompting, via a user interface, a proposed editing of the video content to include the selected audio content; and
performing the proposed editing or a variation thereof based on input received via the user interface.
9 . The method of claim 1 , further comprising transmitting, by the computing system, the edited video content to a content-presentation device for presentation.
10 . The method of claim 1 , further comprising presenting, by a content-presentation device of the computing system, the edited video content.
11 . A computing system configured for performing a set of acts comprising:
obtaining, by a computing system, video data representing video content;
providing the video data and closed-captioning data associated with the video content to a trained model, wherein the trained model is configured to use at least video data and closed-captioning data as runtime input data to generate event identification data as runtime output data;
responsive to providing the video data and the closed-captioning data associated with the video data to the trained model, receiving from the trained model, corresponding generated event identification data;
using, by the computing system, the generated event identification data as a basis to select audio content; and
performing, by the computing system, an operation that facilitates editing the video content to include the selected audio content.
12 . The computing system of claim 11 , wherein the received event identification data further includes event position data that specifies a temporal position of the event within the video content.
13 . The computing system of claim 11 , wherein the received event identification data further includes event duration data that specifies a duration of the event within the video content.
14 . The computing system of claim 11 , wherein the model was trained by providing to the model as training data, multiple instances of training video data and for each instance of training video data, corresponding training event identification data.
15 . The computing system of claim 11 , wherein using the generated event identification data as the basis to select the audio content comprises using mapping data to map the generated event identification data to the selected audio content.
16 . The computing system of claim 11 , wherein performing the operation facilitates editing the video content to include the selected audio content comprises:
editing the video content by adding the selected audio content to the video content.
17 . The computing system of claim 11 , wherein performing the operation facilitates editing the video content to include the selected audio content comprises:
editing the video content by replacing an existing audio content portion of the video content with the selected audio content.
18 . The computing system of claim 11 , wherein performing the operation that facilitates editing the video content to include the selected audio content comprises:
prompting, via a user interface, a proposed editing of the video content to include the selected audio content; and
performing the proposed editing or a variation thereof based on input received via the user interface.
19 . The computing system of claim 11 , further comprising transmitting, by the computing system, the edited video content to a content-presentation device for presentation.
20 . A non-transitory computer-readable medium having stored thereon program instructions that upon execution by a computing system, cause performance of a set of acts comprising:
obtaining, by a computing system, video data representing video content;
providing the video data and closed-captioning data associated with the video content to a trained model, wherein the trained model is configured to use at least video data and closed-captioning data as runtime input data to generate event identification data as runtime output data;
responsive to providing the video data and the closed-captioning data associated with the video data to the trained model, receiving from the trained model, corresponding generated event identification data;
using, by the computing system, the generated event identification data as a basis to select audio content; and
performing, by the computing system, an operation that facilitates editing the video content to include the selected audio content.