Generating and enhancing digital video components
The technology is directed to artificial intelligence (AI) powered tools that can enhance existing digital video components and simplify and automate the creation of new digital video components. The technology includes a digital video component creation tool that leverages existing assets to generate digital video components, a voice-over tool that can add voice-overs, generated from text to video components, and a video component evaluation tool that can evaluate video components for conformity with attributes associated with metrics for video creatives.
1 . A method for assessing a video component comprising:
receiving, by one or more processors, the video component;
evaluating, by the one or more processors, using a multimodal video component evaluation tool, the received video component relative to audio and video metrics, the evaluation comparing attributes within the video component to the audio and video metrics to determine whether each of the audio and video metrics are met by the received video component, wherein the multimodal video component evaluation tool comprises a product/brand engine configured to:
receive as input, pixel information corresponding to individual video frames of the received video component, and
process, using a pixel-level image models, the pixel information by matching objects identified in the individual video frames to a taxonomy of brands and/or products to identify brands and/or products within the video component;
generating, by the one or more processors, based on the evaluation, a result indicating that one or more of the audio and video metrics are not met, wherein the result comprises an indication that the product/brand engine did not detect a particular brand and/or product;
generating, by a generative artificial intelligence (AI) model, content representing the particular brand and/or product based on the result;
integrating, by the generative AI model, the generated content representing the particular brand and/or product into the video component; and
outputting, by the one or more processors, the video component including the particular brand and/or product.
2 . The method of claim 1 , wherein the multimodal video component evaluation tool further comprises a logo detection engine configured to detect logos within the video component and the result comprises an indication of whether the logo detection engine detects one or more logos.
3 . The method of claim 1 , wherein the multimodal video component evaluation tool further comprises an audio annotation engine configured to detect audio annotations within an audio track of the video component and the result comprises an indication of whether the audio annotation engine detects one or more predefined audio annotations comprising pieces of music or speech, lengths of music or speech, or music or speech having a particular volume.
4 . The method of claim 1 , wherein the multimodal video component evaluation tool further comprises an audio transcript engine configured to detect keywords from a taxonomy based on a transcript of an audio track of the video component and the result comprises an indication of whether the audio transcript engine detects one or more brands and/or products within audio of the video component.
5 . The method of claim 1 , wherein the multimodal video component evaluation tool further comprises a promotion engine configured to detect promotions within a transcript of an audio track of the video component and the result comprises an indication of whether the promotion engine detects one or more promotions within the transcript.
6 . The method of claim 1 , further comprising:
identifying, by the one or more processors, the attributes within the video component; and
summarizing, by the one or more processors, the attributes within the video component.
7 . The method of claim 1 , wherein the attributes within the video component comprise one or more of content within the received video component or one or more visual elements within the received video component.
8 . The method of claim 1 , wherein the attributes within the video component comprise one or more visual elements within the received video component, the one or more visual elements comprising at least one of duration of received video content, aspect ratio of received video content, or visual effects within received video content.
9 . A system comprising:
one or more processors configured to:
receive a video component;
evaluate, using a multimodal video component evaluation tool, the received video component relative to audio and video metrics, the evaluation comparing attributes within the video component to the audio and video metrics to determine whether each of the audio and video metrics are met by the received video component, wherein the multimodal video component evaluation tool comprises a product/brand engine configured to:
receive as input, pixel information corresponding to individual video frames of the received video component, and
process, using a pixel-level image models, the pixel information by matching objects identified in the individual video frames to a taxonomy of brands and/or products to identify brands and/or products within the video component;
generate, based on the evaluation, a result indicating that one or more of the audio and video metrics are not met, wherein the result comprises an indication that the product/brand engine did not detect a particular brand and/or product;
generate, by a generative artificial intelligence (AI) model, content representing the particular brand and/or product based on the result;
integrate, by the generative AI model, the generated content representing the particular brand and/or product into the video component; and
output the video component including the particular brand and/or product.
10 . The system of claim 9 , wherein the multimodal video component evaluation tool further comprises a logo detection engine configured to detect logos within the video component and the result comprises an indication of whether the logo detection engine detects one or more logos.
11 . The system of claim 9 , wherein the multimodal video component evaluation tool further comprises an audio annotation engine configured to detect audio annotations within an audio track of the video component and the result comprises an indication of whether the audio annotation engine detects one or more predefined audio annotations comprising pieces of music or speech, lengths of music or speech, or music or speech having a particular volume.
12 . The system of claim 9 , wherein the multimodal video component evaluation tool further comprises an audio transcript engine configured to detect keywords from a taxonomy based on a transcript of an audio track of the video component and the result comprises an indication of whether the audio transcript engine detects one or more brands and/or products within audio of the video component.
13 . The system of claim 9 , wherein the multimodal video component evaluation tool further comprises a promotion engine is configured to detect promotions within a transcript of an audio track of the video component and the result comprises an indication of whether the promotion engine detects one or more promotions within the transcript.
14 . The system of claim 9 , wherein the one or more processors are further programmed to:
identify the attributes within the video component; and
summarize the attributes within the video component.
15 . The system of claim 9 , wherein the attributes within the video component comprise one or more of content within the received video component or one or more visual elements within the received video component.
16 . The system of claim 9 , wherein the attributes within the video component comprise one or more visual elements within the received video component, the one or more visual elements comprising at least one of duration of received video content, aspect ratio of received video content, or visual effects within received video content.
17 . A non-transitory computer readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to:
receive a video component;
evaluate, using a multimodal video component evaluation tool, the received video component relative to audio and video metrics, the evaluation comparing attributes within the video component to the audio and video metrics to determine whether each of the audio and video metrics are met by the received video component, wherein the multimodal video component evaluation tool comprises a product/brand engine configured to:
receive as input, pixel information corresponding to individual video frames of the received video component, and
process, using a pixel-level image models, the pixel information by matching objects identified in the individual video frames to a taxonomy of brands and/or products to identify brands and/or products within the video component;
generate, based on the evaluation, a result indicating that one or more of the audio and video metrics are not met, wherein the result comprises an indication that the product/brand engine did not detect a particular brand and/or product;
generate, by a generative artificial intelligence (AI) model, content representing the particular brand and/or product based on the result;
add, by the generative AI model, the generated content representing the particular brand and/or product to the video component; and
output the video component including the particular brand and/or product.
18 . The non-transitory computer readable medium of claim 17 , wherein the multimodal video component evaluation tool further comprises a logo detection engine configured to detect logos within the video component and the result comprises an indication of whether the logo detection engine detects one or more logos.
19 . The non-transitory computer readable medium of claim 17 , wherein the multimodal video component evaluation tool further comprises an audio annotation engine configured to detect audio annotations within an audio track of the video component and the result comprises an indication of whether the audio annotation engine detects one or more predefined audio annotations comprising pieces of music or speech, lengths of music or speech, or music or speech having a particular volume.
20 . The non-transitory computer readable medium of claim 17 , wherein the multimodal video component evaluation tool further comprises:
an audio transcript engine configured to detect keywords from a taxonomy based on a transcript of an audio track of the video component and the result comprises an indication of whether the audio transcript engine detects one or more brands and/or products within the transcript; and
a promotion engine is configured to detect promotions within the transcript and the result comprises an indication of whether the promotion engine detects one or more promotions within the transcript.