Systems and methods for generating customized media content
Systems and methods for generating customized media content are provided (for example, trailers or recaps for a movie or television show). For example, the system may generate a trailer to a television show that is tailored to a specific user, such as a trailer that includes action-specific scenes for a user who often views content in an action-based genre. The system may receive as an input a text or voice-based query from the user (or the system may be automated and may receive user historical data as an input). Based on the input, one or more computing models may generate a text-based narrative to be used with the customized content. Once the narrative is generated, the one or more models may then identify specific video frames to include in the customized content. The video frames may then be stitched together and the customized content may be generated using the stitched video frames and the narrative.
1 . A method comprising:
training one or more machine learning models using a set of media content stored in memory, the set of media content including at least: one or more movies and one or more televisions shows;
generating, using a first machine learning model, a plurality of video segments of the set of media content;
receiving, using one or more processors, a first input associated with a first user, the first input being a natural language input that is indicative of first customized media content to be generated based on at least a first movie or first television show of the set of media content,
wherein the first input includes at least one of: a text-based or voice query received from a user device and historical data associated with the first user,
and wherein the first customized media content includes a custom trailer or recap for the first movie or first television show;
generating, by a second machine learning model and based on the first input, a first text-based narrative associated with the first movie or first television show;
receiving, by a third machine learning model, the first text-based narrative, wherein the first text-based narrative comprises at least some unique text that differs from text included in an existing text-based narrative of the media content and is based on the natural language input instead of the existing text-based narrative of the first movie or first television show;
identifying, by the third machine learning model and based on the first text-based narrative, one or more first video frames of the first movie or first television show;
identifying, by the third machine learning model, one or more video segments of the plurality of video segments that includes the one or more first video frames; and
generating, by the one or more machine learning models, the first customized media content using the first text-based narrative and the one or more video segments, wherein the first machine learning model, second machine learning model, and third machine learning model are different models.
2 . The method of claim 1 , further comprising:
receiving a second input associated with a second user, the second input being indicative of second customized media content to be generated based on the first movie or first television show;
generating, by the second machine learning model and based on the second input, a second text-based narrative;
identifying, by the third machine learning model and based on the second text-based narrative, one or more second video frames of the set of media content; and
generating, by the one or more machine learning models, the second customized media content using the second text-based narrative and the one or more second video frames, wherein the second customized media content is different than the first customized media content, and wherein the second customized media content is tailored to the second user.
3 . The method of claim 1 , further comprising:
generating a computer-generated voiceover based on the first text-based narrative, wherein the first customized media content includes the computer-generated voiceover.
4 . The method of claim 1 , wherein the machine learning models comprise at least one of: a large language model, a large multimodal visual model, and a shot detection model.
5 . A method comprising:
receiving, using one or more processors, a first input associated with a first user, the first input being a natural language input that is indicative of first customized media content to be generated for first media content;
generating, by a first machine learning model and based on the first input, a first text-based narrative for the first customized media content, wherein the first text-based narrative comprises at least some unique text that differs from text included in an existing text-based narrative of the first media content and is based on the natural language input instead of the existing text-based narrative of the first media content;
identifying, by a second machine learning model and based on the first text-based narrative, one or more first video frames of the first media content; and
generating the first customized media content using the first text-based narrative and the one or more first video frames, wherein the first machine learning model and second machine learning model are different models.
6 . The method of claim 5 , further comprising:
receiving, using the one or more processors, a second input associated with a second user, the second input being indicative of second customized media content to be generated based on the first media content;
generating, by the first machine learning model and based on the second input, a second text-based narrative;
identifying, by the second machine learning model and based on the second text-based narrative, one or more second video frames of the first media content; and
generating the second customized media content using the second text-based narrative and the one or more second video frames, wherein the second customized media content is different than the first customized media content.
7 . The method of claim 5 , further comprising:
generating a first video segment and a second video segment for first media content of the first media content; and
determining that a first video frame of the one or more first video frames is included within the first video segment,
wherein the first customized media content includes the first video segment instead of the second video segment based on the determination that the first video frame is included within the first video segment.
8 . The method of claim 5 , further comprising:
generating, a computer-generated voiceover based on the first text-based narrative, wherein the first customized media content includes the computer-generated voiceover.
9 . The method of claim 5 , wherein the first input is a text-based or voice query received from a user device.
10 . The method of claim 5 , wherein the first input is historical data associated with the first user.
11 . The method of claim 5 , further comprising:
training one or more computing models using a set of media content.
12 . A system comprising:
memory that stores computer-executable instructions; and
one or more processors configured to access the memory and execute the computer-executable instructions to:
receive, using one or more processors, a first input associated with a first user, the first input being a natural language input that is indicative of first customized media content to be generated for first media content;
generate, by a first machine learning model and based on the first input, a first text-based narrative for the first customized media content, wherein the first text-based narrative comprises at least some unique text that differs from text included in an existing text-based narrative of the first media content and is based on the natural language input instead of the existing text-based narrative of the first media content;
identify, by a second machine learning model and based on the first text-based narrative, one or more first video frames of the first media content; and
generate the first customized media content using the first text-based narrative and the one or more first video frames, wherein the first machine learning model and second machine learning model are different models.
13 . The system of claim 12 , wherein the one or more processors are further configured to execute the computer-executable instructions to:
receive a second input associated with a second user, the second input being indicative of second customized media content to be generated based on the first media content;
generate, by the first machine learning model and based on the second input, a second text-based narrative;
identify, by the second machine learning model and based on the second text-based narrative, one or more second video frames of the first media content; and
generate second customized media content using the second text-based narrative and the one or more second video frames, wherein the second customized media content is different than the first customized media content.
14 . The system of claim 12 , wherein the one or more processors are further configured to execute the computer-executable instructions to:
generate a first video segment and a second video segment for first media content of the first media content; and
determine that a first video frame of the one or more first video frames is included within the first video segment,
wherein the first customized media content includes the first video segment instead of the second video segment based on the determination that the first video frame is included within the first video segment.
15 . The system of claim 12 , wherein the one or more processors are further configured to execute the computer-executable instructions to:
generate a computer-generated voiceover based on the first text-based narrative, wherein the first customized media content includes the computer-generated voiceover.
16 . The system of claim 12 , wherein the first input is a text-based or voice query received from a user device.
17 . The system of claim 12 , wherein the first input is historical data associated with the first user.
18 . The system of claim 12 , wherein the one or more processors are further configured to execute the computer-executable instructions to:
train one or more computing models using a set of media content.