IP Library Granted Patent US 12671882
Granted Patent B1
US 12671882 · App. 18/489,251 · Granted Jun 30, 2026

Systems and methods for generating customized media content

Inventors: Vladislav Isaev (West Vancouver, CA); Yash Chaturvedi (Issaquah, WA); Steven James Cox (Mill Creek, WA); Jacobus Hendrik du Preez (Liberty Hill, TX); Yongjun Wu (Bellevue, WA)
Assignee: Amazon Technologies, Inc.
H04N21/8549H04N21/234H04N21/4667H04N21/482H04N21/8106
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12671882
App. No.
18/489,251
Granted
Jun 30, 2026
Kind
B1
Abstract

Systems and methods for generating customized media content are provided (for example, trailers or recaps for a movie or television show). For example, the system may generate a trailer to a television show that is tailored to a specific user, such as a trailer that includes action-specific scenes for a user who often views content in an action-based genre. The system may receive as an input a text or voice-based query from the user (or the system may be automated and may receive user historical data as an input). Based on the input, one or more computing models may generate a text-based narrative to be used with the customized content. Once the narrative is generated, the one or more models may then identify specific video frames to include in the customized content. The video frames may then be stitched together and the customized content may be generated using the stitched video frames and the narrative.

Claims (61)

1 . A method comprising:

training one or more machine learning models using a set of media content stored in memory, the set of media content including at least: one or more movies and one or more televisions shows;

generating, using a first machine learning model, a plurality of video segments of the set of media content;

receiving, using one or more processors, a first input associated with a first user, the first input being a natural language input that is indicative of first customized media content to be generated based on at least a first movie or first television show of the set of media content,

wherein the first input includes at least one of: a text-based or voice query received from a user device and historical data associated with the first user,

and wherein the first customized media content includes a custom trailer or recap for the first movie or first television show;

generating, by a second machine learning model and based on the first input, a first text-based narrative associated with the first movie or first television show;

receiving, by a third machine learning model, the first text-based narrative, wherein the first text-based narrative comprises at least some unique text that differs from text included in an existing text-based narrative of the media content and is based on the natural language input instead of the existing text-based narrative of the first movie or first television show;

identifying, by the third machine learning model and based on the first text-based narrative, one or more first video frames of the first movie or first television show;

identifying, by the third machine learning model, one or more video segments of the plurality of video segments that includes the one or more first video frames; and

generating, by the one or more machine learning models, the first customized media content using the first text-based narrative and the one or more video segments, wherein the first machine learning model, second machine learning model, and third machine learning model are different models.

2 . The method of claim 1 , further comprising:

receiving a second input associated with a second user, the second input being indicative of second customized media content to be generated based on the first movie or first television show;

generating, by the second machine learning model and based on the second input, a second text-based narrative;

identifying, by the third machine learning model and based on the second text-based narrative, one or more second video frames of the set of media content; and

generating, by the one or more machine learning models, the second customized media content using the second text-based narrative and the one or more second video frames, wherein the second customized media content is different than the first customized media content, and wherein the second customized media content is tailored to the second user.

3 . The method of claim 1 , further comprising:

generating a computer-generated voiceover based on the first text-based narrative, wherein the first customized media content includes the computer-generated voiceover.

4 . The method of claim 1 , wherein the machine learning models comprise at least one of: a large language model, a large multimodal visual model, and a shot detection model.

5 . A method comprising:

receiving, using one or more processors, a first input associated with a first user, the first input being a natural language input that is indicative of first customized media content to be generated for first media content;

generating, by a first machine learning model and based on the first input, a first text-based narrative for the first customized media content, wherein the first text-based narrative comprises at least some unique text that differs from text included in an existing text-based narrative of the first media content and is based on the natural language input instead of the existing text-based narrative of the first media content;

identifying, by a second machine learning model and based on the first text-based narrative, one or more first video frames of the first media content; and

generating the first customized media content using the first text-based narrative and the one or more first video frames, wherein the first machine learning model and second machine learning model are different models.

6 . The method of claim 5 , further comprising:

receiving, using the one or more processors, a second input associated with a second user, the second input being indicative of second customized media content to be generated based on the first media content;

generating, by the first machine learning model and based on the second input, a second text-based narrative;

identifying, by the second machine learning model and based on the second text-based narrative, one or more second video frames of the first media content; and

generating the second customized media content using the second text-based narrative and the one or more second video frames, wherein the second customized media content is different than the first customized media content.

7 . The method of claim 5 , further comprising:

generating a first video segment and a second video segment for first media content of the first media content; and

determining that a first video frame of the one or more first video frames is included within the first video segment,

wherein the first customized media content includes the first video segment instead of the second video segment based on the determination that the first video frame is included within the first video segment.

8 . The method of claim 5 , further comprising:

generating, a computer-generated voiceover based on the first text-based narrative, wherein the first customized media content includes the computer-generated voiceover.

9 . The method of claim 5 , wherein the first input is a text-based or voice query received from a user device.

10 . The method of claim 5 , wherein the first input is historical data associated with the first user.

11 . The method of claim 5 , further comprising:

training one or more computing models using a set of media content.

12 . A system comprising:

memory that stores computer-executable instructions; and

one or more processors configured to access the memory and execute the computer-executable instructions to:

receive, using one or more processors, a first input associated with a first user, the first input being a natural language input that is indicative of first customized media content to be generated for first media content;

generate, by a first machine learning model and based on the first input, a first text-based narrative for the first customized media content, wherein the first text-based narrative comprises at least some unique text that differs from text included in an existing text-based narrative of the first media content and is based on the natural language input instead of the existing text-based narrative of the first media content;

identify, by a second machine learning model and based on the first text-based narrative, one or more first video frames of the first media content; and

generate the first customized media content using the first text-based narrative and the one or more first video frames, wherein the first machine learning model and second machine learning model are different models.

13 . The system of claim 12 , wherein the one or more processors are further configured to execute the computer-executable instructions to:

receive a second input associated with a second user, the second input being indicative of second customized media content to be generated based on the first media content;

generate, by the first machine learning model and based on the second input, a second text-based narrative;

identify, by the second machine learning model and based on the second text-based narrative, one or more second video frames of the first media content; and

generate second customized media content using the second text-based narrative and the one or more second video frames, wherein the second customized media content is different than the first customized media content.

14 . The system of claim 12 , wherein the one or more processors are further configured to execute the computer-executable instructions to:

generate a first video segment and a second video segment for first media content of the first media content; and

determine that a first video frame of the one or more first video frames is included within the first video segment,

wherein the first customized media content includes the first video segment instead of the second video segment based on the determination that the first video frame is included within the first video segment.

15 . The system of claim 12 , wherein the one or more processors are further configured to execute the computer-executable instructions to:

generate a computer-generated voiceover based on the first text-based narrative, wherein the first customized media content includes the computer-generated voiceover.

16 . The system of claim 12 , wherein the first input is a text-based or voice query received from a user device.

17 . The system of claim 12 , wherein the first input is historical data associated with the first user.

18 . The system of claim 12 , wherein the one or more processors are further configured to execute the computer-executable instructions to:

train one or more computing models using a set of media content.