Systems and methods for identifying a design template matching a media item
A method for automatically generating one or more digital designs is disclosed. The method includes identifying an input media item; processing the input media item to generate an input media item descriptor; and identifying a first target media item from a set of target media items. Each target media item in the set of target media items is associated with a target media item descriptor and a candidate design template, and the first target media item is identified based on a similarity between the input media item descriptor and the target media item descriptor of the first target media item. The method further includes generating a new digital design. The new digital design being based on the candidate design template associated with the first target media item, and generated to replace the first target media item with the input media item.
1 . A method for automatically generating one or more digital designs, the method including:
identifying an input media item;
processing the input media item to generate an input media item descriptor;
identifying a first target media item from a set of target media items, wherein each target media items in the set of target media items is associated with a target media item descriptor and a candidate design template, and wherein the first target media item is identified based on a similarity between the input media item descriptor and the target media item descriptor of the first target media item;
selecting the candidate design template associated with the first target media item; and
generating a new digital design, wherein:
the new digital design is based on the selected candidate design template; and
the new digital design is generated to replace the first target media item in the selected candidate design template with the input media item.
2 . The method of claim 1 , wherein the set of target media items and the input media item is at least one of an image, a video, or an audio file.
3 . The method of claim 2 , wherein generating the input media item descriptor comprising:
analyzing the input media item using a machine learning model, and
generating a vector embedding representing the input media item in a vector space or generating a natural language caption representing content of the media item.
4 . The method of claim 3 , wherein the machine learning model is a contrastive language-image pretraining (CLIP) model.
5 . The method of claim 3 , wherein identifying the first target media item from the set of target media items comprises:
determining a distance between the vector embedding of the input media item and the vector embeddings of the set of target media items; and
identifying the first target media item from the set of target media items as the target media item that has a vector embedding nearest to the vector embedding of the input media item.
6 . The method of claim 1 , wherein the set of target media items is selected from a superset of media items based on a frequency with which a target media item is replaced from a design template.
7 . A computer-implemented method for identifying one or more design templates matching an input media item, including:
receiving the input media item, the input media item selected by a user;
processing the input media item to generate an input media item descriptor;
identifying one or more target media items from a set of target media items, wherein each target media item in the set of target media items is associated with a target media item descriptor and is an existing media item of a candidate design template, and wherein the one or more target media items are identified based on a similarity between the input media item descriptor and the target media item descriptors of the one or more target media items;
identifying candidate design templates associated with the one or more identified target media items; and
causing display of the identified candidate design templates on a display of a user device that selected the input media item.
8 . The computer-implemented method of claim 7 , further comprising generating one or more new digital designs, wherein each new digital design is based on a candidate design template of the candidate design templates associated with the one or more target media items, and wherein each new digital design is generated to replace the one or more target media items of the candidate design templates with the input media item.
9 . The computer-implemented method of claim 7 , wherein the set of target media items and the input media item is at least one of an image, a video, or an audio file.
10 . The computer-implemented method of claim 7 , further comprises generating media item descriptors for each media item in the set of target media items.
11 . The computer-implemented method of claim 10 , wherein generating the input media item descriptor or generating the media item descriptors for the set of media items comprising analyzing the input media item or each media item in the set of media items using a machine learning model, and the media item descriptors includes a vector embedding representing the input media item or each media item in the set of media items in a vector space.
12 . The computer-implemented method of claim 11 , wherein identifying the one or more target media items from the set of target media items comprises:
determining a distance between the vector embedding of the input media item and the vector embeddings of the set of target media items; and
identifying the one or more target media items from the set of target media items as the one or more target media items that have a vector embedding within a threshold distance of the vector embedding of the input media item.
13 . The computer-implemented method of claim 11 , wherein identifying the one or more target media items from the set of target media items comprises:
determining a distance between the vector embedding of the input media item and the vector embeddings of the set of target media items; and
identifying the one or more target media items from the set of target media items by selecting a maximum or minimum number of target media items that have a vector embeddings near the vector embedding of the input media item.
14 . The computer-implemented method of claim 7 , wherein the set of target media items is selected from a superset of media items based on a frequency with which a target media item is replaced from a design template.
15 . The computer-implemented method of claim 10 , further comprising generating an index based on the media item descriptors of the set of media items.
16 . The computer-implemented method of claim 15 , wherein identifying the one or more target media items from the set of target media items comprises performing a search in the index for the one or more target media items that have descriptors that are the most similar to the descriptor of the input media item.
17 . A non-transitory computer readable medium comprising instructions, which when executed by a processing unit of a computer processing system, cause the computer processing system to:
identify an input media item;
process the input media item to generate an input media item descriptor;
identify a first target media item from a set of target media items, wherein each target media items in the set of target media items is associated with a target media item descriptor and a candidate design template, and wherein the first target media item is identified based on a similarity between the input media item descriptor and the target media item descriptor of the first target media item;
select the candidate design template associated with the first target media item; and
generate a new digital design, wherein:
the new digital design is based on the selected candidate design template; and
the new digital design is generated to replace the first target media item in the selected candidate design template with the input media item.
18 . The non-transitory computer readable medium of claim 17 , wherein generating the input media item descriptor comprising analyzing the input media item using a machine learning model, and the input media item descriptor includes a vector embedding representing the input media item in a vector space or a natural language caption representing content of the input media item.
19 . The non-transitory computer readable medium of claim 18 , wherein identifying the first target media item from the set of target media items comprises:
determining a distance between the vector embedding of the input media item and the vector embeddings of the set of target media items and identifying the first target media item from the set of target media items as the media item that has a vector embedding nearest to the vector embedding of the input media item.
20 . The non-transitory computer readable medium of claim 17 , wherein the set of target media items is selected from a superset of media items based on a frequency with which a target media item is replaced from a design template.