Systems and methods for modification of machine learning model-generated text and images based on user queries and profiles
Systems and methods for generating user-specific textual and image-based outputs corresponding to existing items, in response to user queries, are disclosed herein. For example, the system may receive a query that includes a textual description. The system may retrieve a user profile for a user associated with the query. Based on the query, the system may obtain a description of an item. Based on the query, the user profile, and the description, the system may generate an output and an image using a machine learning model. Based on the output and the image, the system may generate a graphical representation of the item. The system may receive a selection of the graphical representation of the item. Based on the selection of the graphical representation, the system may enable access to the first item.
1 . A system for reducing image-based hallucination generation in machine-learning models for generating user-specific images on features identified within user queries, the system comprising:
one or more processors; and
one or more non-transitory, computer-readable media storing instructions that, when executed by the one or more processors, cause operations comprising:
obtaining, based on a query, an item description of a first item;
providing previous queries of a user profile, the query, and the first item description to a one or more machine learning n artificial intelligence-models to cause the one or more machine learning models to generate a first synthetic image and a text output that comprises a plurality of features corresponding to the item description and the user profile, wherein characteristics of the first synthetic image correspond to the plurality of features;
determining that one or more characteristics of the first synthetic image do not contain features that are relevant to the query or the first item by detecting the one or more characteristics of the first synthetic image do not satisfy a set of matching criteria associated with the first item;
generating a second synthetic image based on the one or more characteristics of the first synthetic image not satisfying the set of matching criteria;
generating a graphical representation of the first item comprising the second synthetic image and a representation of the text output;
in response to receiving a selection of the graphical representation of the first item, updating the one or more machine learning models to reduce hallucination generation likelihood by retraining the one or more machine learning models based on the text output and the second synthetic image in lieu of the first synthetic image; and
updating the user profile by generating user metadata based on additional outputs of the retrained machine learning model associated with a later query.
2 . A method for reducing image-based hallucination generation in machine-learning models, the method comprising:
obtaining by one or more processors, based on a query, an item description of a first item;
providing, by the one or more processors, previous queries of a user profile, the query, and the item description to one or more machine learning models to cause the one or more machine learning models to generate a first synthetic image and a text output that comprises an indication of a plurality of features corresponding to the item description and the user profile, wherein characteristics of the first synthetic image correspond to the plurality of features;
determining, by the one or more processors, that one or more characteristics of the first synthetic image do not satisfy a set of matching criteria associated with the first item;
generating, by the one or more processors, a second synthetic image based on the one or more characteristics of the first synthetic image not satisfying the set of matching criteria;
generating, by the one or more processors, a graphical representation of the first item comprising the second synthetic image and a representation of the text output;
receiving, by the one or more processors, a selection of the graphical representation of the first item;
in response to receiving the selection of the graphical representation of the first item, updating, by the one or more processors, the one or more machine learning models to reduce hallucination generation likelihood by retraining the one or more machine learning models based on the text output and the second synthetic image; and
updating, by the one or more processors, the user profile by generating user metadata based on additional outputs of the one or more machine learning models associated with a later query.
3 . The method of claim 2 , wherein obtaining the first-item description of the first item comprises:
determining, based on the query, a plurality of requested characteristics describing a requested item;
obtaining a plurality of characteristic sets, wherein each characteristic set of the plurality of characteristic sets comprises a corresponding plurality of item characteristics associated with a corresponding item;
comparing each characteristic set of the plurality of characteristic sets with the plurality of requested characteristics;
in response to comparing each characteristic set of the plurality of characteristic sets with the plurality of requested characteristics, determining the first item; and
retrieving, from an item database, the item description of the first item.
4 . The method of claim 3 , wherein, in response to comparing each characteristic set of the plurality of characteristic sets with the plurality of requested characteristics, determining the first item comprises:
generating a plurality of similarity metrics, wherein each similarity metric of the plurality of similarity metrics indicates a similarity between a requested characteristic of the plurality of requested characteristics and a corresponding characteristic set, for a corresponding item, of the plurality of characteristic sets;
determining that a first similarity metric of the plurality of similarity metrics meets a threshold similarity metric; and
determining the first item, wherein the first item corresponds to the first similarity metric.
5 . The method of claim 4 , wherein generating the graphical representation of the first item comprising the second synthetic image and a representation of the text output further comprises:
determining a subset of similarity metrics of the plurality of similarity metrics, wherein each similarity metric of the subset of similarity metrics meets the threshold similarity metric; and
generating a ranked list of a plurality of graphical representations of items comprising the graphical representation of the first item, wherein the plurality of graphical representations are ranked according to the plurality of similarity metrics.
6 . The method of claim 2 , wherein generating the second synthetic image based on the one or more characteristics of the first synthetic image not satisfying the set of matching criteria in response to providing the query and the item description of the first item to the one or more machine learning models, wherein the second synthetic image distinct from the first synthetic image.
7 . The method of claim 2 , wherein generating the text output comprises:
generating, via the one or more machine learning models, a plurality of semantic tokens corresponding to the item description and the query, wherein the plurality of semantic tokens includes words, phrases, or sentences associated with the item description; and
generating a first feature of the plurality of features to comprise a first semantic token of the plurality of semantic tokens.
8 . The method of claim 2 , further comprising:
generating, based on the selection of the graphical representation of the first item, training data comprising the first synthetic image, the representation of the text output, and the user profile; and
providing the training data to the one or more machine learning models to train the one or more machine learning models to generate images and text outputs based on input user profiles.
9 . The method of claim 8 , further comprising:
receiving a second query from a user;
obtaining, based on the second query, a second item description of a second item; and
providing the user profile, the second query, and the second item description to the one or more machine learning models to cause the one or more machine learning models to generate a second output and a second image.
10 . The method of claim 9 , further comprising:
generating, for display on a user interface, the graphical representation of the first item and a second graphical representation of the second item;
receiving a selection of the second graphical representation of the second item; and
based on receiving the selection of the second graphical representation of the second item, enabling access to the second item.
11 . The method of claim 2 , further comprising:
generating a second text output based on the one or more characteristics of the first synthetic image not satisfying the set of matching criteria, wherein the second text output is distinct from the text output;
generating a second graphical representation of the first item, wherein the second graphical representation of the first item comprises the second synthetic image and a representation of the second text output;
receiving a selection of the second graphical representation of the first item; and
in response to receiving the selection of the second graphical representation of the first item, enabling access to the first item.
12 . The method of claim 2 , further comprising:
obtaining, based on the query, a second item description of a second item;
providing previous queries of a user profile, the query, and the second item description to the one or more machine learning models to cause the one or more machine learning models to generate third synthetic image and a second text output;
determining that the one or more characteristics of the third synthetic image do not satisfy a set of matching criteria associated with the second item; and
generating a fourth synthetic image based on the one or more characteristics of the third synthetic image not satisfying the set of matching criteria.
13 . The method of claim 12 , further comprising:
generating a second graphical representation of the second item, wherein the second graphical representation of the second item comprises the third synthetic image and a representation of the second text output;
receiving a selection of the second graphical representation of the second item; and
in response to receiving the selection of the second graphical representation of the second item, updating the one or more machine learning models to reduce hallucination generation likelihood by retraining the one or more machine learning models based on the second text output and the third synthetic image in lieu of the third synthetic image; and
updating the user profile by generating user metadata based on additional outputs of the one or more machine learning models associated with a later query.
14 . One or more non-transitory, computer-readable media storing instructions that, when executed by one or more processors, cause operations comprising:
obtaining, based on a query, an item description of a first item;
providing previous queries of a user profile, the query, and the item description to one or more machine learning models to cause the one or more machine learning models to generate a first synthetic image and a text output that comprises an indication of a plurality of features corresponding to the item description, wherein characteristics of the first synthetic image correspond to the plurality of features;
determining that one or more characteristics of the first synthetic image do not satisfy a set of matching criteria associated with the first item;
generating a second synthetic image based on the one or more characteristics of the first synthetic image not satisfying the set of matching criteria;
generating a graphical representation of the first item comprising the second synthetic image and a representation of the text output;
receiving a selection of the graphical representation of the first item; and
in response to receiving the selection of the graphical representation of the first item, updating the one or more machine learning models by retraining the one or more machine learning models based on the text output and the second synthetic image.
15 . The one or more non-transitory, computer-readable media of claim 14 , wherein obtaining the item description comprises:
determining, based on the query, a plurality of requested characteristics describing a requested item;
obtaining a plurality of characteristic sets, wherein each characteristic set of the plurality of characteristic sets comprises a corresponding plurality of item characteristics associated with a corresponding item;
comparing each characteristic set of the plurality of characteristic sets with the plurality of requested characteristics;
in response to comparing each characteristic set of the plurality of characteristic sets with the plurality of requested characteristics, determining the first item; and
retrieving, from an item database, the item description of the first item.
16 . The one or more non-transitory, computer-readable media of claim 15 , wherein, in response to comparing each characteristic set of the plurality of characteristic sets with the plurality of requested characteristics, determining the first item comprises:
generating a plurality of similarity metrics, wherein each similarity metric of the plurality of similarity metrics indicates a similarity between a requested characteristic of the plurality of requested characteristics and a corresponding characteristic set, for a corresponding item, of the plurality of characteristic sets;
determining that a first similarity metric of the plurality of similarity metrics meets a threshold similarity metric; and
determining the first item, wherein the first item corresponds to the first similarity metric.
17 . The one or more non-transitory, computer-readable media of claim 16 , the operations further comprising:
determining a subset of similarity metrics of the plurality of similarity metrics, wherein each similarity metric of the subset of similarity metrics meets the threshold similarity metric;
generating a ranked list of a plurality of graphical indications of items comprising a graphical indication, wherein the plurality of graphical indications of items are ranked according to the plurality of similarity metrics; and
presenting the graphical indication.
18 . The one or more non-transitory, computer-readable media of claim 14 , further comprising:
obtaining a first characteristic set corresponding to a first item, wherein the first characteristic set includes characteristics describing the first item;
determining whether the characteristics of the first synthetic image correspond to the first characteristic set; and
based on determining that a first characteristic of the first synthetic image does not correspond to the first characteristic set, providing, the query, and the first item description to the one or more machine learning models to generate a second synthetic image distinct from the first synthetic image.