Responsible prompt recommendation
An embodiment generates, by analyzing prompts, a prompt template. Each prompt includes a text description of a content to be generated by a model. An embodiment classifies, using a first trained classification model, a variant portion of a prompt into a category in a set of categories. An embodiment selects, from a repository of prompt templates including the prompt template, a selected prompt template having a similarity above a threshold similarity to a first prompt. An embodiment classifies, using the selected prompt template, a variant portion of the first prompt into a first category in the set of categories. An embodiment adjusts, responsive to determining that the first category is designated as a harmful category, the variant portion of the first prompt. An embodiment causes, using the adjusted first prompt, the model to produce a first content.
1 . A computer-implemented method comprising:
generating, by executing a natural language processing algorithm analyzing a plurality of prompts, a prompt template, wherein each prompt in the plurality of prompts comprises a text description of a content to be generated by a large language model;
segmenting, using the prompt template, each prompt in the plurality of prompts into an invariant portion and a variant portion;
training, using prompt text data, a machine learning model to classify variant portions of prompts, the training generating a first trained classification model;
classifying, using the first trained classification model, a variant portion of a prompt in the plurality of prompts into a category in a set of categories, wherein each category of the set of categories comprises a level of harm and wherein each level corresponds to a different type of adjustment;
selecting, using the first trained classification model, from a repository of prompt templates including the prompt template, a selected prompt template, the selected prompt template having a similarity above a threshold similarity to a first prompt;
classifying, using the first trained classification model using the selected prompt template, a variant portion of the first prompt into a first category in the set of categories, the first category comprising a first level of harm;
automatically adjusting, responsive to determining by the first trained classification model that the first category is designated as a harmful category, the variant portion of the first prompt based on the first level of harm, the adjusting resulting in an adjusted first prompt; and
causing, using the adjusted first prompt, the large language model to produce a first content.
2 . The computer-implemented method of claim 1 , wherein a first category in the set of categories is designated as a responsible category.
3 . The computer-implemented method of claim 1 , wherein a second category in the set of categories is designated as a harmful category.
4 . The computer-implemented method of claim 1 , wherein the adjusting removes the variant portion of the first prompt.
5 . The computer-implemented method of claim 1 , wherein the adjusting replaces the variant portion of the first prompt with a replacement variant portion.
6 . The computer-implemented method of claim 1 , further comprising:
selecting, from the repository of prompt templates including the prompt template, a second selected prompt template, the second selected prompt template having a similarity above a threshold similarity to a second prompt;
classifying, using the second selected prompt template, a variant portion of the second prompt into a second category in the set of categories; and
identifying, to a user, responsive to determining that the second category is designated as a responsible category, the variant portion of the first prompt as a responsible portion.
7 . A computer program product comprising one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by a processor to cause the processor to perform operations comprising:
generating, by executing a natural language processing algorithm analyzing a plurality of prompts, a prompt template, wherein each prompt in the plurality of prompts comprises a text description of a content to be generated by a large language model;
segmenting, using the prompt template, each prompt in the plurality of prompts into an invariant portion and a variant portion;
training, using prompt text data, a machine learning model to classify variant portions of prompts, the training generating a first trained classification model;
classifying, using the first trained classification model, a variant portion of a prompt in the plurality of prompts into a category in a set of categories, wherein each category of the set of categories comprises a level of harm and wherein each level corresponds to a different type of adjustment;
selecting, using the first trained classification model, from a repository of prompt templates including the prompt template, a selected prompt template, the selected prompt template having a similarity above a threshold similarity to a first prompt;
classifying, using the first trained classification model using the selected prompt template, a variant portion of the first prompt into a first category in the set of categories, the first category comprising a first level of harm;
automatically adjusting, responsive to determining by the first trained classification model that the first category is designated as a harmful category, the variant portion of the first prompt based on the first level of harm, the adjusting resulting in an adjusted first prompt; and
causing, using the adjusted first prompt, the large language model to produce a first content.
8 . The computer program product of claim 7 , wherein the stored program instructions are stored in a computer readable storage device in a data processing system, and wherein the stored program instructions are transferred over a network from a remote data processing system.
9 . The computer program product of claim 7 , wherein the stored program instructions are stored in a computer readable storage device in a server data processing system, and wherein the stored program instructions are downloaded in response to a request over a network to a remote data processing system for use in a computer readable storage device associated with the remote data processing system, further comprising:
program instructions to meter use of the program instructions associated with the request; and
program instructions to generate an invoice based on the metered use.
10 . The computer program product of claim 7 , wherein a first category in the set of categories is designated as a responsible category.
11 . The computer program product of claim 7 , wherein a second category in the set of categories is designated as a harmful category.
12 . The computer program product of claim 7 , wherein the adjusting removes the variant portion of the first prompt.
13 . The computer program product of claim 7 , wherein the adjusting replaces the variant portion of the first prompt with a replacement variant portion.
14 . The computer program product of claim 7 , further comprising:
selecting, from the repository of prompt templates including the prompt template, a second selected prompt template, the second selected prompt template having a similarity above a threshold similarity to a second prompt;
classifying, using the second selected prompt template, a variant portion of the second prompt into a second category in the set of categories; and
identifying, to a user, responsive to determining that the second category is designated as a responsible category, the variant portion of the first prompt as a responsible portion.
15 . A computer system comprising a processor and one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by the processor to cause the processor to perform operations comprising:
generating, by executing a natural language processing algorithm analyzing a plurality of prompts, a prompt template, wherein each prompt in the plurality of prompts comprises a text description of a content to be generated by a large language model;
segmenting, using the prompt template, each prompt in the plurality of prompts into an invariant portion and a variant portion;
training, using prompt text data, a machine learning model to classify variant portions of prompts, the training generating a first trained classification model;
classifying, using the first trained classification model, a variant portion of a prompt in the plurality of prompts into a category in a set of categories, wherein each category of the set of categories comprises a level of harm and wherein each level corresponds to a different type of adjustment;
selecting, using the first trained classification model, from a repository of prompt templates including the prompt template, a selected prompt template, the selected prompt template having a similarity above a threshold similarity to a first prompt;
classifying, using the first trained classification model using the selected prompt template, a variant portion of the first prompt into a first category in the set of categories, the first category comprising a first level of harm;
automatically adjusting, responsive to determining by the first trained classification model that the first category is designated as a harmful category, the variant portion of the first prompt based on the first level of harm, the adjusting resulting in an adjusted first prompt; and
causing, using the adjusted first prompt, the large language model to produce a first content.
16 . The computer system of claim 15 , wherein a first category in the set of categories is designated as a responsible category.
17 . The computer system of claim 15 , wherein a second category in the set of categories is designated as a harmful category.
18 . The computer system of claim 15 , wherein the adjusting removes the variant portion of the first prompt.
19 . The computer system of claim 15 , wherein the adjusting replaces the variant portion of the first prompt with a replacement variant portion.
20 . The computer system of claim 15 , further comprising:
selecting, from the repository of prompt templates including the prompt template, a second selected prompt template, the second selected prompt template having a similarity above a threshold similarity to a second prompt;
classifying, using the second selected prompt template, a variant portion of the second prompt into a second category in the set of categories; and
identifying, to a user, responsive to determining that the second category is designated as a responsible category, the variant portion of the first prompt as a responsible portion.