Filtering for harmful generative artificial intelligence results
Techniques for filtering for harmful generative artificial intelligence (AI) results are described. An example of filtering includes receiving a request for an input phrase to be responded to by a generative AI model; comparing to the received input phrase to at least one known harmful input phrase to determine that the received input phrase is to be provided to the generative AI model; providing the received input phrase to the generative AI model; and generating a response by at least in part on an output of the generative AI model.
1 . A computer-implemented method comprising:
receiving a request for an input phrase to be responded to by a generative artificial intelligence (AI) model;
comparing the received input phrase to a known harmful input phrase that is most similar to the received input phrase and a paraphrase of a plurality of paraphrases associated with the known harmful input phrase in a multi-stage process to determine that the received input phrase is to be provided to the generative AI model, the comparing comprising:
generating a first set of similarity values between the received input phrase and the paraphrase; and
determining that the first set of similarity values is below a first similarity threshold;
providing the received input phrase to the generative AI model; and
generating a response based at least in part on an output of the generative AI model.
2 . The computer-implemented method of claim 1 , wherein comparing the received input phrase to the known harmful input phrase and the paraphrase associated with the known harmful input phrase in the multi-stage process to determine that the received input phrase is to be provided to the generative AI model further comprises:
generating a second set of similarity values between the received input phrase and at least a subset of stored known harmful input phrases including the known harmful input phrase;
determining that the second set of similarity values is below a second similarity threshold; and
determining that the second set of similarity values is above a third second similarity threshold.
3 . The computer-implemented method of claim 1 , wherein the received input phrase is in a format of one of text, audio, or video.
4 . A computer-implemented method comprising:
receiving a request for an input phrase to be responded to by a generative artificial intelligence (AI) model;
comparing the received input phrase to at least one paraphrase associated with a known harmful input phrase that is most similar to the received input phrase to determine that the received input phrase is to be provided to the generative AI model, the comparing comprising:
generating a first set of similarity values between the received input phrase and the at least one paraphrase associated with the known harmful input phrase that is most similar to the received input phrase; and
determining that the first set of similarity values is below a first similarity threshold;
providing the received input phrase to the generative AI model; and
generating a response based at least in part on an output of the generative AI model.
5 . The computer-implemented method of claim 4 , further comprising:
comparing the received input phrase to at least one known harmful input phrase to determine that the received input phrase is to be provided to the generative AI model by:
generating a second set of similarity values between the received input phrase and at least a subset of stored known harmful input phrases; and
determining that the second set of similarity values is below a second similarity threshold.
6 . The computer-implemented method of claim 5 , wherein comparing the received input phrase to the at least one known harmful input phrase to determine that the received input phrase is to be provided to the generative AI model further comprises:
determining that the second set of similarity values is above a third similarity threshold.
7 . The computer-implemented method of claim 6 , wherein at least one of the first, second, and third similarity thresholds are user configurable.
8 . The computer-implemented method of claim 6 , wherein comparing the received input phrase to the at least one known harmful input phrase to determine that the received input phrase is to be provided to the generative AI model further comprises:
determining that none of the second set of similarity values is above the second similarity threshold and is above the third similarity threshold;
generating a third set of similarity values between at least a proper subset of stored paraphrases associated with one or more of the stored known harmful input phrases; and
determining that none of the third set of similarity values indicates that the received input phrase is similar to the at least proper subset of stored paraphrases associated with the one or more of the stored known harmful input phrases.
9 . The computer-implemented method of claim 4 , further comprising:
receiving one or more known harmful input phrases including the known harmful input phrase;
generating one or more paraphrases for each known harmful input phrase of the one or more known harmful input phrases using one or more paraphrase generator models; and
storing the one or more known harmful input phrases and the generated one or more paraphrases for each known harmful input phrase of the one or more known harmful input phrases.
10 . The computer-implemented method of claim 9 , further comprising:
storing a denial response for one or more of the received one or more harmful input phrases.
11 . The computer-implemented method of claim 4 , further comprising:
verifying the output of the generative AI model is not a hallucination.
12 . The computer-implemented method of claim 4 , further comprising:
determining that the received input phrase is not out-of-scope, the determining comprising:
generating an embedding for the received input phrase to produce classification logits;
summing the logits in classification logits; and
comparing the sum to an out-of-scope threshold, wherein when the sum is greater than or equal to the out-of-scope threshold the received input phrase is not out-of-scope.
13 . The computer-implemented method of claim 4 , wherein the received input phrase is in a format of one of text, audio, or video.
14 . The computer-implemented method of claim 4 , wherein the generative AI model is one of a Transformer-based model, a generative adversarial network, or an auto-regressive convolutional neural networks (AR-CNN).
15 . The computer-implemented method of claim 4 , further comprising:
translating the received input phrase to a different language and using the translation as the received input phrase.
16 . A system comprising:
a first one or more computing devices to implement a storage service in a multi-tenant provider network; and
a second one or more computing devices to implement a generative artificial intelligence (AI) service in the multi-tenant provider network, the generative AI service including instructions that upon execution cause the generative AI service to: receive a request for an input phrase to be responded to by a generative AI model;
compare the received input phrase to at least one paraphrase associated with a known harmful input phrase that is most similar to the received input phrase, wherein the known harmful input phrase is to be stored by the storage service, the comparing to determine that the received input phrase is to be provided to the generative AI model, the comparing comprising:
generating a first set of similarity values between the received input phrase and the at least one paraphrase associated with the known harmful input phrase that is most similar to the received input phrase; and
determining that the first set of similarity values is below a first similarity threshold;
provide the received input phrase to the generative AI model; and
generate a response based at least in part on an output of the generative AI model.
17 . The system of claim 16 , wherein the generative AI service is further to compare the received input phrase to at least one known harmful input phrase.
18 . The system of claim 16 , wherein the generative AI service is further to verify the output of the generative AI model is not a hallucination.
19 . The system of claim 16 , wherein the generative AI model is one of a Transformer-based model, a generative adversarial network, or an auto-regressive convolutional neural networks (AR-CNN).
20 . The system of claim 16 , wherein the received input phrase is in a format of one of text, audio, or video.