Real-time identification of fact hallucinations in artificial intelligence (AI)
An AI query response processed by an artificial intelligence (AI) module using a neural model is received as an input. The query response is compared to one or more tokenized facts of the compressed blocks at a data level. Comparing includes computing a similarity score of a vector derived from the tokenized query response to one or more vectors derived from the one or more tokenized facts. Responsive to a verification result failing to meet a similarity score threshold, it can be determined that a hallucination exists in the AI query response and a policy-based action, such as blocking the AI query response, can be taken. Responsive to the verification result meeting the similarity threshold, the AI query response can be allowed to proceed.
1 . A method in an artificial intelligence (AI) validation server, for real-time identification of fact hallucinations in query results produced by AI, the method comprising:
receiving, in real-time, an AI query response processed by an AI module using a neural model as an input, wherein the AI query response involves at least two facts;
performing real-time verification of the query response by locating and retrieving compressed blocks from a token database having word-level similarity to the AI query response, wherein word-level similarity is computed from a weighted sum of Euclidian vector distances, by tokenizing data of the AI query response, wherein tokenizing comprises retrieving Kaggle tokenIDs associated with one or more words of the query response;
comparing the AI query response to one or more tokenized facts of the compressed blocks at a data level, wherein comparing includes computing a similarity score of a vector derived from the tokenized query response to one or more vectors derived from the one or more tokenized facts;
responsive to a verification result failing to meet a similarity score threshold, determining that a hallucination exists in the AI and taking a policy-based action on the AI query response; and
responsive to the verification result meeting the similarity threshold, allowing the AI query response to proceed.
2 . The method of claim 1 , wherein the tokenizing one or more facts comprises Kaggle tokenIDs associated with one or more words of the one or more facts.
3 . The method of claim 1 , wherein comparing comprises performing an AND operation on the tokenized query response and the one or more tokenized facts.
4 . The method of claim 1 , wherein the policy-based action comprises blocking the query response.
5 . The method of claim 1 , wherein the policy-based action comprises logging the verification result.
6 . The method of claim 1 , further comprising:
logging the fact hallucination.
7 . The method of claim 1 , further comprising: tracking performance of the AI module.
8 . The method of claim 1 , wherein the AI module producing the AI query response is located remotely from the AI validation server.
9 . The method of claim 1 , wherein a plurality of AI query responses are received from a plurality of different AI modules.
10 . The method of claim 1 , wherein the AI validation server and the AI module are both communicatively coupled to a data communication network.
11 . The method of claim 1 , wherein the AI validation server and the AI module are embedded in a common physical device.
12 . The method of claim 1 , further comprising:
receiving an AI query from a user over a data communication network, and allowing the AI query response to be transmitted to the user after successful validation.
13 . A non-transitory computer-readable media in an artificial intelligence (AI) validation server, implemented at least partially in hardware, when executed by a processor, for real-time identification of fact hallucinations in query results produced by AI, the method comprising the steps of:
receiving, in real-time, an AI query response processed by an AI module using a neural model as an input, wherein the AI query response involves at least two facts;
performing real-time verification of the AI query response by locating and retrieving compressed blocks from a token database having word-level similarity to the AI query response, by tokenizing data of the AI query response, wherein word-level similarity is computed from a weighted sum of Euclidian vector distances, by tokenizing data of the AI query response, wherein tokenizing comprises retrieving Kaggle tokenIDs associated with one or more words of the query response;
comparing the AI query response to one or more tokenized facts of the compressed blocks at a data level, wherein comparing includes computing a similarity score of a vector derived from the tokenized query response to one or more vectors derived from the one or more tokenized facts;
responsive to a verification result failing to meet a similarity score threshold, determining that a hallucination exists in the AI query response and taking a policy-based action on the AI query response; and
responsive to the verification result meeting the similarity threshold, allowing the AI query response to proceed.
14 . An artificial intelligence (AI) validation server, for real-time identification of hallucinations in query results produced by AI, the AI validation server:
a processor;
a network gateway communicatively coupled to the processor and to the data communication network; and
a memory communicatively coupled to the processor and storing modules, comprising:
an API module configured to receive, in real-time, a query response processed by an AI module using a neural model as an input, wherein the AI query response involves at least two facts;
a sentence retrieval module configured to perform real-time verification of the AI query response by locating and retrieving compressed blocks from a token database having word-level similarity to the AI query response, by tokenizing data of the AI query response, wherein word-level similarity is computed from a weighted sum of Euclidian vector distances, by tokenizing data of the AI query response, wherein tokenizing comprises retrieving Kaggle tokenIDs associated with one or more words of the query response;
a sentence score module configured to compare the AI query response to one or more tokenized facts of the compressed blocks at a data level, wherein comparing includes computing a similarity score of a vector derived from the tokenized query response to one or more vectors derived from the one or more tokenized facts,
wherein the sentence scoring module is configured to, responsive to a verification result failing to meet a similarity score threshold, determine that a hallucination exists in the AI query response; and
a veracity policy module configured to, responsive to the verification result failing, take a policy-based action on the AI query response,
wherein responsive to the verification result meeting the similarity threshold, allow the AI query response to proceed.