IP Library › Granted Patent US 12,505,311
Granted Patent B2
US 12,505,311 · App. 18/452,298 · Granted Dec 23, 2025

Hallucination detection and handling for a large language model based domain-specific conversation system

Inventors: Zhengyu Zhou (Fremont, CA); Arsalan Gundroo (San Francisco, CA)
Assignee: Robert Bosch GmbH
G06F40/40G06F16/3329G06F40/205
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,311
App. No.
18/452,298
Filed
Aug 18, 2023
Granted
Dec 23, 2025
Kind
B2
Art Unit
2688
USPC
704/9
Abstract

Systems and methods are described herein for detecting and reducing hallucinations and improving the reliability of an LLM-based QA system. Particularly, the disclosure provides a set of hallucination detection and handling approaches that address different types of hallucinations of the QA system. The hallucination detection approaches described herein include comparing a natural language answer with context information by way of sentence similarity estimation and keyword matching. The hallucination handling approaches described herein include removing hallucinated sentences from the natural language answer or regenerating the natural language answer using better context information, depending on a level of hallucination detected in the natural language answer. A hybrid framework is also provided that systematically combines the hallucination detection approaches and hallucination handling approaches into one system to achieve an optimal hallucination-reduction performance for the QA system.

Claims (60)

1 . A method of answering natural language questions, the method comprising:

receiving, with a processor, a natural language question;

retrieving, with the processor, natural language context information from an information source using a first information retrieval technique based on the natural language question;

generating, with the processor, a natural language answer using a language model based on the natural language question and the natural language context information; and

detecting, with the processor, hallucinated content in the natural language answer based on the natural language context information, the detecting the hallucinated including (i) comparing the natural language answer with the natural language context information, (ii) determining similarities between sentences in the natural language answer and sentences in the natural language context information, and (iii) labelling at least one sentence in the natural language answer as the hallucinated content in response to the at least one sentence having less than a threshold similarity with every sentence in the natural language context information.

2 . The method according to claim 1 , the determining the similarities further comprising:

determining a sentence embedding for each sentence in the natural language answer;

determining a sentence embedding for each sentence in the natural language context information; and

determining, for each respective sentence in the natural language answer, a respective embedding similarity between the respective sentence in the natural language answer and each respective sentence in the natural language context information, based on the sentence embedding for the respective sentence in the natural language answer and the sentence embedding for the respective sentence in the natural language context information.

3 . The method according to claim 2 , the determining the respective embedding similarity further comprising:

determining a respective cosine similarity between the sentence embedding for the respective sentence in the natural language answer and the sentence embedding for the respective sentence in the natural language context information.

4 . The method according to claim 2 , the determining the similarities further comprising:

determining, for each respective sentence in the natural language answer, a respective word pattern overlap rate between the respective sentence in the natural language answer and each respective sentence in the natural language context information.

5 . The method according to claim 4 , the determining the respective word pattern overlap rate further comprising:

determining a respective mapping between the respective sentence in the natural language answer and the respective in the natural language context information; and

determining the respective word pattern overlap rate between the respective sentence in the natural language answer and the respective in the natural language context information based on the respective mapping.

6 . The method according to claim 5 , the determining the respective word pattern overlap rate further comprising:

determining a respective overlap length as a sum of overlapping words between the respective sentence in the natural language answer and the respective in the natural language context information; and

determining the respective word pattern overlap rate between the respective sentence in the natural language answer and the respective in the natural language context information by dividing the respective overlap length by a number of words in a shorter sentence of the respective sentence in the natural language answer and the respective in the natural language context information.

7 . The method according to claim 6 , the determining the respective overlap length further comprising:

determining sum of overlapping words between the respective sentence in the natural language answer and the respective in the natural language context information as only including words in a sequence of overlapping words having at least a predetermined minimum number of words.

8 . The method according to claim 4 , the labelling the at least one sentence in the natural language answer as the hallucinated content further comprising:

labelling the at least one sentence in the natural language answer as the hallucinated content in response to:

(i) the respective embedding similarity between the at least one sentence in the natural language answer and each sentence in the natural language context information being less than a first predetermined threshold; and

(ii) the respective word pattern overlap rate between the at least one sentence in the natural language answer and each sentence in the natural language context information being less than a second predetermined threshold.

9 . The method according to claim 1 , the detecting the hallucinated content further comprising:

identifying keywords in the natural language answer;

determining whether the keywords are also present in the natural language context information; and

labelling the natural language answer as including the hallucinated content in response an amount of the keywords in the natural language answer not being present in the natural language context information exceeding a third predetermined threshold.

10 . The method according to claim 1 further comprising:

generating, with the processor, a modified natural language answer that reduces the hallucinated content in the natural language answer.

11 . The method according to claim 10 , the generating the modified the natural language answer further comprising:

generating the modified natural language answer by deleting the hallucinated content from the natural language answer.

12 . The method according to claim 11 further comprising:

determining, with the processor, a hallucination level that quantifies an amount of the hallucinated content in the natural language answer,

wherein the deleting the hallucinated content from the natural language answer is performed in response to the hallucination level being less than a predetermined threshold.

13 . The method according to claim 10 , the generating the modified the natural language answer further comprising:

retrieving modified natural language context information using a second information retrieval technique based on the natural language question, the second information retrieval technique being different than the first information retrieval technique; and

generating the modified natural language answer using the language model based on the natural language question and the modified natural language context information.

14 . The method according to claim 13 further comprising:

determining, with the processor, a hallucination level that classifies an amount of the hallucinated content in the natural language answer,

wherein the retrieving modified natural language context information and the generating the modified natural language answer based on the modified natural language context information are performed in response to the hallucination level being greater than a predetermined threshold.

15 . The method according to claim 13 , wherein the retrieving modified natural language context information and the generating the modified natural language answer based on the modified natural language context information are repeated using different information retrieval techniques in each case, until the modified natural language answer includes less than a threshold amount of hallucinated content.

16 . The method according to claim 13 , wherein the retrieving modified natural language context information and the generating the modified natural language answer based on the modified natural language context information are repeated using different information retrieval techniques in each case, until no further information retrieval techniques result a reduced amount of the hallucinated content in the modified natural language answer.

17 . The method according to claim 10 further comprising:

outputting, with an output device, the modified natural language answer to a user.

18 . The method according to claim 17 further comprising:

identifying, with the processor, sentences in the natural language context information that provide support for sentences in the modified natural language answer; and

outputting, with the output device, indications of the sentences in the natural language context information that provide support for sentences in the modified natural language answer.

19 . A method of answering natural language questions, the method comprising:

receiving, with a processor, a natural language question;

retrieving, with the processor, natural language context information from an information source using a first information retrieval technique based on the natural language question;

generating, with the processor, a natural language answer using a language model based on the natural language question and the natural language context information; and

detecting, with the processor, hallucinated content in the natural language answer based on the natural language context information, the detecting the hallucinated including (i) comparing the natural language answer with the natural language context information, (ii) identifying keywords in the natural language answer, (iii) determining whether the keywords are also present in the natural language context information, and (iv) labelling the natural language answer as including the hallucinated content in response an amount of the keywords in the natural language answer not being present in the natural language context information exceeding a third predetermined threshold.

20 . A method of answering natural language questions, the method comprising:

receiving, with a processor, a natural language question;

retrieving, with the processor, natural language context information from an information source using a first information retrieval technique based on the natural language question;

generating, with the processor, a natural language answer using a language model based on the natural language question and the natural language context information;

detecting, with the processor, hallucinated content in the natural language answer based on the natural language context information; and

generating, with the processor, a modified natural language answer that reduces the hallucinated content in the natural language answer, the generating the modified the natural language answer further including (i) retrieving modified natural language context information using a second information retrieval technique based on the natural language question, the second information retrieval technique being different than the first information retrieval technique, and (ii) generating the modified natural language answer using the language model based on the natural language question and the modified natural language context information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 25, 2023
From: ZHOU, ZHENGYU; GUNDROO, ARSALAN
To: ROBERT BOSCH GMBH
Reel/Frame 064703/0617 →
Continuity (1)
Related Publication 20250061286A1 · Feb 20, 2025
References Cited (26)
US 4914590A · Loatman · 1990 [cited by examiner]
US 11508372B1 · Schwartz · 2022 [cited by examiner]
US 11605376B1 · Hoover · 2023 [cited by examiner]
US 20080140387A1 · Linker · 2008 [cited by examiner]
US 20130297293A1 · Di Cristo · 2013 [cited by examiner]
US 20180189267A1 · Takiel · 2018 [cited by examiner]
US 20190327330A1 · Natarajan · 2019 [cited by examiner]
US 20210049476A1 · Davis · 2021 [cited by examiner]
US 20210056970A1 · Jain · 2021 [cited by examiner]
US 20210158175A1 · Trim · 2021 [cited by examiner]
US 20230359654A1 · Johannhoerster · 2023 [cited by examiner]
US 20230419051A1 · Sabapathy · 2023 [cited by examiner]
US 20240378399A1 · Gandhi · 2024 [cited by examiner]
US 20250061286A1 · Zhou · 2025 [cited by examiner]
Ouyang, Long et al., “Training language models to follow instructions with human feedback”, 36th Conference on Neural Information Processing Systems (NeurIPS 2022), 15 pages. [cited by applicant]
Thoppilan, Romal et al., “LaMDA: Language Models for Dialog Applications”, arXiv:2201.08239v3 [cs.CL] Feb. 10, 2022, 47 pages. [cited by applicant]
Touvron, Hugo et al., “LLaMA: Open and Efficient Foundation Language Models”, arXiv:2302.1397v1 [cs/CL] Feb. 27, 2023, 27 pages. [cited by applicant]
Ji, Ziwei et al., “Survey of Hallucination in Natural Language Generation”, ACM Computing Surveys, vol. 1, No. 1, Publication Date: Feb. 2022, 47 pages. [cited by applicant]
Chen, Sihao et al., “Improving Faithfulness in Abstractive Summarization with Contrast Candidate Generation and Selection”, arXiv:2104.09061v1 [cs/CL] Apr. 19, 2021, 7 pages. [cited by applicant]
Manakul, Potsawee et al., “Selfcheckgpt: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models”, arXiv:2303.08896v2 [cs/CL] May 8, 2023, 13 pages. [cited by applicant]
Shuster, Kurt et al., Retrieval Augmentation Reduces Hallucination in Conversation, arXiv:2104.07567v1 [cs/CL] Apr. 15, 2021, 21 pages. [cited by applicant]
Zhou, Chunting et al., “Detecting Hallucinated Content in Conditional Neural Sequence Generation”, arXiv:2011.02593v3 [cs/CL] Jun. 2, 2021, 16 pages. [cited by applicant]
Wan, David et al., “Faithfulness-Aware Decoding Strategies for Abstractive Summarization”, arXiv:2303.03278v1 [cs/CL] Mar. 6, 2023, 17 pages. [cited by applicant]
Huang, Kung-Hsiang et al., “Swing: Balancing Coverage and Faithfulness for Dialogue Summarization”, arXiv:2301.10483v1 [cs/CL] Jan. 25, 2023, 14 pages. [cited by applicant]
Reimers, Nils et al., “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks”, arXiv:1908.10084v1 [cs/CL] Aug. 27, 2019, 11 pages. [cited by applicant]
Firoozeh, Nazanin et al., “Keyword extraction: Issues and methods”, Natural Language Engineering, 2019, 33 pages. [cited by applicant]