IP Library › Granted Patent US 12,737,544
Granted Patent B1
US 12,737,544 · App. 19/264,632 · Granted Sep 15, 2026

Unsupervised detection of hallucinations in language model outputs

Inventors: Shai Ardazi (Petach-Tikva, IL); Matan Vetzler (Givat-Shmuel, IL); Amir Bialer (Tel-Aviv, IL); Noa Haas (Tel Aviv, IL)
Assignee: Intuit Inc.
G06F40/279
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,544
App. No.
19/264,632
Granted
Sep 15, 2026
Kind
B1
Abstract

Certain aspects of the disclosure provide a method for unsupervised detection of hallucinations in generative language model outputs. The method includes receiving a language model output comprising tokens and corresponding confidence scores, identifying low-confidence tokens whose scores fall below a confidence score threshold, and determining unsupervised attributes of these tokens. An anomaly score is then calculated based on the unsupervised attributes, and if this score meets or exceeds a set threshold, the system predicts that the language model output contains at least one hallucination.

Claims (54)

1 . A method for performing unsupervised detection of hallucinations in generative language model output, comprising:

receiving a language model output associated with a language model, the language model output comprising a plurality of tokens and a plurality of confidence scores, wherein each confidence score of the plurality of confidence scores corresponds to a respective token of the plurality of tokens and represents how confident the language model is in a prediction of the respective token of the plurality of tokens;

identifying one or more low-confidence tokens, each having a corresponding confidence score that does not meet a confidence score threshold;

applying a sliding window across the language model output of a window size and a window step size;

calculating one or more local densities of one or more windows based on applying the sliding window across the language model output;

determining one or more unsupervised attributes of the one or more low-confidence tokens, wherein the one or more unsupervised attributes comprise a maximum window density;

determining the maximum window density, wherein the maximum window density is of the one or more local densities;

determining an anomaly score of the language model output based on the one or more unsupervised attributes of the one or more low-confidence tokens;

determining that the anomaly score of the language model output at least meets an anomaly score threshold; and

generating a prediction that the language model output comprises at least one hallucination output by the language model.

2 . The method of claim 1 , wherein:

the one or more unsupervised attributes further comprise an average window density, and

the method further comprises determining the average window density, wherein the average window density is of the one or more local densities.

3 . The method of claim 1 , wherein:

the one or more unsupervised attributes comprise a density of low-confidence tokens, and

the method further comprises determining the density of low-confidence tokens by determining a ratio of low-confidence tokens to a total number of tokens of the language model output.

4 . The method of claim 1 , wherein the one or more unsupervised attributes comprise a positional encoding of each of the one or more low-confidence tokens relative to the language model output.

5 . The method of claim 4 , further comprising:

segmenting the language model output into two or more segments; and

determining the positional encoding of each of the one or more low-confidence tokens by determining a ratio of low-confidence tokens to a total number of tokens in each segment of the two or more segments.

6 . The method of claim 1 , wherein:

the one or more unsupervised attributes comprise one or more cluster attributes, and the method further comprises:

identifying one or more clusters of the one or more low-confidence tokens; and

identifying the one or more cluster attributes, wherein the one or more cluster attributes are of the one or more clusters.

7 . The method of claim 6 , wherein the one or more cluster attributes comprise a number of clusters of low-confidence tokens.

8 . The method of claim 6 , wherein the one or more cluster attributes comprise an average cluster size of the one or more clusters.

9 . The method of claim 1 , wherein the anomaly score of the language model output is determined based on applying a local outlier factor algorithm.

10 . An apparatus comprising a processing system that includes one or more processors and one or more memories coupled with the one or more processors, the processing system configured to:

receive a language model output associated with a language model, the language model output comprising a plurality of tokens and a plurality of confidence scores, wherein each confidence score of the plurality of confidence scores corresponds to a respective token of the plurality of tokens and represents how confident the language model is in a prediction of the respective token of the plurality of tokens;

identify one or more low-confidence tokens, each having a corresponding confidence score that does not meet a confidence score threshold;

apply a sliding window across the language model output on a window size and a window step size;

calculate one or more local densities of one or more windows based on applying the sliding window across the language model output;

determine one or more unsupervised attributes of the one or more low-confidence tokens, wherein the one or more unsupervised attributes comprise an average window density;

determine the average window density, wherein the average window density is of the one or more local densities;

determine an anomaly score of the language model output based on the one or more unsupervised attributes of the one or more low-confidence tokens;

determine that the anomaly score of the language model output at least meets an anomaly score threshold; and

generate a prediction that the language model output comprises at least one hallucination output by the language model.

11 . The apparatus of claim 10 , wherein:

the one or more unsupervised attributes further comprise a maximum window density, and

the processing system is further configured to determine the maximum window density, wherein the maximum window density is of the one or more local densities.

12 . The apparatus of claim 10 , wherein:

the one or more unsupervised attributes comprise a density of low-confidence tokens, and

the processing system is further configured to determine the density of low-confidence tokens by determining a ratio of low-confidence tokens to a total number of tokens of the language model output.

13 . The apparatus of claim 10 , wherein the one or more unsupervised attributes comprise a positional encoding of each of the one or more low-confidence tokens relative to the language model output.

14 . The apparatus of claim 13 , wherein the processing system is further configured to:

segment the language model output into two or more segments; and

determine the positional encoding of each of the one or more low-confidence tokens by determining a ratio of low-confidence tokens to a total number of tokens in each segment of the two or more segments.

15 . The apparatus of claim 10 , wherein:

the one or more unsupervised attributes comprise one or more cluster attributes, and the processing system is further configured to:

identify one or more clusters of the one or more low-confidence tokens; and

identify the one or more cluster attributes, wherein the one or more cluster attributes are of the one or more clusters.

16 . The apparatus of claim 15 , wherein the one or more cluster attributes comprise a number of clusters of low-confidence tokens.

17 . The apparatus of claim 15 , wherein the one or more cluster attributes comprise an average cluster size of the one or more clusters.

18 . The apparatus of claim 10 , wherein the anomaly score of the language model output is determined based on applying a local outlier factor algorithm.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2025
From: ARDAZI, SHAI; VETZLER, MATAN; BIALER, AMIR; HAAS, NOA
To: INTUIT INC.
Reel/Frame 071928/0264 →
References Cited (17)
US 12373649B1 · Mathews · 2025 [cited by examiner]
US 12443638B1 · Sinha · 2025 [cited by examiner]
US 20170337181A1 · Belov · 2017 [cited by examiner]
US 20220414320A1 · Dolan · 2022 [cited by examiner]
US 20230062177A1 · Landry · 2023 [cited by examiner]
US 20230068145A1 · Landry · 2023 [cited by examiner]
US 20240184988A1 · Sridhar · 2024 [cited by examiner]
US 20240386207A1 · Emrey · 2024 [cited by examiner]
US 20240394512A1 · Cunningham · 2024 [cited by examiner]
US 20240394600A1 · Cunningham · 2024 [cited by examiner]
US 20240419912A1 · Somech · 2024 [cited by examiner]
US 20250061286A1 · Zhou · 2025 [cited by examiner]
US 20250111151A1 · Huang · 2025 [cited by examiner]
US 20250259115A1 · Jeyashekar et al. · 2025 [cited by applicant]
US 20250291828A1 · Wood · 2025 [cited by examiner]
US 20260050807A1 · Watson · 2026 [cited by examiner]
US 20260065029A1 · Lee · 2026 [cited by examiner]