IP Library Granted Patent US 12,737,538
Granted Patent B2
US 12,737,538 · App. 18/924,625 · Granted Sep 15, 2026

Detection of artificial authors

Inventor: Claude-Alain Roten (Vevey, CH)
G06F40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,538
App. No.
18/924,625
Granted
Sep 15, 2026
Kind
B2
Abstract

A computer implemented method for determining if a questioned text has been produced by a human or by a generative artificial intelligence based conversational agent. The method includes the steps of: retrieving from the test text a feature (y) representing the redundancy of the test text; using this feature for determining if the test text has been produced by a human or by a generative artificial intelligence based conversational agent.

Claims (44)

1 . A computer implemented method for determining if a test text has been produced by a human or by a generative artificial intelligence based conversational agent, comprising the step of:

retrieving from the test text a feature (y) representing the redundancy of the test text;

using this feature for determining if the test text has been produced by a human or by a generative artificial intelligence based conversational agent,

wherein the method comprises the step of computing a Bayes factor (BF)

BF

=

f

(

y

|

H

1

)

f

(

y

|

H

2

)

.

where H 1 corresponds to the hypothesis that the author of the test text is a human,

H 2 corresponds to the hypothesis that the author of a given questioned document is a GenAI-based conversational agent, and

y represents the redundancy feature,

and wherein a value of BF less than one is used as an indicator that the test text might have been produced by generative artificial intelligence based conversational agent.

2 . The method of claim 1 , wherein said redundancy (y) is computed by measuring the number of repetitions of feature N-grams in the test text.

3 . The method of claim 2 , said feature N-grams comprising one or more of words, word N-Grams, pattern N-grams of characters, word length N-grams, and punctuation N-grams, wherein N represent an integer such as 2, 3, 4, 5 or more.

4 . The method according to claim 1 , wherein support of said stylometric features related to the corresponding Bayes factor value is defined according to predetermined scale comprising the levels “weak”, “moderate”, “moderately strong”, “strong”, “very strong” and “extremely strong”.

5 . The method according to claim 1 , further comprising the step of selecting the most representative feature N-grams, such as the 20% most representation feature N-grams, said most representative feature N-grams being of rare or of moderate usage.

6 . The method according to claim 1 , further comprising a classification step wherein a first decision (d1) is taken for a test text considered as written by a human being and a second decision (d2) is taken for a test text considered generated by a generative artificial intelligence based conversational agent, wherein a loss parameter (l1, l2) represents the loss incurred when one of the first (d1) or second (d2) decision is false, and wherein the first (d1) and second (d2) decisions are taken so as to minimize the loss.

7 . The method according to claim 6 , wherein the probability of hypothesis H1 (Pr(H1)) equals the probability of the hypothesis (H2) and wherein the loss function l1 equals the loss function l2.

8 . The method according to claim 6 , wherein the probability of hypothesis H1 (Pr(H1)) differs from the probability of the hypothesis (H2) and/or wherein the loss function 11 differs from the loss function 12 so as to consider more severely a falsely classified text generated by a generative artificial intelligence based conversational agent, than a falsely classified text written by a human being.

9 . The method according to claim 8 , wherein l2=T×l1 so that falsely classifying a text as written by an artificial intelligence is considered T times as serious as the opposite, wherein T denotes an integer such as 5, 10 or 20.

10 . The method according to claim 1 , further comprising measuring the variability of stylometric features within said test text, using this variability, in combination with said feature representing the redundancy of the test text, for determining if the test text has been produced by a human or by a generative artificial intelligence based conversational agent.

11 . The method according to claim 10 , wherein reference texts are grouped in a first cluster (X) related to stylometric features frequently used in a first style of text and rarely used in the second style of text, or in a second cluster (Y) related to stylometric features frequently used in the second style of text and rarely used in the first one.

12 . The method according to claim 11 , wherein relevant stylometric features of a test text are evaluated so that the test text can be considered as being part of a first cluster (X), corresponding to a human text, or a second cluster (Y) generated from a generative artificial intelligence based conversational agent.

13 . Method of claim 10 , wherein the accuracy of the analysis is defined according the closest neighbour.

14 . Method according to claim 1 , wherein each sequence of test text and reference text is represented on a histogram by a coloured bar, each sequence of test text and reference text is compared to texts of reference style, and represented in a multidimensional space wherein each dimension relates to the frequency of a given stylometric feature, wherein each of said colours refers to a reference style.

15 . Method according to claim 13 , wherein the size of each of said bars is proportional to the normalised distance of the text to the barycenter of the closest reference texts.

16 . The method according to claim 14 , further comprising a step of accuracy of the calibration based on comparison with closest neighbours.

17 . The method according to claim 13 , wherein a text having an homogenous style, shows an accuracy of more than 90%, or more than 95% or more than 97%, and a text combining several known styles shows a lower accuracy.

18 . Method according to claim 1 , wherein any stylometric feature and analysis is validated by a Bayesian process.

Priority Claims (1)
CH 001181/2023 · Oct 24, 2023 · national
Continuity (1)
Related Publication 20250131193A1 · Apr 24, 2025
References Cited (18)
US 12238322B2 · Luo · 2025 [cited by examiner]
US 12321831B1 · Karpman · 2025 [cited by examiner]
US 20190050388A1 · Myriam et al. · 2019 [cited by applicant]
US 20210174017A1 · Genilloud et al. · 2021 [cited by applicant]
US 20210390447A1 · Cheruvu · 2021 [cited by examiner]
US 20240296288A1 · Bitton · 2024 [cited by examiner]
US 20240331707A1 · Siekman · 2024 [cited by examiner]
US 20240406003A1 · Murialdo · 2024 [cited by examiner]
WO 2008036059A1 · 2008 [cited by applicant]
WO 2017144939A1 · 2017 [cited by applicant]
Keerthichowdary, Naive Bayes Classification | DetectAI-Generated Text, Dec. 2, 2023, Medium, All pages. [cited by examiner]
Swiss Search Report Issued in Corresponding Swiss Patent Application No. CH001181/2023, dated Jun. 14, 2024, 5 pages. [cited by applicant]
Roten et al. “Detecter par stylometrie la fraude academique utilisant ChatGPT”, Cahiers IRAFPA, vol. 1, No. 1, Jul. 14, 2023, 11 pages, Accessed via Internet URL: https://cahiers.irafpa.org/article/view/4126/3703. [cited by applicant]
Bozza et al., “A model independent redundancy measure for human versus ChatGPT authorship discrimination using a Bayesian probabilistic approach”, Scientific Reports, vol. 13, No. 1, Nov. 6, 2023, 8 pages, Accessed via … [cited by applicant]
Anonymous, “GPTZero Wikipedia,” Oct. 22, 2023, 5 pages, Accessed via Internet URL: https://en.wikipedia.org/wiki/GPTZero. [cited by applicant]
Gehrmann et al., “GLTR: Statistical Detection and Visualization of Generated Text,” Jun. 10, 2019, 6 pages, Accessed via Internet URL: https://aclanthology.org/P19-3019.pdf. [cited by applicant]
Badaskar et al., “Identifying Real or Fake Articles: Towards better Language Modeling,” Third International Joint Conference On Natural Language Processing, Jan. 7, 2008, 6 pages, Accessed via Internet URL: https://acla… [cited by applicant]
Anonymous, “Stylometry Wikipedia,” Oct. 17, 2023, 17 pages, Accessed via Internet URL: https://en.wikipedia.org/wiki/Stylometry. [cited by applicant]