IP Library Granted Patent US 12694211
Granted Patent B2
US 12694211 · App. 18/645,563 · Granted Jul 28, 2026

System and method for intelligent evaluation of artificial intelligence generated texts

Inventor: Danny Butvinik (Haifa, IL)
Assignee: Actimize Ltd.
G06F40/284G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12694211
App. No.
18/645,563
Granted
Jul 28, 2026
Kind
B2
Abstract

A system and method for automatically evaluating computer generated content may include: calculating a plurality of metrics for an input text, where the plurality of metrics may include one or more perplexity scores describing a prediction of the input text by a large language model (LLM); determining, based on one or more of the calculated metrics, whether to accept or reject the input text; and performing an exchange of data between remotely connected computer devices based on the determining to accept or reject the text. In some embodiments, calculating of metrics and determining whether to accept or reject the input text may be performed without relying on any information received subsequent to the initial receiving of the input text. Some embodiments may perform automated computerized actions such as, e.g., deploy or discard an update to the LLM based on the determining whether to accept or reject the input text.

Claims (34)

1 . A computerized method of automatically evaluating computer generated content, the method comprising, using a computer processor:

calculating a plurality of metrics for an input text, the plurality of metrics comprising one or more perplexity scores, the one or more perplexity scores describing a prediction of the input text by a large language model (LLM);

determining, based on one or more of the calculated plurality of metrics, whether to accept or reject the input text; and

performing an exchange of data between remotely connected computer devices based on the determining to accept or reject the input text.

2 . The method of claim 1 , wherein the one or more perplexity scores describe predicting of one or more individual sentences by the LLM, and wherein one or more of the calculated plurality of metrics comprises a perplexity table, the perplexity table describing the one or more individual sentences.

3 . The method of claim 1 , wherein the plurality of metrics include at least one of: a text sentiment metric, a text diversity metric, a text readability metric, and an entity consistency metric.

4 . The method of claim 1 , wherein one or more of the plurality of metrics are calculated using a plurality of vector embeddings, wherein one or more of the plurality of vector embeddings describes a plurality of sentences in the input text, and wherein one or more of the plurality of vector embeddings describes a plurality of words in the input text.

5 . The method of claim 1 , wherein the calculating the plurality of metrics and the determining whether to accept or reject the input text are performed without relying on any information received subsequent to a receiving of the input text.

6 . The method of claim 1 , comprising creating an assessment table, the assessment table comprising one or more of the calculated plurality of metrics and one or more interpretations of one or more of the calculated plurality of metrics; and

uploading the assessment table to a web page.

7 . The method of claim 1 , comprising deploying an update to the LLM based on the determining whether to accept or reject the input text.

8 . A computerized system for automatically evaluating computer generated content, the system comprising:

a memory,

and a computer processor configured to:

calculate a plurality of metrics for an input text, the plurality of metrics comprising one or more perplexity scores, the one or more perplexity scores describing a prediction of the input text by a large language model (LLM);

determine, based on one or more of the calculated plurality of metrics, whether to accept or reject the input text; and

perform an exchange of data between remotely connected computer devices based on the determining to accept or reject the input text.

9 . The computerized system of claim 8 , wherein the one or more perplexity scores describe predicting of one or more individual sentences by the LLM, and wherein one or more of the calculated plurality of metrics comprises a perplexity table, the perplexity table describing the one or more individual sentences.

10 . The computerized system of claim 8 , wherein the plurality of metrics include at least one of: a text sentiment metric, a text diversity metric, a text readability metric, and an entity consistency metric.

11 . The computerized system of claim 8 , wherein one or more of the plurality of metrics are calculated using a plurality of vector embeddings, wherein one or more of the plurality of vector embeddings describes a plurality of sentences in the input text, and wherein one or more of the plurality of vector embeddings describes a plurality of words in the input text.

12 . The computerized system of claim 8 , wherein the calculating the plurality of metrics and the determining whether to accept or reject the input text are performed without relying on any information received subsequent to a receiving of the input text.

13 . The computerized system of claim 8 , wherein the processor is to create an assessment table, the assessment table comprising one or more of the calculated plurality of metrics and one or more interpretations of one or more of the calculated plurality of metrics; and

uploading the assessment table to a web page.

14 . The computerized system of claim 8 , wherein the processor is to deploy an update to the LLM based on the determining whether to accept or reject the input text.

15 . A computerized method of automatically evaluating an input text, the method comprising, using a computer processor:

computing a plurality of scores for the input text, the plurality of computed scores comprising one or more perplexity metrics, the one or more perplexity metrics describing a generation of the input text by a generative artificial intelligence (GenAI) model;

determining, based on one or more of the plurality of computed scores, whether to approve or dismiss the input text; and

saving an update to the GenAI model based on the determining whether to approve or dismiss the input text.

16 . The method of claim 15 , wherein the one or more perplexity metrics describe generating of one or more individual sentences by the GenAI model, and wherein one or more of the plurality of computed scores comprises a perplexity table, the perplexity table describing the one or more individual sentences.

17 . The method of claim 15 , wherein the plurality of scores include at least one of: a text sentiment score, a text diversity score, a text readability score, and an entity consistency score.

18 . The method of claim 15 , wherein one or more of the plurality of scores are computed using a plurality of vector representations, wherein one or more of the plurality of vector representations describe a plurality of sentences in the input text, and wherein one or more of the plurality of vector representations describes a plurality of words in the input text.

19 . The method of claim 15 , wherein the computing the plurality of scores and the determining whether to approve or dismiss the input text are performed without relying on any data received following a receiving of the input text.

20 . The method of claim 15 , comprising creating an output assessment, the output assessment comprising a table, the table including one or more of the plurality of computed scores and one or more descriptions of one or more of the plurality of computed scores; and

uploading the output assessment to a web page.