IP Library › Granted Patent US 12,436,973
Granted Patent B2
US 12,436,973 · App. 18/383,557 · Granted Oct 7, 2025

Data tagging and prompt generation system

Inventors: Jing-tao Li (Shaanxi, CN); Jian Song (Xi'an, CN); Ming Yan (Shaanxi, CN); Jingyuan Li (Xi'an, CN); Bo Dang (Xi'an, CN)
Assignee: SAP SE
G06F16/285G06F16/2462
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,436,973
App. No.
18/383,557
Filed
Oct 25, 2023
Granted
Oct 7, 2025
Kind
B2
Art Unit
2167
USPC
707/738
Abstract

System, method, and various embodiments for data tagging and prompt generation are described herein. An embodiment operates by receiving input data, identifying metadata, generating one or more statistics based on the input data, calculating a sample size for the input data based on the one or more statistics and extracting a sample of the input data of the sample size. A prompt is generated based on a prompt template, and the prompt is provided to a language model configured to tag the input in accordance with the prompt. The output including tagged input data is received, and a query is executed against the tagged input data.

Claims (76)

1. A method comprising:

receiving, by one or more processors, input data comprising data to be tagged by a language model;

identifying metadata associated with the input data, wherein the metadata comprises a name by which to refer to the input data;

generating one or more statistics based on the input data, the one or more statistics comprising a total number of data items in the input data;

calculating a sample size for the input data based on the one or more statistics, wherein the sample size is less than the total number of data items in the input data;

extracting a sample of the input data in accordance with the sample size, wherein the sample of the input data comprises a subset of the input data;

generating a prompt based on a prompt template, the prompt template comprising an input segment comprising the metadata and the sample of the input data, and an output segment identifying a format for an output;

providing the prompt to the language model configured to generate one or more tags based on the sample of the input data, and tag the input data with the one or more tags in accordance with the prompt;

receiving the output comprising tagged input data which was tagged with one or more tags generated based on the sample of the input data and in accordance with the format, wherein the tagged input data includes a semantic meaning or semantic context of the input data;

storing the tagged input data in a database;

executing a query against the tagged input data stored in the database; and

returning a result of the query.

2. The method of claim 1 , wherein the receiving comprises:

identifying sensitive data and non-sensitive data from the input data;

extracting the sensitive data, wherein only the non-sensitive data is provided to the language model; and

tagging the sensitive data independently of the output.

3. The method of claim 1 , wherein the language model comprises an artificial intelligence language model configured to perform a variety of tasks including tagging the input data, and wherein the artificial intelligence language model is operating one or more different processors.

4. The method of claim 1 , further comprising:

receiving a request for additional data, after providing the prompt and prior to receiving the output;

extracting a second sample of the input data in accordance with the sample size; and

generating a second prompt comprising the second sample; and

providing the second prompt including the second sample to the language model.

5. The method of claim 4 , wherein the second sample is a same size as the sample size.

6. The method of claim 1 , wherein the input data comprises a table from a database, the table comprising a plurality of columns, each column including a plurality of rows.

7. The method of claim 6 , wherein at least a subset of the plurality of columns from the table include tags generated by the language model.

8. A system comprising:

a memory; and

at least one processor coupled to the memory and configured to perform operations comprising:

receiving input data comprising data to be tagged by a language model;

identifying metadata associated with the input data, wherein the metadata comprises a name by which to refer to the input data;

generating one or more statistics based on the input data, the one or more statistics comprising a total number of data items in the input data;

calculating a sample size for the input data based on the one or more statistics, wherein the sample size is less than the total number of data items in the input data;

extracting a sample of the input data in accordance with the sample size, wherein the sample of the input data comprises a subset of the input data;

generating a prompt based on a prompt template, the prompt template comprising an input segment comprising the metadata and the sample of the input data, and an output segment identifying a format for an output;

providing the prompt to the language model configured to generate one or more tags based on the sample of the input data, and tag the input data with the one or more tags in accordance with the prompt;

receiving the output comprising tagged input data which was tagged with one or more tags generated based on the sample of the input data and in accordance with the format, wherein the tagged input data includes a semantic meaning or semantic context of the input data;

storing the tagged input data in a database;

executing a query against the tagged input data stored in the database; and

returning a result of the query.

9. The system of claim 8 , wherein the receiving comprises:

identifying sensitive data and non-sensitive data from the input data;

extracting the sensitive data, wherein only the non-sensitive data is provided to the language model; and

tagging the sensitive data independently of the output.

10. The system of claim 8 , wherein the language model comprises an artificial intelligence language model configured to perform a variety of tasks including tagging the input data, and wherein the artificial intelligence language model is operating one or more different processors.

11. The system of claim 8 , the operations further comprising:

receiving a request for additional data, after providing the prompt and prior to receiving the output;

extracting a second sample of the input data in accordance with the sample size; and

generating a second prompt comprising the second sample; and

providing the second prompt including the second sample to the language model.

12. The system of claim 11 , wherein the second sample is a same size as the sample size.

13. The system of claim 8 , wherein the input data comprises a table from a database, the table comprising a plurality of columns, each column including a plurality of rows.

14. The system of claim 13 , wherein at least a subset of the plurality of columns from the table include tags generated by the language model.

15. A non-transitory computer-readable device having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:

receiving input data comprising data to be tagged by a language model;

identifying metadata associated with the input data, wherein the metadata comprises a name by which to refer to the input data;

generating one or more statistics based on the input data, the one or more statistics comprising a total number of data items in the input data;

calculating a sample size for the input data based on the one or more statistics, wherein the sample size is less than the total number of data items in the input data;

extracting a sample of the input data in accordance with the sample size, wherein the sample of the input data comprises a subset of the input data;

generating a prompt based on a prompt template, the prompt template comprising an input segment comprising the metadata and the sample of the input data, and an output segment identifying a format for an output;

providing the prompt to the language model configured to generate one or more tags based on the sample of the input data, and tag the input data with the one or more tags in accordance with the prompt;

receiving the output comprising tagged input data which was tagged with one or more tags generated based on the sample of the input data and in accordance with the format, wherein the tagged input data includes a semantic meaning or semantic context of the input data;

storing the tagged input data in a database;

executing a query against the tagged input data stored in the database; and

returning a result of the query.

16. The non-transitory computer-readable device of claim 15 , wherein the receiving comprises:

identifying sensitive data and non-sensitive data from the input data;

extracting the sensitive data, wherein only the non-sensitive data is provided to the language model; and

tagging the sensitive data independently of the output.

17. The non-transitory computer-readable device of claim 15 , wherein the language model comprises an artificial intelligence language model configured to perform a variety of tasks including tagging the input data, and wherein the artificial intelligence language model is operating one or more different processors.

18. The non-transitory computer-readable device of claim 15 , the operations further comprising:

receiving a request for additional data, after providing the prompt and prior to receiving the output;

extracting a second sample of the input data in accordance with the sample size; and

generating a second prompt comprising the second sample; and

providing the second prompt including the second sample to the language model.

19. The non-transitory computer-readable device of claim 18 , wherein the second sample is a same size as the sample size.

20. The method of claim 7 , wherein a first tag of the one or more tags is used as a column name of a first column of the plurality of columns.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2023
From: LI, JING-TAO; SONG, JIAN; YAN, MING; LI, JINGYUAN; DANG, BO
To: SAP SE
Reel/Frame 065353/0591 →
Continuity (1)
Related Publication 20250139128A1 · May 1, 2025
References Cited (2)
US 11516158B1 · Luzhnica · 2022 [cited by examiner]
US 20190258904A1 · Ma · 2019 [cited by examiner]