IP Library › Granted Patent US 12,730,970
Granted Patent B2
US 12,730,970 · App. 18/671,681 · Granted Sep 8, 2026

Generation method for sample data, device and storage medium

Inventors: Wenqi Xie (Beijing, CN); Zhaosha Fan (Beijing, CN); Xiaodong Su (Beijing, CN); Minglei Li (Beijing, CN)
Assignee: BEIJING VOLCANO ENGINE TECHNOLOGY CO., LTD.
G06F40/284G06F40/253
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,730,970
App. No.
18/671,681
Granted
Sep 8, 2026
Kind
B2
Abstract

A generation method and apparatus for sample data, a method and apparatus for information detection, a device and a storage medium are provided, and the generation method includes: acquiring first reference data, the first reference data including target information matching a target information type, and the target information type being a preset information type with a security requirement; performing analysis processing on the target information in the first reference data to generate an analysis result corresponding to the target information, the analysis processing including semantic analysis, lexical structure analysis and grammatical structure analysis; generating a plurality of positive sample information and a plurality of negative sample information based on the analysis result corresponding to the target information; generating a sample data set including positive sample data and negative sample data based on the plurality of positive sample information, the plurality of negative sample information, and second reference data.

Claims (77)

1 . A generation method for sample data, executed on a computer device and comprising:

acquiring, by a processor, first reference data, wherein the first reference data comprises target information matching a target information type, and the target information type is a preset information type with a security requirement;

performing, by the processor, an analysis processing on the target information in the first reference data to generate an analysis result corresponding to the target information, wherein the analysis processing comprises semantic analysis, lexical structure analysis and grammatical structure analysis;

generating a plurality of positive sample information and a plurality of negative sample information based on the analysis result corresponding to the target information; and

generating a sample data set comprising positive sample data and negative sample data based on the plurality of positive sample information, the plurality of negative sample information, and second reference data,

wherein the performing analysis processing on the target information in the first reference data to generate an analysis result corresponding to the target information comprises:

performing the semantic analysis on the target information in the first reference data to generate at least one first keyword respectively corresponding to at least one target information type;

performing the lexical structure analysis on the target information in the first reference data to generate at least one first regular expression respectively corresponding to the at least one target information type, wherein each of the at least one first regular expression is used to characterize a lexical structure matching the target information type;

performing the grammatical structure analysis on the target information in the first reference data to generate an information template matching a data type of the first reference data; and

generating the analysis result corresponding to the target information based on the at least one first keyword and the at least one first regular expression respectively corresponding to the at least one target information type and the information template matching the data type of the first reference data,

wherein the sample data set is used to train a specific data detection model.

2 . The generation method according to claim 1 , wherein the generating the plurality of positive sample information and the plurality of negative sample information based on the analysis result corresponding to the target information comprises:

for each target information type of the at least one target information type, generating a plurality of first information sample values which correspond to the target information type and satisfy a lexical structure of the target information type based on a first regular expression corresponding to the target information type; and

according to an information template indicated by the analysis result, generating the plurality of positive sample information of the target information type based on a first keyword and the plurality of first information sample values corresponding to the target information type.

3 . The generation method according to claim 1 , wherein the generating the plurality of positive sample information and the plurality of negative sample information based on the analysis result corresponding to the target information comprises:

for the each target information type of the at least one target information type, performing preset operation on a first keyword corresponding to the target information type to generate a second keyword, wherein the preset operation comprises at least one selected from the group consisting of a truncation operation and a character addition operation;

generating a second information sample value which does not satisfy a lexical structure of the target information type based on a first regular expression corresponding to the target information type; and

according to an information template indicated by the analysis result, generating a plurality of negative sample information of the target information type based on the second keyword and the second information sample value corresponding to the target information type.

4 . The generation method according to claim 3 , wherein the generating the second information sample value which does not satisfy a lexical structure of the target information type based on a first regular expression corresponding to the target information type comprises:

generating a first information sample value corresponding to the target information type based on the first regular expression corresponding to the target information type; performing a preset operation on the first information sample value corresponding to the target information type to generate the second information sample value; and/or

generating a second regular expression which does not satisfy the lexical structure of the target information type based on the first regular expression corresponding to the target information type; generating the second information sample value corresponding to the target information type based on the second regular expression.

5 . The generation method according to claim 1 , wherein the first reference data further comprises confusing information, the confusing information is information that interferes with detection of the target information, and the generation method further comprises:

performing the semantic analysis on the confusing information in the first reference data to generate a third keyword corresponding to at least one target information type;

determining a third information sample value corresponding to the third keyword from the confusing information; and

generating a plurality of negative sample information of the target information type based on the third keyword and the third information sample value corresponding to the at least one target information type.

6 . The generation method according to claim 1 , wherein a number of pieces of second reference data is more than one, and the generating a sample data set comprising positive sample data and negative sample data based on the plurality of positive sample information, the plurality of negative sample information, and the second reference data comprises:

for each piece of the second reference data, determining an insertion scheme of the second reference data based on a set proportional parameter and a random number generated for the second reference data, wherein the insertion scheme comprises inserting positive sample information, inserting negative sample information, and not inserting sample information;

when the insertion scheme of the second reference data is inserting target sample information, inserting the target sample information into the second reference data to generate target sample data, wherein the target sample information comprises at least one selected from a group consisting of the positive sample information and the negative sample information, when the target sample information comprises the positive sample information, the target sample data comprises the positive sample data, and when the target sample information comprises the negative sample information, the target sample data comprises the negative sample data;

determining labeling information of the positive sample data, wherein the labeling information comprises the target information type, information sample value, an initial index position of the information sample value in the positive sample data, and content information of the information sample value in the positive sample data; and

forming the sample data set based on a plurality of pieces of negative sample data and positive sample data associated with the labeling information.

7 . The generation method according to claim 6 , wherein the inserting the target sample information into the second reference data to generate target sample data comprises:

determining an insertion parameter corresponding to the second reference data, wherein the insertion parameter comprises a number of insertion positions, a number of samples corresponding to each of the insertion positions, and a target information type corresponding to the each of the number of insertion positions;

determining insertion positions matching the number of insertion positions from the second reference data;

acquiring sample information to be inserted corresponding to each of the insertion positions according to the number of samples corresponding to each of the insertion positions and the target information type corresponding to each of the insertion positions; and

inserting the sample information to be inserted corresponding to each of the insertion positions into the second reference data to generate target sample data.

8 . The generation method according to claim 1 , wherein the sample data set generated according to the generation method is used to train an information detection model, the information detection model is used to detect information content comprised in data to be detected to obtain a detection result corresponding to the data to be detected, and when the detection result indicates that the data to be detected comprises target information belonging to a target information type, prompt information is generated.

9 . A computer device, comprising:

at least one processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the at least one processor; the at least one processor communicates with the memory through the bus upon running of the computer device, and the machine-readable instructions, upon being executed by the at least one processor, execute a generation method for sample data, and the generation method for sample data comprises:

acquiring first reference data, wherein the first reference data comprises target information matching a target information type, and the target information type is a preset information type with a security requirement;

performing analysis processing on the target information in the first reference data to generate an analysis result corresponding to the target information, wherein the analysis processing comprises semantic analysis, lexical structure analysis and grammatical structure analysis;

generating a plurality of positive sample information and a plurality of negative sample information based on the analysis result corresponding to the target information; and

generating a sample data set comprising positive sample data and negative sample data based on the plurality of positive sample information, the plurality of negative sample information, and second reference data,

wherein the performing analysis processing on the target information in the first reference data to generate an analysis result corresponding to the target information comprises:

performing the semantic analysis on the target information in the first reference data to generate at least one first keyword respectively corresponding to at least one target information type;

performing the lexical structure analysis on the target information in the first reference data to generate at least one first regular expression respectively corresponding to the at least one target information type, wherein each of the at least one first regular expression is used to characterize a lexical structure matching the target information type;

performing the grammatical structure analysis on the target information in the first reference data to generate an information template matching a data type of the first reference data; and

generating the analysis result corresponding to the target information based on the at least one first keyword and the at least one first regular expression respectively corresponding to the at least one target information type and the information template matching the data type of the first reference data,

wherein the sample data set is used to train a specific data detection model.

10 . The computer device according to claim 9 , wherein the generating the plurality of positive sample information and the plurality of negative sample information based on the analysis result corresponding to the target information comprises:

for each target information type of the at least one target information type, generating a plurality of first information sample values which correspond to the target information type and satisfy a lexical structure of the target information type based on a first regular expression corresponding to the target information type; and

according to an information template indicated by the analysis result, generating the plurality of positive sample information of the target information type based on a first keyword and the plurality of first information sample values corresponding to the target information type.

11 . The computer device according to claim 9 , wherein the generating the plurality of positive sample information and the plurality of negative sample information based on the analysis result corresponding to the target information comprises:

for the each target information type of the at least one target information type, performing preset operation on a first keyword corresponding to the target information type to generate a second keyword, wherein the preset operation comprises at least one selected from the group consisting of a truncation operation and a character addition operation;

generating a second information sample value which does not satisfy a lexical structure of the target information type based on a first regular expression corresponding to the target information type; and

according to an information template indicated by the analysis result, generating the plurality of negative sample information of the target information type based on the second keyword and the second information sample value corresponding to the target information type.

12 . The computer device according to claim 11 , wherein the generating a second information sample value which does not satisfy a lexical structure of the target information type based on a first regular expression corresponding to the target information type comprises:

generating a first information sample value corresponding to the target information type based on the first regular expression corresponding to the target information type; performing a preset operation on the first information sample value corresponding to the target information type to generate the second information sample value; and/or

generating a second regular expression which does not satisfy the lexical structure of the target information type based on the first regular expression corresponding to the target information type; generating the second information sample value corresponding to the target information type based on the second regular expression.

13 . The computer device according to claim 9 , wherein the first reference data further comprises confusing information, the confusing information is information that interferes with detection of the target information, and the generation method further comprises:

performing the semantic analysis on the confusing information in the first reference data to generate a third keyword corresponding to at least one target information type;

determining a third information sample value corresponding to the third keyword from the confusing information; and

generating the plurality of negative sample information of the target information type based on the third keyword and the third information sample value corresponding to the at least one target information type.

14 . The computer device according to claim 9 , wherein a number of pieces of second reference data is more than one, and the generating a sample data set comprising positive sample data and negative sample data based on the plurality of positive sample information, the plurality of negative sample information, and the second reference data comprises:

for each piece of the second reference data, determining an insertion scheme of the second reference data based on a set proportional parameter and a random number generated for the second reference data, wherein the insertion scheme comprises inserting positive sample information, inserting negative sample information, and not inserting sample information;

when the insertion scheme of the second reference data is inserting target sample information, inserting the target sample information into the second reference data to generate target sample data, wherein the target sample information comprises at least one selected from a group consisting of the positive sample information and the negative sample information, when the target sample information comprises the positive sample information, the target sample data comprises the positive sample data, and when the target sample information comprises the negative sample information, the target sample data comprises the negative sample data;

determining labeling information of the positive sample data, wherein the labeling information comprises the target information type, information sample value, an initial index position of the information sample value in the positive sample data, and content information of the information sample value in the positive sample data; and

forming the sample data set based on a plurality of pieces of negative sample data and positive sample data associated with the labeling information.

15 . The computer device according to claim 14 , wherein the inserting the target sample information into the second reference data to generate target sample data comprises:

determining an insertion parameter corresponding to the second reference data, wherein the insertion parameter comprises a number of insertion positions, a number of samples corresponding to each of the insertion positions, and a target information type corresponding to each of the insertion positions;

determining insertion positions matching the number of insertion positions from the second reference data;

acquiring sample information to be inserted corresponding to each of the insertion positions according to the number of samples corresponding to each of the insertion positions and the target information type corresponding to each of the insertion positions; and

inserting the sample information to be inserted corresponding to each of the insertion positions into the second reference data to generate target sample data.

16 . A non-transitory computer-readable storage medium storing computer programs, wherein the computer programs, upon being run by at least one processor, execute a generation method for sample data, and the generation method for sample data comprises:

acquiring first reference data, wherein the first reference data comprises target information matching a target information type, and the target information type is a preset information type with a security requirement;

performing analysis processing on the target information in the first reference data to generate an analysis result corresponding to the target information, wherein the analysis processing comprises semantic analysis, lexical structure analysis and grammatical structure analysis;

generating a plurality of positive sample information and a plurality of negative sample information based on the analysis result corresponding to the target information; and

generating a sample data set comprising positive sample data and negative sample data based on the plurality of positive sample information, the plurality of negative sample information, and second reference data.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2026
From: XIE, WENQI
To: BEIJING DOUYIN INFORMATION SERVICE CO., LTD.
Reel/Frame 075451/0667 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2026
From: FAN, ZHAOSHA; SU, XIAODONG; LI, MINGLEI
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 075451/0678 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2026
From: BEIJING DOUYIN INFORMATION SERVICE CO., LTD.; BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
To: BEIJING VOLCANO ENGINE TECHNOLOGY CO., LTD.
Reel/Frame 075451/0690 →
Priority Claims (1)
CN 202310582856.3 · May 22, 2023 · national
Continuity (1)
Related Publication 20240394475A1 · Nov 28, 2024
References Cited (14)
US 8488916B2 · Terman · 2013 [cited by examiner]
US 10075384B2 · Shear · 2018 [cited by examiner]
US 12229676B2 · Charnock · 2025 [cited by examiner]
US 12293289B2 · Charnock · 2025 [cited by examiner]
US 12401835B2 · Govindarajan · 2025 [cited by examiner]
US 20220083813A1 · Du · 2022 [cited by examiner]
US 20240070457A1 · Charnock · 2024 [cited by examiner]
US 20240070458A1 · Charnock · 2024 [cited by examiner]
US 20240394475A1 · Xie · 2024 [cited by examiner]
US 20250071040A1 · Wang · 2025 [cited by examiner]
US 20250193462A1 · Govindarajan · 2025 [cited by examiner]
CN 112163633A · 2021 [cited by examiner]
CN 115759043A · 2023 [cited by examiner]
WO WO2023030513A1 · 2023 [cited by examiner]