IP Library › Granted Patent US 12,608,610
Granted Patent B1
US 12,608,610 · App. 19/219,261 · Granted Apr 21, 2026

Synthetic data generation system and method

Inventors: Nikhil Pareek (Lewes, DE); Rishav Hada (Lewes, DE); Srikanth Malyala (Lewes, DE); N.V.J.K Kartik (Lewes, DE)
Assignee: FUTURE AGI INC.
G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,610
App. No.
19/219,261
Filed
May 27, 2025
Granted
Apr 21, 2026
Kind
B1
Art Unit
2129
USPC
706/16
Abstract

A computer-implemented method for generating synthetic data is provided. The method includes receiving user input specifying domain-specific requirements for synthetic data generation and selecting a scenario type. The scenario type is one of Seedless, Seeded, or combination of Seeded and Knowledge Base (KB). The method defines a structured schema based on the user input. The structured schema includes data fields, relationships between data fields, and distributional targets. Based on the structured schema, the method generates an initial set of synthetic data samples using a neural template-driven generation model trained on domain-specific data. The method applies adversarial contrastive sampling. This involves training a discriminator neural network to distinguish between the initial set of synthetic samples and real data samples. The discriminator neural network is used to identify generated samples similar to real data. A contrastive set of samples dissimilar to those identified as similar by the discriminator is generated. The method integrates the initial set of synthetic data samples with the contrastive set to create the synthetic data.

Claims (42)

1 . A computer-implemented method for generating synthetic data, comprising:

receiving user input specifying domain-specific requirements for synthetic data generation and selecting a scenario type, wherein the scenario type is one of Seedless, Seeded, or combination of Seeded and Knowledge Base (KB);

defining a structured schema based on the user input, wherein the structured schema includes data fields, relationships between data fields, and distributional targets; and

based on the structured schema, generating an initial set of synthetic data samples using a neural template-driven generation model trained on domain-specific data;

applying adversarial contrastive sampling by

training a discriminator neural network to distinguish between the initial set of synthetic data samples and real data samples,

using the discriminator neural network to identify generated samples similar to the real data samples, and

generating a contrastive set of samples dissimilar to those identified as similar by the discriminator, and

integrating the initial set of synthetic data samples with the contrastive set to create the synthetic data,

wherein, in response to the scenario type being Seeded or the combination, the method further comprising:

performing statistical tests, including Chi-Square or Kolmogorov-Smirnov (KS) tests, to detect divergences between the synthetic data and reference data received for the scenario type being Seeded or the combination; and

upon detection of the divergences, adjusting one or more parameters of the synthetic data creation process, wherein the adjustments include modifying a sampling temperature, increasing template variety, or refining the specificity of a retrieval process to enhance statistical alignment between the synthetic data and the reference data.

2 . The computer-implemented method of claim 1 , further comprising iteratively refining the synthetic data by repeating the generation and adversarial contrastive sampling steps until a predetermined diversity threshold is met, the diversity being measured using embedding space distance metrics between samples.

3 . The computer-implemented method of claim 1 , wherein the neural template-driven generation model is a Large Language Model, wherein the Large Language Model is trained on a domain-specific corpus annotated with examples of valid synthetic data corresponding to the structured schema.

4 . The computer-implemented method of claim 1 , further comprising segmenting and indexing large documents into smaller chunks, and storing the chunks in a vector database and/or knowledge graph if the selected scenario involves the use of the knowledge base.

5 . The computer-implemented method of claim 4 , further comprising performing partial classification on a subset of the chunks, allowing for updates to a taxonomy based on the domain-specific data.

6 . The computer-implemented method of claim 1 , further comprising validating the structured schema to ensure compliance with domain constraints and resolving conflicts.

7 . The computer-implemented method of claim 1 , further comprising classifying the user input into entities, relationships, and categories based on the defined schema using a multi-layered classification system when the selected scenario is one of Seeded or combination.

8 . The computer-implemented method of claim 1 , wherein the synthetic data based on the structured schema is synthesized by performing a retrieval-augmented generation process by conducting a semantic similarity search within the KB using a vector database, and extracting contextual information through knowledge graph traversal to ground the generated synthetic data in relevant domain context.

9 . The computer-implemented method of claim 1 , wherein the synthetic data based on the structured schema is synthesized by generating counterfactual data points using causal inference techniques to model complex causal relationships and explore “what-if” scenarios.

10 . The computer-implemented method of claim 1 , wherein the synthetic data based on the structured schema is synthesized by validating the generated synthetic data through local consistency checks, ensuring attribute coherence and adherence to specified constraints, including date and format consistency.

11 . The computer-implemented method of claim 1 , wherein the statistical tests further include distributional tests such as Wasserstein distance and Maximum Mean Discrepancy.

12 . The computer-implemented method of claim 1 , further comprising: re-synthesizing by adjusting retrieval queries, constraints, and re-running relevant steps in the synthetic data generation process for iterative refinement of the data.

13 . The computer-implemented method of claim 1 , wherein the created synthetic data is domain-grounded, privacy-compliant, and ready for machine learning training, testing, and analysis.

14 . The computer-implemented method of claim 1 , further comprising performing adversarial validation and robustness testing to assess the synthetic data against edge cases, ensuring its readiness for deployment in real-world machine learning applications.

15 . The computer-implemented method of claim 1 , further comprising evaluating the synthetic data by applying statistical tests and performing embedding space analysis to assess semantic quality.

16 . The computer-implemented method of claim 1 , further comprising performing anomaly detection on the synthetic data using explainable AI techniques, comprising one or more of SHAP (SHapley Additive explanations), Chain-of-Thought prompting, and, LIME (Local Interpretable Model-agnostic Explanations), to identify gaps or inconsistencies in the synthetic data and provide insights for targeted refinement.

17 . The computer-implemented method of claim 16 , further comprising re-synthesizing the synthetic data upon detection of the anomaly or when the quality of the synthetic data falls below a defined threshold to generate a final synthetic dataset.

18 . A system for generating synthetic data, comprising:

a processor; and

a memory storing instructions that, when executed by the processor, cause the processor to:

receive user input specifying domain-specific requirements for synthetic data generation and selecting a scenario type, wherein the scenario type is one of Seedless, Seeded, or combination of Seeded and Knowledge Base (KB);

define a structured schema based on the user input, wherein the structured schema includes data fields, relationships between data fields, and distributional targets; and

based on the structured schema, generate an initial set of synthetic data samples using a neural template-driven generation model trained on domain-specific data,

apply adversarial contrastive sampling by

training a discriminator neural network to distinguish between the initial set of synthetic data samples and real data samples,

using the discriminator neural network to identify generated samples similar to the real data samples, and

generating a contrastive set of samples dissimilar to those identified as similar by the discriminator, and

integrate the initial set of synthetic data samples with the contrastive set to create the synthetic data,

wherein, in response to the scenario type being Seeded or the KB being integrated, the method further comprising:

performing statistical tests, including Chi-Square or Kolmogorov-Smirnov (KS) tests, to detect divergences between the synthetic data and reference data received for the scenario type being Seeded or the combination; and

upon detection of the divergences, adjusting one or more parameters of the synthetic data creation process, wherein the adjustments include modifying a sampling temperature, increasing template variety, or refining the specificity of a retrieval process to enhance statistical alignment between the synthetic data and the reference data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2025
From: PAREEK, NIKHIL; MALYALA, SRIKANTH; HADA, RISHAV; KARTIK, N.V.J.K
To: FUTURE AGI INC.
Reel/Frame 071442/0678 →
References Cited (11)
US 20220058437A1 · Soni · 2022 [cited by examiner]
US 20230073695A1 · Goodsitt · 2023 [cited by examiner]
US 20230082365A1 · Shiarlis · 2023 [cited by examiner]
US 20230140834A9 · Hazard · 2023 [cited by examiner]
US 20230186026A1 · Arthur · 2023 [cited by examiner]
US 20230409607A1 · Pratik · 2023 [cited by examiner]
US 20240371015A1 · Ansari · 2024 [cited by examiner]
US 20250124332A1 · Nakamura Sakai · 2025 [cited by examiner]
US 20250139418A1 · Saplicki · 2025 [cited by examiner]
WO 2024164723A1 · 2024 [cited by applicant]
Of Song et al. (“Discriminator Contrastive Divergence: Semi-Amortized Generative Modeling by Exploring Energy of the Discriminator”, arXiv, Apr. 5, 2020) (Year: 2020). [cited by examiner]
Cited By (1)
US 12,724,759