IP Library › Granted Patent US 12,282,502
Granted Patent B2
US 12,282,502 · App. 17/659,138 · Granted Apr 22, 2025

Generating synthesized user data

Inventors: Emmet James Whitehead, Jr. (Santa Cruz, CA); Lena Reed (Palo Alto, CA); Afshin Mobramaein Kano (San Francisco, CA)
Assignee: Sauce Labs Inc.
G06F16/3329G06F16/22G06F16/2228G06F16/2477G06F16/252G06F16/285G06F16/316G06F16/338G06F16/367G06F40/20G06N3/006G06N7/01G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,282,502
App. No.
17/659,138
Granted
Apr 22, 2025
Kind
B2
Abstract

Disclosed are examples of systems, apparatuses, methods, and computer program products for generating synthesized user data. A method may involve receiving a data specification schema. A method may involve determining a number of test data objects to be generated. A method may involve defining the test data objects, the defining of each test data object including: determining, from the data specification schema, a number of fields of the test data object to be populated, the fields representing categories of simulated user data; and determining values for the fields, the values simulating user data. The method may involve storing the test data objects in a database. The method may involve generating a tabular data file including or identifying the test data objects, the tabular data file configured to be processed by one or more processors of a computing system during a user data testing procedure of the computing system.

Claims (48)

1. A method for generating synthesized user data, comprising:

receiving a data specification schema that specifies characteristics of test data objects comprising synthetic user data, wherein the characteristics comprise a third-party data source that includes a natural language text generation algorithm;

generating the test data objects, the generating of each test data object including:

determining, from the data specification schema, a number of fields of the test data object to be populated, the fields representing categories of simulated user data, and

determining values for the fields, the values simulating user data, wherein determining the values comprises:

determining values for a first field by querying the third-party data source such that the third-party data source returns, responsive to the query, realistic synthetic user data used to populate the values for the at least one field comprising a sequence of a plurality of words, wherein the third-party data source is associated with an entity other than an entity generating the test data objects,

determining values for a third field by generating a value associated with a second field and determining the value for the third field based on the generated value associated with the second field, wherein the data specification schema indicates that the third field is to have a value that is a function of the value associated with the second field, and

determining values for a fourth field by querying a second third-party data source that is a database, wherein the data specification schema indicates a query provided to the database, and wherein a response to the query corresponds to the value associated with the fourth field;

storing the test data objects in a database, the storing of each test data object including populating the fields with the determined values;

generating a tabular data file including the test data objects; and

transmitting the tabular data file to a computing system, wherein the computing system is configured to utilize the tabular data file comprising the realistic synthetic user data during a user data testing procedure of an application performed by the computing system.

2. The method of claim 1 , wherein the data specification schema indicates a data type associated with each field of the number of fields.

3. The method of claim 1 , wherein querying the third-party data source comprises requesting authentication to use the third-party data source using an authentication token specified in the data specification schema.

4. The method of claim 1 , wherein the third-party data source comprises a third-party machine learning algorithm.

5. The method of claim 1 , wherein the sequence of the plurality of words comprises at least one of: conversational text, a list, or instructions.

6. A system for generating synthesized user data, the system comprising:

a memory; and

one or more processors operatively coupled to the memory, the one or more processors configured to cause:

receiving a data specification schema that specifies characteristics of test data objects comprising synthetic user data, wherein the characteristics comprise a third-party data source that includes a natural language text generation algorithm;

generating the test data objects, the generating of each test data object including:

determining, from the data specification schema, a number of fields of the test data object to be populated, the fields representing categories of simulated user data, and

determining values for the fields, the values simulating user data, wherein determining the values comprises:

determining values for a first field by querying the third-party data source such that the third-party data source returns, responsive to the query, realistic synthetic user data used to populate the values for the at least one field comprising a sequence of a plurality of words, wherein the third-party data source is associated with an entity other than an entity generating the test data objects,

determining values for a third field by generating a value associated with a second field and determining the value for the third field based on the generated value associated with the second field, wherein the data specification schema indicates that the third field is to have a value that is a function of the value associated with the second field, and

determining values for a fourth field by querying a second third-party data source that is a database, wherein the data specification schema indicates a query provided to the database, and wherein a response to the query corresponds to the value associated with the fourth field;

storing the test data objects in a database, the storing of each test data object including populating the fields with the determined values;

generating a tabular data file including the test data objects; and

transmitting the tabular data file to a computing system, wherein the computing system is configured to utilize the tabular data file comprising the realistic synthetic user data during a user data testing procedure of an application performed by the computing system the tabular data file configured to be processed by one or more processors of a computing system during a user data testing procedure of the computing system.

7. The system of claim 6 , wherein the data specification schema indicates a data type associated with each field of the number of fields.

8. The system of claim 6 ,, wherein querying the third-party data source comprises requesting authentication to use the third-party data source using an authentication token specified in the data specification schema.

9. The system of claim 6 , wherein the third-party data source comprises a third-party machine learning algorithm.

10. The system of claim 6 , wherein the sequence of the plurality of words comprises at least one of: conversational text, a list, or instructions.

11. A computer program product comprising computer readable program code capable of being executed by one or more processors when retrieved from a non-transitory computer-readable medium, the program code comprising instructions configurable to cause:

receiving a data specification schema that specifies characteristics of test data objects comprising synthetic user data, wherein the characteristics comprise a third-party data source that includes a natural language text generation algorithm;

generating the test data objects, the generating of each test data object including:

determining, from the data specification schema, a number of fields of the test data object to be populated, the fields representing categories of simulated user data, and

determining values for the fields, the values simulating user data, wherein determining the values comprises:

determining values for a first field by querying the third-party data source such that the third-party data source returns, responsive to the query, realistic synthetic user data used to populate the values for the at least one field comprising a sequence of a plurality of words, wherein the third-party data source is associated with an entity other than an entity generating the test data objects,

determining values for a third field by generating a value associated with a second field and determining the value for the third field based on the generated value associated with the second field, wherein the data specification schema indicates that the third field is to have a value that is a function of the value associated with the second field, and

determining values for a fourth field by querying a second third-party data source that is a database, wherein the data specification schema indicates a query provided to the database, and wherein a response to the query corresponds to the value associated with the fourth field;

storing the test data objects in a database, the storing of each test data object including populating the fields with the determined values;

generating a tabular data file including the test data objects; and

transmitting the tabular data file to a computing system, wherein the computing system is configured to utilize the tabular data file comprising the realistic synthetic user data during a user data testing procedure of an application performed by the computing system.

12. The computer program product of claim 11 , wherein the data specification schema indicates a data type associated with each field of the number of fields.

13. The computer program product of claim 11 , wherein querying the third-party data source comprises requesting authentication to use the third-party data source using an authentication token specified in the data specification schema.

14. The computer program product of claim 11 , wherein the third-party data source comprises a third-party machine learning algorithm.

15. The computer program product of claim 11 , wherein the sequence of the plurality of words comprises at least one of:

conversational text, a list, or instructions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2022
From: WHITEHEAD, EMMET JAMES, JR; REED, LENA; KANO, AFSHIN MOBRAMAEIN
To: SAUCE LABS INC.
Reel/Frame 059899/0649 →
Continuity (1)
Related Publication 20230334071A1 · Oct 19, 2023
References Cited (43)
US 7584178B2 · Dettinger · 2009 [cited by examiner]
US 7912862B2 · Vaschillo · 2011 [cited by examiner]
US 8239374B2 · Narula · 2012 [cited by examiner]
US 8549480B2 · Cohen · 2013 [cited by examiner]
US 9286290B2 · Allen · 2016 [cited by examiner]
US 9836795B2 · Roberts · 2017 [cited by examiner]
US 10013679B1 · Lewis · 2018 [cited by examiner]
US 10311055B2 · Thombre · 2019 [cited by examiner]
US 11341033B2 · Mohankumar · 2022 [cited by examiner]
US 20020128993A1 · Jevons · 2002 [cited by examiner]
US 20060235835A1 · Dettinger · 2006 [cited by examiner]
US 20110276944A1 · Bergman · 2011 [cited by examiner]
US 20150066986A1 · Piecko · 2015 [cited by examiner]
US 20150082277A1 · Champlin-Scharff · 2015 [cited by examiner]
US 20180293115A1 · Skeem · 2018 [cited by examiner]
US 20180365316A1 · Liang · 2018 [cited by examiner]
US 20190163790A1 · Jennings · 2019 [cited by examiner]
US 20210073110A1 · Vidal · 2021 [cited by examiner]
US 20210193146A1 · Kirazci · 2021 [cited by examiner]
US 20210279165A1 · Rubin · 2021 [cited by examiner]
US 20220138004A1 · Nandakumar · 2022 [cited by examiner]
US 20230007894A1 · Bussa · 2023 [cited by examiner]
US 20230012264A1 · Venugopal · 2023 [cited by examiner]
US 20230073312A1 · Portisch · 2023 [cited by examiner]
US 20230139783A1 · Garib · 2023 [cited by examiner]
WO WO2010038019A2 · 2010 [cited by examiner]
WO WO2018156551A1 · 2018 [cited by examiner]
WO WO2019023509A1 · 2019 [cited by examiner]
WO WO2021086669A1 · 2021 [cited by examiner]
Iyad Alazzam et al., “Test Cases Selection Based on Source Code Features Extraction”, International Journal of Software Engineering and Its Applications, vol. 8, No. 1 (2014), pp. 203-214. [cited by examiner]
Dessislava Petrova-Antonova et al.,Automatic generation of test data for XML schema-based testing of web services , 2015 10th International Joint Conference on Software Technologies (ICSOFT), 2015, pp. 1-8. [cited by examiner]
Chennamsetty Madhusudhana Rao et al., “A Comparative Study of NLP based Semantic Web Standard model using SPARQL database”, 2021 International Conference on Computing Sciences (ICCS) (pp. 1-6). [cited by examiner]
J. Leopold et al., “A visual query system for the specification and scientific analysis of continual queries”, Proceedings IEEE Symposia on Human-Centric Computing Languages and Environments (Cat. No.01TH8587) (2001, pp… [cited by examiner]
Eliza Kuttner et al., “Processing natural language with schema constraint networks”, Computers & Mathematics with Applications, vol. 24, Issue 11, Dec. 1992, pp. 3-10. [cited by examiner]
“Announcing AI21 Studio and Jurassic-1 Language Models,” AI21 labs, 5 pages, URL: https://www.ai21.com/blog/announcing-ai21-studio-and-jurassic-1. [cited by applicant]
Cellat S., “Fine-Tuning Transformer-Based Language Models,” Y Meadows, 2021, 8 pages. [cited by applicant]
Garbade M. J., “Understanding few-shot learning in machine learning,” Medium, 2018, 5 pages. [cited by applicant]
“GPT-Neo,” EleutherAI, 2 pages, URL: https://www.eleuther.ai/projects/gpt-neo/. [cited by applicant]
Kulshrestha R., “Transformers,” Transformers in NLP: A beginner friendly explanation, Towards Data Science, 2020, 14 pages. [cited by applicant]
Lieber, et al., “Jurassic-1: Technical Details and Evaluation,” White Paper, 9 pages. [cited by applicant]
Osborne S., “Learning NLP Languages Models with Real Data,” Towards Data Science, 2019, 24 pages, URL: https://towardsdatascience.com/learning-nlp-language-models-with-real-data-cdff04c51c25. [cited by applicant]
“Prompt Engineering Tips and Tricks with GPT-3,” andrew makes things, Apr. 21, 2021, 6 pages. [cited by applicant]
“Zero-Shot Learning in Modern NLP,” State-of-the-art NLP models for text classification without annotated data, Joe Davison Blog, May 29, 2020, 14 pages. [cited by applicant]
Cited By (2)
US 12,493,615 US 12,608,370