IP Library Granted Patent US 12,510,888
Granted Patent B2
US 12,510,888 · App. 17/676,629 · Granted Dec 30, 2025

Model reduction and training efficiency in computer-based reasoning and artificial intelligence systems

Inventor: Christopher James Hazard (Raleigh, NC)
Assignee: Howso Incorporated
G05B23/0281G06F18/214G06F18/22G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,510,888
App. No.
17/676,629
Granted
Dec 30, 2025
Kind
B2
Abstract

Techniques are provided herein for creating well-balanced computer-based reasoning systems and using those to control systems. The techniques include receiving a request to determine whether to use one or more particular data elements, features, cases, etc. in a computer-based reasoning model (e.g., as data elements, cases or features are being added, or as part of pruning existing features or cases). Conviction measures are determined and inclusivity conditions are tested. The result of comparing the conviction measure can be used to determine whether to include or exclude the feature, case, etc. in the model and/or whether there are anomalies in the model. A controllable system may then be controlled using the computer-based reasoning model. Examples controllable systems include self-driving cars, image labeling systems, manufacturing and assembly controls, federated systems, smart voice controls, automated control of experiments, energy transfer systems, health care systems, cybersecurity systems, and the like.

Claims (94)

1 . A method comprising:

training a computer-based reasoning model, wherein the computer-based reasoning model includes a plurality of data elements, wherein each data element comprises context data paired with action data, wherein the action data is descriptive of an action taken in response to a context described by the context data;

receiving a request to determine whether one or more particular data elements of the plurality of data elements included in the computer-based reasoning model meet inclusivity conditions;

determining one or more conviction scores for the one or more particular data elements,

wherein determining the one or more conviction scores for the one or more particular data elements comprises determining an excluding-type surprisal score for the one or more particular data elements and determining a including-type surprisal score for the one or more particular data elements;

wherein:

the excluding-type surprisal score is calculated based on a first probability density or mass functions (PDMF) for a first set of data elements associated with the computer-based reasoning model where the one or more particular data elements are excluded from the first set of data elements, and

the including-type surprisal score is calculated based on a second PDMF for a second set of data elements associated with the computer-based reasoning model where the one or more particular data elements are included in the second set of data elements;

determining whether the one or more conviction scores meet one or more inclusivity conditions;

in response to determining that the one or more conviction scores meet the one or more inclusivity conditions:

including the one or more particular data elements in the computer-based reasoning model when the inclusivity conditions comprise an inclusion condition; and

excluding the one or more particular data elements in the computer-based reasoning model when the inclusivity conditions comprise an exclusion condition,

wherein determining whether the one or more conviction scores meet the inclusivity conditions comprises determining that the one or more particular data elements meet the inclusion condition when a difference between the excluding-type surprisal score and the including-type surprisal score is beyond a threshold; and

causing control of a controllable system with the computer-based reasoning model, wherein causing control of the controllable system with the computer-based reasoning model comprises receiving a current context for the controllable system and determining a control action to perform based on the current context and at least one of the plurality of data elements of the computer-based reasoning model;

wherein the method is performed on one or more computing devices.

2 . The method of claim 1 , wherein determining that the one or more particular data elements meet the inclusion condition when the difference between the excluding-type surprisal score and the including-type surprisal score is beyond the threshold comprises determining that the difference between the excluding-type surprisal score and the including-type surprisal score is above the threshold.

3 . The method of claim 1 , wherein determining that the one or more particular data elements meet the inclusion condition when the difference between the excluding-type surprisal score and the including-type surprisal score is beyond the threshold comprises determining that the difference between the excluding-type surprisal score and the including-type surprisal score is below the threshold.

4 . The method of claim 1 ,

wherein receiving the request comprises receiving a request to reduce the computer-based reasoning model to a particular size;

and the method further comprises:

determining a number of data elements to exclude in the computer-based reasoning model to reduce the computer-based reasoning model to the particular size;

determining a subset of data elements to exclude in the computer-based reasoning model based at least in part on the one or more conviction scores for data elements in the computer-based reasoning model; and

excluding the subset of data elements from the computer-based reasoning model to reduce the size of the computer-based reasoning model to the particular size.

5 . The method of claim 1 , further comprising:

initially receiving the one or more particular data elements as part of training for the computer-based reasoning model;

in response to determining that the one or more conviction scores meet the inclusion condition, sending an indication to a trainer associated with the training for the computer-based reasoning model to continue to train related to the one or more particular data elements;

in response to determining that the one or more conviction scores meet the exclusion condition, sending the indication to the trainer associated with the training for the computer-based reasoning model that training is no longer needed related to the one or more particular data elements.

6 . The method of claim 1 , wherein causing control of the controllable system comprises:

receiving a request for the control action to perform in the current context;

determining the control action to perform based on comparing the current context to the context data associated with multiple data elements included in the computer-based reasoning model; and

responding to the request for the action to take with the determined control action.

7 . The method of claim 6 , further comprising:

receiving an indication that there was an anomaly associated with the determined control action;

removing one or more data elements associated with the determined control action from the computer-based reasoning model.

8 . The method of claim 1 , further comprising:

continuing to determine the one or more conviction scores for new data elements and including or excluding those data elements based on whether the one or more conviction scores meet the inclusivity conditions until a termination condition for inclusion or exclusion is met.

9 . A system for executing instructions, wherein said instructions are instructions which, when executed by one or more computing devices, cause performance of a process including:

training a computer-based reasoning model, wherein the computer-based reasoning model includes a plurality of data elements, wherein each data element comprises context data paired with action data, wherein the action data is descriptive of an action taken in response to a context described by the context data;

receiving a request to determine whether one or more particular data elements of the plurality of data elements in the computer-based reasoning model meet inclusivity conditions;

determining one or more conviction scores for the one or more particular data elements,

wherein determining the one or more conviction scores for the one or more particular data elements comprises determining a excluding-type surprisal score for the one or more particular data elements and determining a including-type surprisal score for the one or more particular data elements;

wherein:

the excluding-type surprisal score is calculated based on a first probability density or mass functions (PDMF) for a first set of data elements associated with the computer-based reasoning model where the one or more particular data elements are excluded from the first set of data elements, and

the including-type surprisal score is calculated based on a second PDMF for a second set of data elements associated with the computer-based reasoning model where the one or more particular data elements are included in the second set of data elements;

determining whether the one or more conviction scores meet one or more inclusivity conditions;

in response to determining that the one or more conviction scores meet the one or more inclusivity conditions:

including the one or more particular data elements in the computer-based reasoning model when the inclusivity conditions comprise an inclusion condition;

excluding the one or more particular data elements in the computer-based reasoning model when the inclusivity conditions comprise an exclusion condition,

wherein determining whether the one or more conviction scores meet the inclusivity conditions comprises determining that the one or more particular data elements meet the inclusion condition when a difference between the excluding-type surprisal score and the including-type surprisal score is beyond a threshold;

causing control of a controllable system with the computer-based reasoning model, wherein causing control of the controllable system with the computer-based reasoning model comprises receiving a current context for the controllable system and determining a control action to perform based on the current context and at least one of the plurality of data elements of the computer-based reasoning model;

wherein the process is performed on one or more computing devices.

10 . The system of claim 9 , wherein determining that the one or more particular data elements meet the inclusion condition when the difference between the excluding-type surprisal score and the including-type surprisal score is beyond the threshold comprises determining that the difference between the excluding-type surprisal score and the including-type surprisal score is above the threshold.

11 . The system of claim 9 , wherein determining that the one or more particular data elements meet the inclusion condition when the difference between the excluding-type surprisal score and the including-type surprisal score is beyond the threshold comprises determining that the difference between the excluding-type surprisal score and the including-type surprisal score is below the threshold.

12 . The system of claim 9 ,

wherein receiving the request comprises receiving a request to reduce the computer-based reasoning model to a particular size;

and the process further comprises:

determining a number of data elements to exclude in the computer-based reasoning model to reduce the computer-based reasoning model to the particular size;

determining a subset of data elements to exclude in the computer-based reasoning model based at least in part on the one or more conviction scores for data elements in the computer-based reasoning model; and

excluding the subset of data elements from the computer-based reasoning model to reduce the size of the computer-based reasoning model to the particular size.

13 . The system of claim 9 , the process further comprising:

initially receiving the one or more particular data elements as part of training for the computer-based reasoning model;

in response to determining that the one or more conviction scores meet the inclusion condition, sending an indication to a trainer associated with the training for the computer-based reasoning model to continue to train related to the one or more particular data elements;

in response to determining that the one or more conviction scores meet the exclusion condition, sending the indication to the trainer associated with the training for the computer-based reasoning model that training is no longer needed related to the one or more particular data elements.

14 . The system of claim 9 , wherein causing control of the controllable system comprises:

receiving a request for the control action to perform in the current context;

determining the control action to perform based on comparing the current context to the context data associated with multiple data elements included in the computer-based reasoning model; and

responding to the request for the control action to perform with the determined control action.

15 . The system of claim 14 , the process further comprising:

receiving an indication that there was an anomaly associated with the determined control action;

removing one or more data elements associated with the determined control action from the computer-based reasoning model.

16 . A non-transitory computer readable medium storing instructions which, when executed by one or more computing devices, cause the one or more computing devices to perform a process of:

training a computer-based reasoning model, wherein the computer-based reasoning model includes a plurality of data features, wherein each data feature comprises context data paired with action data, wherein the action data is descriptive of an action taken in response to a context described by the context data;

receiving a request to determine whether one or more particular data features in the computer-based reasoning model meet inclusivity conditions;

determining one or more conviction scores for the one or more particular data features,

wherein determining the one or more conviction scores for the one or more particular data features comprises determining a excluding-type surprisal score for the one or more particular data features and determining a including-type surprisal score for the one or more particular data features;

wherein:

the excluding-type surprisal score is calculated based on a first probability density or mass functions (PDMF) for a first set of data features associated with the computer-based reasoning model where the one or more particular data features are excluded from the first set of data features, and

the including-type surprisal score is calculated based on a second PDMF for a second set of data features associated with the computer-based reasoning model where the one or more particular data features are included in the second set of data features;

determining whether the one or more conviction scores meet one or more inclusivity conditions;

in response to determining that the one or more conviction scores meet the one or more inclusivity conditions:

including the one or more particular data features in the computer-based reasoning model when the inclusivity conditions comprise an inclusion condition;

excluding the one or more particular data features in the computer-based reasoning model when the inclusivity conditions comprise an exclusion condition,

wherein determining whether the one or more conviction scores meet the inclusivity conditions comprises determining that the one or more particular data features meet the inclusion condition when a difference between the excluding-type surprisal score and the including-type surprisal score is beyond a threshold;

causing control of a controllable system with the computer-based reasoning model, wherein causing control of the controllable system with the computer-based reasoning model comprises receiving a current context for the controllable system and determining a control action to perform based on the current context and at least one of the plurality of data features of the computer-based reasoning model.

17 . The non-transitory computer readable medium of claim 16 , wherein determining that the one or more particular data features meet the inclusion condition when the difference between the excluding-type surprisal score and the including-type surprisal score is beyond the threshold comprises determining that the difference between the excluding-type surprisal score and the including-type surprisal score is above the threshold.

18 . The non-transitory computer readable medium of claim 16 , wherein determining that the one or more particular data features meet the inclusion condition when the difference between the excluding-type surprisal score and the including-type surprisal score is beyond the threshold comprises determining that the difference between the excluding-type surprisal score and the including-type surprisal score is below the threshold.

19 . The non-transitory computer readable medium of claim 16 , further comprising:

initially receiving the one or more particular data features as part of training for the computer-based reasoning model;

in response to determining that the one or more conviction scores meet the inclusion condition, sending an indication to a trainer associated with the training for the computer-based reasoning model to continue to train related to the one or more particular data features;

in response to determining that the one or more conviction scores meet the exclusion condition, sending the indication to the trainer associated with the training for the computer-based reasoning model that training is no longer needed related to the one or more particular data features.

20 . The non-transitory computer readable medium of claim 16 , wherein causing control of the controllable system comprises:

receiving a request for the control action to perform in the current context;

determining the control action to perform based on comparing the current context to the context data associated with multiple data features included in the computer-based reasoning model; and

responding to the request for the control action to perform with the determined control action.

Assignments (5)
TERMINATION AND RELEASE OF INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jan 22, 2025
From: WESTERN ALLIANCE BANK
To: HOWSO INCORPORATED
Reel/Frame 069988/0038 →
CHANGE OF NAME Recorded Sep 28, 2023
From: DIVEPLANE CORPORATION
To: HOWSO INCORPORATED
Reel/Frame 065081/0559 →
CHANGE OF NAME Recorded Sep 22, 2023
From: DIVEPLANE CORPORATION
To: HOWSO INCORPORATED
Reel/Frame 065021/0691 →
SECURITY INTEREST Recorded Jan 31, 2023
From: DIVEPLANE CORPORATION
To: WESTERN ALLIANCE BANK
Reel/Frame 062554/0106 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2022
From: HAZARD, CHRISTOPHER JAMES
To: DIVEPLANE CORPORATION
Reel/Frame 059075/0351 →
Continuity (7)
Continuation 16992842 · Aug 13, 2020
Continuation 16992876 · Aug 13, 2020
Continuation In Part 16376509 · Apr 5, 2019
Continuation In Part 16220986 · Dec 14, 2018
Continuation In Part 15948805 · Apr 9, 2018
Provisional Application 63038335 · Jun 12, 2020
Related Publication 20220179408A1 · Jun 9, 2022
References Cited (167)
US 4935877A · Koza · 1990 [cited by applicant]
US 5581664A · Allen et al. · 1996 [cited by applicant]
US 6282527B1 · Gounares et al. · 2001 [cited by applicant]
US 6741972B1 · Girardi et al. · 2004 [cited by applicant]
US 7873587B2 · Baum · 2011 [cited by applicant]
US 9489635B1 · Zhu · 2016 [cited by applicant]
US 9858496B2 · Sun et al. · 2018 [cited by applicant]
US 9922286B1 · Hazard · 2018 [cited by applicant]
US 10158658B1 · Sharifi Mehr · 2018 [cited by applicant]
US 10459444B1 · Kentley-Klay · 2019 [cited by applicant]
US 10816980B2 · Hazard et al. · 2020 [cited by applicant]
US 10816981B2 · Hazard et al. · 2020 [cited by applicant]
US 10817750B2 · Hazard et al. · 2020 [cited by applicant]
US 20010049595A1 · Plumer et al. · 2001 [cited by applicant]
US 20040019851A1 · Purvis et al. · 2004 [cited by applicant]
US 20050137992A1 · Polak · 2005 [cited by applicant]
US 20060195204A1 · Bonabeau et al. · 2006 [cited by applicant]
US 20080153098A1 · Rimm et al. · 2008 [cited by applicant]
US 20080307399A1 · Zhou et al. · 2008 [cited by applicant]
US 20090006299A1 · Baum · 2009 [cited by applicant]
US 20090144704A1 · Niggemann et al. · 2009 [cited by applicant]
US 20100106603A1 · Dey et al. · 2010 [cited by applicant]
US 20100287507A1 · Paquette et al. · 2010 [cited by applicant]
US 20110060895A1 · Solomon · 2011 [cited by applicant]
US 20110161264A1 · Cantin · 2011 [cited by applicant]
US 20110225564A1 · Biswas et al. · 2011 [cited by applicant]
US 20130006901A1 · Cantin · 2013 [cited by applicant]
US 20130339365A1 · Balasubramanian et al. · 2013 [cited by applicant]
US 20140324339A1 · Adam et al. · 2014 [cited by applicant]
US 20150058982A1 · Eskin et al. · 2015 [cited by applicant]
US 20160055427A1 · Adjaoute · 2016 [cited by applicant]
US 20170010106A1 · Shashua et al. · 2017 [cited by applicant]
US 20170012772A1 · Mueller · 2017 [cited by applicant]
US 20170053211A1 · Heo et al. · 2017 [cited by applicant]
US 20170091645A1 · Matus · 2017 [cited by applicant]
US 20170161640A1 · Shamir · 2017 [cited by applicant]
US 20170236060A1 · Ignatyev · 2017 [cited by applicant]
US 20180018590A1 · Szeto et al. · 2018 [cited by applicant]
US 20180072323A1 · Gordon et al. · 2018 [cited by applicant]
US 20180089563A1 · Redding et al. · 2018 [cited by applicant]
US 20180235649A1 · Elkadi · 2018 [cited by applicant]
US 20180336018A1 · Lu et al. · 2018 [cited by applicant]
US 20190101924A1 · Styler et al. · 2019 [cited by applicant]
US 20190147331A1 · Arditi · 2019 [cited by applicant]
US 20200005467A1 · Yamada et al. · 2020 [cited by applicant]
US 20200371512A1 · Srinivasamurthy et al. · 2020 [cited by applicant]
WO WO2017057528 · 2017 [cited by applicant]
WO WO2017189859 · 2017 [cited by applicant]
Abdi, “Cardinality Optimization Problems”, The University of Birmingham, PhD Dissertation, May 2013, 197 pages. [cited by applicant]
Aboulnaga, “Generating Synthetic Complex-structured XML Data”, Proceedings of the Fourth International Workshop on the Web and Databases, WebDB 2001, Santa Barbara, California, USA, May 24-25, 2001, 6 pages. [cited by applicant]
Abramson, “The Expected-Outcome Model of Two-Player Games”, PhD Thesis, Columbia University, New York, New York, 1987, 125 pages. [cited by applicant]
Abuelaish et al., “Analysis and Modelling of Groundwater Salinity Dynamics in the Gaza Strip”, Cuadernos Geograficos, vol. 57, Issue 2, pp. 72-91. [cited by applicant]
Agarwal et al., “Nearest-Neighbor Searching Under Uncertainty II”, ACM Transactions on Algorithms, vol. 13, Issue 1, Article 3, 2016, 25 pages. [cited by applicant]
Aggarwal et al., “On the Surprising Behavior of Distance Metrics in High Dimensional Space”, International Conference on Database Theory, London, United Kingdom, Jan. 4-6, 2001, pp. 420-434. [cited by applicant]
Akaike, “Information Theory and an Extension of the Maximum Likelihood Principle”, Proceedings of the 2nd International Symposium on Information Theory, Sep. 2-8, 1971, Tsahkadsor, Armenia, pp. 267-281. [cited by applicant]
Alhaija, “Augmented Reality Meets Computer Vision: Efficient Data Generation for Urban Driving Scenes”, arXiv:1708.01566v1, Aug. 4, 2017, 12 pages. [cited by applicant]
Alpaydin, “Machine Learning: The New AI”, MIT Press, Cambridge, Massachusetts, 2016, 225 pages. [cited by applicant]
Alpaydin, “Voting Over Multiple Condensed Nearest Neighbor”, Artificial Intelligence Review, vol. 11, 1997, pp. 115-132. [cited by applicant]
Altman, “An Introduction to Kernel and Nearest-Neighbor Nonparametric Regression”, The American Statistician, vol. 46, Issue 3, 1992, pp. 175-185. [cited by applicant]
Anderson, “Synthetic data generation for the internet of things,” 2014 IEEE International Conference on Big Data (Big Data), Washington, DC, 2014, pp. 171-176. [cited by applicant]
Archer et al., “Empirical Characterization of Random Forest Variable Importance Measures”, Computational Statistics & Data Analysis, vol. 52, 2008, pp. 2249-2260. [cited by applicant]
Beura, “Development of Features and Feature Reduction Techniques for Mammogram Classification,” Department of Computer Science and Engineering of National Institute of Technology Rourkela, Jun. 2016, 164 pages. [cited by applicant]
Beyer et al., “When is 'Nearest Neighbor' Meaningful?” International Conference on Database Theory, Springer, Jan. 10-12, 1999, Jerusalem, Israel, pp. 217-235. [cited by applicant]
Bull, “Haploid-Diploid Evolutionary Algorithms: The Baldwin Effect and Recombination Nature's Way”, Artificial Intelligence and Simulation of Behaviour Convention, Apr. 19-21, 2017, Bath, United Kingdom, pp. 91-94. [cited by applicant]
Cano et al., “Evolutionary Stratified Training Set Selection for Extracting Classification Rules with Tradeoff Precision-Interpretability” Data and Knowledge Engineering, vol. 60, 2007, pp. 90-108. [cited by applicant]
Chawla, “SMOTE: Synthetic Minority Over-sampling Technique”, arXiv:1106.1813v1, Jun. 9, 2011, 37 pages. [cited by applicant]
Chen, “DropoutSeer: Visualizing Learning Patterns in Massive Open Online Courses for Dropout Reasoning and Prediction”, 2016 IEEE Conference on Visual Analytics Science and Technology (VAST), Oct. 23-28, 2016, Baltimore… [cited by applicant]
Chomboon et al., An Empirical Study of Distance Metrics for k-Nearest Neighbor Algorithm, 3rd International Conference on Industrial Application Engineering, Kitakyushu, Japan, Mar. 28-31, 2015, pp. 280-285. [cited by applicant]
Colakoglu, “A Generalization of the Minkowski Distance and a New Definition of the Ellipse”, arXiv:1903.09657v1, Mar. 2, 2019, 18 pages. [cited by applicant]
Dernoncourt, “Mooc Viz: A Large Scale, Open Access, Collaborative, Data Analytics Platform for MOOCs”, NIPS 2013 Education Workshop, Nov. 1, 2013, Lake Tahoe, Utah, USA, 8 pages. [cited by applicant]
Ding, “Generating Synthetic Data for Neural Keyword-to-Question Models”, arXiv:1807.05324v1, Jul. 18, 2018, 12 pages. [cited by applicant]
Dwork et al., “The Algorithmic Foundations of Differential Privacy”, Foundations and Trends in Theoretical Computer Science, vol. 9, Nos. 3-4, 2014, pp. 211-407. [cited by applicant]
Efros et al., “Texture Synthesis by Non-Parametric Sampling”, International Conference on Computer Vision, Sep. 20-25, 1999, Corfu, Greece, 6 pages. [cited by applicant]
Esener et al., “A New Feature Ensemble with a Multistage Classification Scheme for Breast Cancer Diagnosis,” Journal of Healthcare Engineering, 2017, 15 pages. [cited by applicant]
Fathony, “Discrete Wasserstein Generative Adversarial Networks (DWGAN)”, OpenReview.net, Feb. 18, 2018, 20 pages. [cited by applicant]
Ganegedara et al., “Self Organising Map Based Region of Interest Labelling for Automated Defect Identification in Large Sewer Pipe Image Collections”, IEEE World Congress on Computational Intelligence, Jun. 10-15, 2012,… [cited by applicant]
Gao et al., “Efficient Estimation of Mutual Information for Strongly Dependent Variables”, 18th International Conference on Artificial Intelligence and Statistics, San Diego, California, May 9-12, 2015, pp. 277-286. [cited by applicant]
Gehr et al., “AI [cited by applicant]
Gemmeke et al., “Using Sparse Representations for Missing Data Imputation in Noise Robust Speech Recognition”, European Signal Processing Conference, Lausanne, Switzerland, Aug. 25-29, 2008, 5 pages. [cited by applicant]
Gholamreza, “Face Recognition Using Color Local Binary Pattern from Mutually Independent Color Channels,” Anbarjafari EURASIP Journal on Image and Video Processing, 2013, 11 pages. [cited by applicant]
Ghosh, “Inferential Privacy Guarantees for Differentially Private Mechanisms”, arXiv:1603:.01508v1, Mar. 4, 2016, 31 pages. [cited by applicant]
Goodfellow et al., “Deep Learning”, 2016, 800 pages. [cited by applicant]
Google AI Blog, “The What-If Tool: Code-Free Probing of Machine Learning Models”, Sep. 11, 2018, https://pair-code,github.io/what-if-tool, retrieved on Mar. 14, 2019, 5 pages. [cited by applicant]
Gottlieb et al., “Near-Optimal Sample Compression for Nearest Neighbors”, Advances in Neural Information Processing Systems, Montreal, Canada, Dec. 8-13, 2014, 9 pages. [cited by applicant]
Gray, “Quickly Generating Billion-Record Synthetic Databases”, SIGMOD '94: Proceedings of the 1994 ACM SIGMOD international conference on Management of data, May 1994, Minneapolis, Minnesota, USA, 29 pages. [cited by applicant]
Hastie et al., “The Elements of Statistical Learning”, 2001, 764 pages. [cited by applicant]
Hazard et al., “Natively Interpretable Machine Learning and Artificial Intelligence: Preliminary Results and Future Directions”, arXiv:1901v1, Jan. 2, 2019, 15 pages. [cited by applicant]
Hinneburg et al., “What is the Nearest Neighbor in High Dimensional Spaces?”, 26th International Conference on Very Large Databases, Cairo, Egypt, Sep. 10-14, 2000, pp. 506-515. [cited by applicant]
Hmeidi et al., “Performance of KNN and SVM Classifiers on Full Word Arabic Articles”, Advanced Engineering Informatics, vol. 22, Issue 1, 2008, pp. 106-111. [cited by applicant]
Hoag “A Parallel General-Purpose Synthetic Data Generator”, ACM SIGMOID Record, vol. 6, Issue 1, Mar. 2007, 6 pages. [cited by applicant]
Houle et al., “Can Shared-Neighbor Distances Defeat the Curse of Dimensionality?”, International Conference on Scientific and Statistical Database Management, Heidelberg, Germany, Jun. 31-Jul. 2, 2010, 18 pages. [cited by applicant]
Imbalanced-Learn, “SMOTE”, 2016-2017, https://imbalanced-learn.readthedocs.io/en/stable/generated/imblearn.over_sampling.SMOTE.html, retrieved on Aug. 11, 2020, 6 pages. [cited by applicant]
Indyk et al., “Approximate Nearest Neighbors: Towards Removing the Curse of Dimensionality”, Procedures of the 30th ACM Symposium on Theory of Computing, Dallas, Texas, May 23-26, 1998, pp. 604-613. [cited by applicant]
International Search Report and Written Opinion for PCT/US2018/047118, mailed on Dec. 3, 2018, 9 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/US2019/026502, mailed on Jul. 24, 2019, 16 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/US2019/066321, mailed on Mar. 19, 2020, 17 pages. [cited by applicant]
Internet Archive, “System Verilog distribution Constraint—Verification Guide”, Aug. 6, 2018, http://web.archive.org/web/20180806225430/https://www.verificationguide.com/p/systemverilog-distribution-constraint.html, retr… [cited by applicant]
Internet Archive, “SystemVerilog Testbench Automation Tutorial”, Nov. 17, 2016, https://web.archive.org/web/20161117153225/http://www.doulos.com/knowhow/sysverilog/tutorial/constraints/, retrieved on Mar. 3, 2020, 6 pag… [cited by applicant]
Kittler, “Feature Selection and Extraction”, Handbook of Pattern Recognition and Image Processing, Jan. 1986, Chapter 3, pp. 115-132. [cited by applicant]
Kohavi et al., “Wrappers for Feature Subset Selection”, Artificial Intelligence, vol. 97, Issues 1-2, Dec. 1997, pp. 273-323. [cited by applicant]
Kontorovich et al., “Nearest-Neighbor Sample Compression: Efficiency, Consistency, Infinite Dimensions”, Advances in Neural Information Processing Systems, 2017, pp. 1573-1583. [cited by applicant]
Kulkarni et al., “Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation”, arXiv:1604.06057v2, May 31, 2016, 14 pages. [cited by applicant]
Kuramochi et al., “Gene Classification using Expression Profiles: A Feasibility Study”, Technical Report TR 01-029, Department of Computer Science and Engineering, University of Minnesota, Jul. 23, 2001, 18 pages. [cited by applicant]
Kushilevitz et al., “Efficient Search for Approximate Nearest Neighbor in High Dimensional Spaces”, Society for Industrial and Applied Mathematics Journal Computing, vol. 30, No. 2, pp. 457-474. [cited by applicant]
Leinster et al., “Maximizing Diversity in Biology and Beyond”, Entropy, vol. 18, Issue 3, 2016, 23 pages. [cited by applicant]
Liao et al., “Similarity Measures for Retrieval in Case-Based Reasoning Systems”, Applied Artificial Intelligence, vol. 12, 1998, pp. 267-288. [cited by applicant]
Lin et al., “Why Does Deep and Cheap Learning Work So Well?” Journal of Statistical Physics, vol. 168, 2017, pp. 1223-1247. [cited by applicant]
Lin, “Development of a Synthetic Data Set Generator for Building and Testing Information Discovery Systems”, Proceedings of the Third International Conference on Information Technology: New Generations, Nevada, USA, Apr… [cited by applicant]
Lukaszyk, “A New Concept of Probability Metric and its Applications in Approximation of Scattered Data Sets”, Computational Mechanics, vol. 33, 2004, pp. 299-304. [cited by applicant]
Lukaszyk, “Probability Metric, Examples of Approximation Applications in Experimental Mechanics”, PhD Thesis, Cracow University of Technology, 2003, 149 pages. [cited by applicant]
Maitre et al., “Wavelet-Based Joint Estimation and Encoding of Depth-Image-Based Representations for Free-Viewpoint Rendering,” IEEE Transactions on Image Processing, vol. 17, No. 6, Jun. 2008, pp. 946-957. [cited by applicant]
Mann et al., “On a Test of Whether One or Two Random Variables is Stochastically Larger than the Other”, The Annals of Mathematical Statistics, 1947, pp. 50-60. [cited by applicant]
Martino et al., “A Fast Universal Self-Tuned Sampler within Gibbs Sampling”, Digital Signal Processing, vol. 47, 2015, pp. 68-83. [cited by applicant]
Mohri et al., Foundations of Machine Learning, 2012, 427 pages—uploaded as Part 1 and Part 2. [cited by applicant]
Montanez, “SDV: An Open Source Library for Synthetic Data Generation”, Massachusetts Institute of Technology, Master's Thesis, Sep. 2018, 105 pages. [cited by applicant]
Negra, “Model of a Synthetic Wind Speed Time Series Generator”, Wind Energy, Wiley Interscience, Sep. 6, 2007, 17 pages. [cited by applicant]
Nguyen et al., “NP-Hardness of { 0 Minimization Problems: Revision and Extension to the Non-Negative Setting”, 13th International Conference on Sampling Theory and Applications, Jul. 8-12, 2019, Bordeaux, France, 4 page… [cited by applicant]
Olson et al., “PMLB: A Large Benchmark Suite for Machine Learning Evaluation and Comparison”, arXiv:1703.00512v1, Mar. 1, 2017, 14 pages. [cited by applicant]
Patki, “The Synthetic Data Vault: Generative Modeling for Relational Databases”, Massachusetts Institute of Technology, Master's Thesis, Jun. 2016, 80 pages. [cited by applicant]
Patki, “The Synthetic Data Vault”, 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA), Oct. 17-19, 2016, Montreal, QC, pp. 399-410. [cited by applicant]
Pedregosa et al., “Machine Learning in Python”, Journal of Machine Learning Research, vol. 12, 2011, pp. 2825-2830. [cited by applicant]
Pei, “A Synthetic Data Generator for Clustering and Outlier Analysis”, The University of Alberta, 2006, 33 pages. [cited by applicant]
Phan et al., “Adaptive Laplace Mechanism: Differential Privacy Preservation in Deep Learning” 2017 IEEE International Conference on Data Mining, New Orleans, Louisiana, Nov. 18-21, 2017, 10 pages. [cited by applicant]
Poerner et al., “Evaluating Neural Network Explanation Methods Using Hybrid Documents and Morphosyntactic Agreement”, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Long Papers)… [cited by applicant]
Prakosa “Generation of Synthetic but Visually Realistic Time Series of Cardiac Images Combining a Biophysical Model and Clinical Images” IEEE Transactions on Medical Imaging, vol. 32, No. 1, Jan. 2013, pp. 99-109. [cited by applicant]
Priyardarshini, “WEDAGEN: A synthetic web database generator”, International Workshop of Internet Data Management (IDM'99), Sep. 2, 1999, Florence, IT, 24 pages. [cited by applicant]
Pudjijono, “Accurate Synthetic Generation of Realistic Personal Information” Advances in Knowledge Discovery and Data Mining, Pacific-Asia Conference on Knowledge Discovery and Data Mining, Apr. 27-30, 2009, Bangkok, Th… [cited by applicant]
Raikwal et al., “Performance Evaluation of SVM and K-Nearest Neighbor Algorithm Over Medical Data Set”, International Journal of Computer Applications, vol. 50, No. 14, Jul. 2012, pp. 975-985. [cited by applicant]
Rao et al., “Cumulative Residual Entropy: A New Measure of Information”, IEEE Transactions on Information Theory, vol. 50, Issue 6, 2004, pp. 1220-1228. [cited by applicant]
Reiter, “Using CART to Generate Partially Synthetic Public Use Microdata” Journal of Official Statistics, vol. 21, No. 3, 2005, pp. 441-462. [cited by applicant]
Ribeiro et al., “‘Why Should I Trust You’: Explaining the Predictions of Any Classifier”, arXiv:1602.04938v3, Aug. 9, 2016, 10 pages. [cited by applicant]
Rosenberg et al., “Semi-Supervised Self-Training of Object Detection Models”, IEEE Workshop on Applications of Computer Vision, 2005, 9 pages. [cited by applicant]
Schaul et al., “Universal Value Function Approximators”, International Conference on Machine Learning, Lille, France, Jul. 6-11, 2015, 9 pages. [cited by applicant]
Schlabach et al., “FOX-GA: A Genetic Algorithm for Generating and ANalayzing Battlefield Courses of Action”, 1999 MIT. [cited by applicant]
Schreck, “Towards An Automatic Predictive Question Formulation”, Massachusetts Institute of Technology, Master's Thesis, Jun. 2016, 121 pages. [cited by applicant]
Schreck, “What would a data scientist ask? Automatically formulating and solving prediction problems”, 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA), Oct. 17-19, 2016, Montreal, QC, pp… [cited by applicant]
Schuh et al., “Improving the Performance of High-Dimensional KNN Retrieval Through Localized Dataspace Segmentation and Hybrid Indexing”, Proceedings of the 17th East European Conference, Advances in Databases and Infor… [cited by applicant]
Schuh et al., “Mitigating the Curse of Dimensionality for Exact KNN Retrieval”, Proceedings of the 26th International Florida Artificial Intelligence Research Society Conference, St. Pete Beach, Florida, May 22-24, 2014… [cited by applicant]
Schwarz et al., “Estimating the Dimension of a Model”, The Annals of Statistics, vol. 6, Issue 2, Mar. 1978, pp. 461-464. [cited by applicant]
Silver et al., “Mastering the Game of Go Without Human Knowledge”, Nature, vol. 550, Oct. 19, 2017, pp. 354-359. [cited by applicant]
Skapura, “Building Neural Networks”, 1996, p. 63. [cited by applicant]
Smith, “FeatureHub: Towards collaborative data science”, 2017 IEEE International Conference on Data Science and Advanced Analytics (DSAA), Oct. 19-21, 2017, Tokyo, pp. 590-600. [cited by applicant]
Stephenson et al., “A Continuous Information Gain Measure to Find the Most Discriminatory Problems for AI Benchmarking”, arxiv.org, arxiv.org/abs/1809.02904v2, retrieved on Aug. 21, 2019, 8 pages. [cited by applicant]
Stoppiglia et al., “Ranking a Random Feature for Variable and Feature Selection” Journal of Machine Learning Research, vol. 3, 2003, pp. 1399-1414. [cited by applicant]
Sun et al., “Fuzzy Modeling Employing Fuzzy Polyploidy Genetic Algorithms”, Journal of Information Science and Engineering, Mar. 2002, vol. 18, No. 2, pp. 163-186. [cited by applicant]
Sun, “Learning Vine Copula Models For Synthetic Data Generation”, arXiv:1812.01226v1, Dec. 4, 2018, 9 pages. [cited by applicant]
Surya et al., “Distance and Similarity Measures Effect on the Performance of K-Nearest Neighbor Classifier”, arXiv:1708.04321v1, Aug. 14, 2017, 50 pages. [cited by applicant]
Tan et al., “Incomplete Multi-View Weak-Label Learning”, 27th International Joint Conference on Artificial Intelligence, 2018, pp. 2703-2709. [cited by applicant]
Tao et al., “Quality and Efficiency in High Dimensional Nearest Neighbor Search”, Proceedings of the 2009 ACM SIGMOD International Conference on Management of Data, Providence, Rhode Island, Jun. 29-Jul. 2, 2009, pp. 56… [cited by applicant]
Tishby et al., “Deep Learning and the Information Bottleneck Principle”, arXiv:1503.02406v1, Mar. 9, 2015, 5 pages. [cited by applicant]
Tockar, “Differential Privacy: The Basics”, Sep. 8, 2014, https://research.neustar.biz/2014/09/08/differential-privacy-the-basics/ retrieved on Apr. 1, 2019, 3 pages. [cited by applicant]
Tomasev et al., “Hubness-aware Shared Neighbor Distances for High-Dimensional k-Nearest Neighbor Classification”, 7th international Conference on Hybrid Artificial Intelligent Systems, Salamanca, Spain, Mar. 28-30, 2012… [cited by applicant]
Tran, “Dist-GAN: An Improved GAN Using Distance Constraints”, arXiv:1803.08887V3, Dec. 15, 2018, 20 pages. [cited by applicant]
Trautmann et al., “On the Distribution of the Desirability Index using Harrington's Desirability Function”, Metrika, vol. 63, Issue 2, Apr. 2006, pp. 207-213. [cited by applicant]
Triguero et al., “Self-Labeled Techniques for Semi-Supervised Learning: Taxonomy, Software and Empirical Study”, Knowledge and Information Systems, vol. 42, Issue 2, 2015, pp. 245-284. [cited by applicant]
Tuomisto, “A Consistent Terminology for Quantifying Species Diversity? Yes, It Does Exist” Oecologia, vol. 164, 2010, pp. 853-860. [cited by applicant]
Vacek et al., “Using Case-Based Reasoning for Autonomous Vehicle Guidance”, International Conference on Intelligent Robots and Systems, San Diego, California, Oct. 29-Nov. 2, 2007, 5 pages. [cited by applicant]
Verleysen et al., “The Curse of Dimensionality in Data Mining and Time Series Prediction” International Work-Conference on Artificial Neural Networks, Barcelona, Spain, Jun. 8-10, 2005, pp. 758-770. [cited by applicant]
Viera, “Generating Synthetic Sequential Data using GANs”, Medium: Toward AI, Jun. 29, 2020, 31 pages. [cited by applicant]
Wachter et al., “Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR”, Harvard Journal of Law and Technology, vol. 31, No. 2, Spring 2018, 47 pages. [cited by applicant]
Wang et al., “Falling Rule Lists” 18th International Conference on Artificial Intelligence and Statistics, San Diego, California, May 9-12, 2015, 10 pages. [cited by applicant]
Wei et al., “An Operation-Time Simulation Framework for UAV Swarm Configuration and Mission Planning”, 2013. [cited by applicant]
Weselkowski et al., “TraDE: Training Device Selection via Multi-Objective Optimization”, IEEE 2014. [cited by applicant]
Wu, “The Synthetic Student: A Machine Learning Model to Simulate MOOC Data”, Massachusetts Institute of Technology, Master's Thesis, May 2015, 103 pages. [cited by applicant]
Xiao, “Towards Automatically Linking Data Elements”, Massachusetts Institute of Technology, Master's Thesis, Jun. 2017, 92 pages. [cited by applicant]
Xu et al., “An Algorithm for Remote Sensing Image Classification Based on Artificial Immune B-Cell Network”, The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. XXXVII… [cited by applicant]
Zhao et al., “Semi-Supervised Multi-Label Learning with Incomplete Labels”, 24th International Joint Conference on Artificial Intelligence, Buenos Aires, Argentina, Jul. 25-31, 2015, pp. 4062-4068. [cited by applicant]