IP Library Granted Patent US 12,260,348
Granted Patent B2
US 12,260,348 · App. 18/364,915 · Granted Mar 25, 2025

Search and query in computer-based reasoning systems

Inventors: Michael Auerbach (Bellevue, WA); Michael Resnick (Raleigh, NC); Christopher James Hazard (Raleigh, NC)
Assignee: Howso Incorporated
G06N5/04G06F16/2264G06F16/245
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,348
App. No.
18/364,915
Granted
Mar 25, 2025
Kind
B2
Abstract

Techniques for improved searching and querying in computer-based reasoning systems are discussed and include receiving multiple new multidimensional data element to store in a computer-based reasoning data model; determining a feature bucket for each feature of each data element and storing a reference identifier in the feature bucket(s). A query on the computer-based reasoning system includes input data element (e.g., an actual data element, or a set of restrictions on features). For each feature in the input data element, feature buckets are determined, candidate results are determined based on whether cases have related feature buckets, and the results are determined based at least in part on the candidate results. In some embodiments, control of controllable systems may be caused based on the results.

Claims (77)

1. A method comprising:

receiving multiple multidimensional data elements to store in a computer-based data model, wherein at least a first dimension of a plurality of dimensions associated with the multiple multidimensional data elements comprises strings;

for each received multidimensional data element:

determining a reference identifier for the multidimensional data element;

for each dimension in the multidimensional data element:

determining a feature bucket from a set of feature buckets for a value of the dimension of the multidimensional data element;

storing in a feature bucketed data structure, in a feature bucket corresponding to the determined feature bucket for the value of the dimension of the multidimensional data element, the reference identifier to the multidimensional data element;

receiving a query for related data elements to an input multidimensional data element;

for each feature in the input multidimensional data element:

determining a feature bucket for a value of the feature of the input multidimensional data element;

determining, as one or more candidate multidimensional data elements and from the feature bucket data structure, one or more multidimensional data elements that have related feature buckets with the input multidimensional data element;

determining the related data elements based at least in part on the respective distances from the one or more candidate multidimensional data elements to the input multidimensional data element;

returning the determined related data elements;

generating instructions for a controllable system based at least in part on the returned related data elements, wherein the controllable system is a type of system for autonomous vehicles, image labeling data, laboratory control, health care decision making, smart voice control, control of federated devices, manufacturing data, energy transfer systems, or smart home data;

causing control of the controllable system by transmitting the instructions to the controllable system.

2. The method of claim 1 , wherein feature buckets in the set of feature buckets are strictly ordered and values in feature bucket i of the set of feature buckets are all greater than values in feature bucket i- 1 in the set of feature buckets.

3. The method of claim 1 , further comprising:

when a number of the one or more candidate multidimensional data elements is less than a threshold:

for each feature in the input multidimensional data element:

determining from the feature bucket data structure, one or more multidimensional data elements that are within a threshold number of feature buckets in the set of feature buckets from the determined feature bucket for the input multidimensional data element;

ranking one or more candidate multidimensional data elements based at least in part on a number of feature buckets within a threshold number of corresponding feature bucket of the input multidimensional data element.

4. The method of claim 1 , further comprising:

determining that a number of items stored in a particular feature bucket is outside a threshold, revising a set of ranges for one or more feature buckets in the set of feature buckets.

5. The method of claim 1 , wherein the strings of the first dimension comprise non-numeric strings.

6. The method of claim 1 , wherein determining from the feature bucket data structure, one or more multidimensional data elements that have related feature buckets with the input multidimensional data element comprises determining from the feature bucket data structure, one or more multidimensional data elements that share the same feature buckets with the input multidimensional data element.

7. The method of claim 1 , wherein determining from the feature bucket data structure, one or more multidimensional data elements that have related feature buckets with the input multidimensional data element comprises determining from the feature bucket data structure, one or more multidimensional data elements that have nearby feature buckets as compared to the input multidimensional data element.

8. The method of claim 1 , wherein receiving the query for related data elements to the input multidimensional data element comprises receiving a query for results to a structured query.

9. The method of claim 1 , wherein receiving the query for related data elements to the input multidimensional data element comprises receiving the query for k nearest neighbor data elements to the input multidimensional data element.

10. A system for performing a structured query, comprising one or more computing devices, which one or more computing devices are configured to perform a method of:

receiving multiple multidimensional data elements to store in a computer-based data model, wherein at least a first dimension of a plurality of dimensions associated with the multiple multidimensional data elements comprises strings;

for each received multidimensional data element:

determining a reference identifier for the multidimensional data element;

for each dimension in the multidimensional data element:

determining a feature bucket from a set of feature buckets for a value of the dimension of the multidimensional data element;

storing in a feature bucketed data structure, in a feature bucket corresponding to the determined feature bucket for the value of the dimension of the multidimensional data element, the reference identifier to the multidimensional data element;

receiving a structured query for related data elements to an input multidimensional data element, wherein the structured query relates to the input multidimensional data element;

for each feature in the input multidimensional data element:

determining a feature bucket for a value of the feature of the input multidimensional data element;

determining, as one or more candidate multidimensional data elements and from the feature bucket data structure, one or more multidimensional data elements that have related feature buckets with the input multidimensional data element;

determining a respective distance from each candidate multidimensional data element to the input multidimensional data using a Damerau-Levenshtein distance for at least the first dimension comprising strings;

determining the related data elements based at least in part on the respective distances from the one or more candidate multidimensional data elements to the input multidimensional data element;

returning the determined related data elements;

generating instructions for a controllable system based at least in part on the returned related data elements, wherein the controllable system is a type of system for autonomous vehicles, image labeling data, laboratory control, health care decision making, smart voice control, control of federated devices, manufacturing data, energy transfer systems, or smart home data;

causing control of the controllable system by transmitting the instructions to the controllable system.

11. The system of claim 10 , wherein feature buckets in the set of feature buckets are strictly ordered and values in feature bucket i of the set of feature buckets are all greater than values in feature bucket i- 1 in the set of feature buckets.

12. The system of claim 10 , the method further comprising:

when a number of the one or more candidate multidimensional data elements is less than a threshold:

for each feature in the input multidimensional data element:

determining from the feature bucket data structure, one or more multidimensional data elements that are within a threshold number of feature buckets in the set of feature buckets from the determined feature bucket for the input multidimensional data element;

ranking one or more candidate multidimensional data elements based at least in part on a number of feature buckets within a threshold number of corresponding feature bucket of the input multidimensional data element.

13. The system of claim 10 , wherein determining from the feature bucket data structure, one or more multidimensional data elements that have related feature buckets with the input multidimensional data element comprises determining from the feature bucket data structure, one or more multidimensional data elements that share the same feature buckets with the input multidimensional data element.

14. The system of claim 10 , wherein determining from the feature bucket data structure, one or more multidimensional data elements that have related feature buckets with the input multidimensional data element comprises determining from the feature bucket data structure, one or more multidimensional data elements that have nearby feature buckets as compared to the input multidimensional data element.

15. One or more non-transitory storage media storing instructions which, when executed by one or more computing devices, cause performance of a method of:

receiving multiple multidimensional data elements to store in a computer-based data model, wherein at least a first dimension of a plurality of dimensions associated with the multiple multidimensional data elements comprises strings;

for each received multidimensional data element:

determining a reference identifier for the multidimensional data element;

for each dimension in the multidimensional data element:

determining a feature bucket from a set of feature buckets for a value of the dimension of the multidimensional data element;

storing in a feature bucketed data structure, in a feature bucket corresponding to the determined feature bucket for the value of the dimension of the multidimensional data element, the reference identifier to the multidimensional data element;

receiving a query for related data elements to an input multidimensional data element;

for each feature in the input multidimensional data element:

determining a feature bucket for a value of the feature of the input multidimensional data element;

determining, as one or more candidate multidimensional data elements and from the feature bucket data structure, one or more multidimensional data elements that have related feature buckets with the input multidimensional data element;

determining a respective distance from each candidate multidimensional data element to the input multidimensional data using a Damerau-Levenshtein distance for at least the first dimension comprising strings;

determining the related data elements based at least in part on the respective distances from the one or more candidate multidimensional data elements to the input multidimensional data element;

returning the determined related data elements;

generating instructions for a controllable system based at least in part on the returned related data elements, wherein the controllable system is a type of system for autonomous vehicles, image labeling data, laboratory control, health care decision making, smart voice control, control of federated devices, manufacturing data, energy transfer systems, or smart home data;

causing control of the controllable system by transmitting the instructions to the controllable system.

16. The one or more non-transitory storage media of claim 15 , wherein feature buckets in the set of feature buckets are strictly ordered and values in feature bucket i of the set of feature buckets are all greater than values in feature bucket i- 1 in the set of feature buckets.

17. The one or more non-transitory storage media of claim 16 , the method further comprising:

when a number of the one or more candidate multidimensional data elements is less than a threshold:

for each feature in the input multidimensional data element:

determining from the feature bucket data structure, one or more multidimensional data elements that are within a threshold number of feature buckets in the set of feature buckets from the determined feature bucket for the input multidimensional data element;

ranking one or more candidate multidimensional data elements based at least in part on a number of feature buckets within a threshold number of corresponding feature bucket of the input multidimensional data element.

18. The one or more non-transitory storage media of claim 16 , wherein determining from the feature bucket data structure, one or more multidimensional data elements that have related feature buckets with the input multidimensional data element comprises determining from the feature bucket data structure, one or more multidimensional data elements that share the same feature buckets with the input multidimensional data element.

19. The one or more non-transitory storage media of claim 16 , wherein determining from the feature bucket data structure, one or more multidimensional data elements that have related feature buckets with the input multidimensional data element comprises determining from the feature bucket data structure, one or more multidimensional data elements that have nearby feature buckets as compared to the input multidimensional data element.

20. The one or more non-transitory storage media of claim 16 , wherein receiving the query for related data elements to the input multidimensional data element comprises receiving a query for results to a structured query.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2024
From: AUERBACH, MICHAEL
To: DIVEPLANE CORPORATION
Reel/Frame 066249/0682 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: RESNICK, MICHAEL; HAZARD, CHRISTOPHER JAMES
To: DIVEPLANE CORPORATION
Reel/Frame 065148/0855 →
CHANGE OF NAME Recorded Sep 28, 2023
From: DIVEPLANE CORPORATION
To: HOWSO INCORPORATED
Reel/Frame 065081/0559 →
Continuity (3)
Continuation 16796258 · Feb 20, 2020
Provisional Application 62848729 · May 16, 2019
Related Publication 20240046125A1 · Feb 8, 2024
References Cited (153)
US 4935877A · Koza · 1990 [cited by applicant]
US 5581664A · Allen et al. · 1996 [cited by applicant]
US 6282527B1 · Gounares et al. · 2001 [cited by applicant]
US 6741972B1 · Girardi et al. · 2004 [cited by applicant]
US 7873587B2 · Baum · 2011 [cited by applicant]
US 9489635B1 · Zhu · 2016 [cited by applicant]
US 9922286B1 · Hazard · 2018 [cited by applicant]
US 10459444B1 · Kentley-Klay · 2019 [cited by applicant]
US 20010049595A1 · Plumer et al. · 2001 [cited by applicant]
US 20040019851A1 · Purvis et al. · 2004 [cited by applicant]
US 20050137992A1 · Polak · 2005 [cited by applicant]
US 20060195204A1 · Bonabeau et al. · 2006 [cited by applicant]
US 20080153098A1 · Rimm et al. · 2008 [cited by applicant]
US 20080307399A1 · Zhou et al. · 2008 [cited by applicant]
US 20090006299A1 · Baum · 2009 [cited by applicant]
US 20090144704A1 · Niggemann et al. · 2009 [cited by applicant]
US 20100106603A1 · Dey et al. · 2010 [cited by applicant]
US 20100287507A1 · Paquette et al. · 2010 [cited by applicant]
US 20110060895A1 · Solomon · 2011 [cited by applicant]
US 20110072206A1 · Ross · 2011 [cited by examiner]
US 20110161264A1 · Cantin · 2011 [cited by applicant]
US 20110225564A1 · Biswas et al. · 2011 [cited by applicant]
US 20120278321A1 · Traub · 2012 [cited by examiner]
US 20130006901A1 · Cantin · 2013 [cited by applicant]
US 20130339365A1 · Balasubramanian et al. · 2013 [cited by applicant]
US 20160055427A1 · Adjaoute · 2016 [cited by applicant]
US 20170010106A1 · Shashua et al. · 2017 [cited by applicant]
US 20170012772A1 · Mueller · 2017 [cited by applicant]
US 20170013547A1 · Skaaksrud · 2017 [cited by examiner]
US 20170053211A1 · Heo et al. · 2017 [cited by applicant]
US 20170091645A1 · Matus · 2017 [cited by applicant]
US 20170161640A1 · Shamir · 2017 [cited by applicant]
US 20170236060A1 · Ignatyev · 2017 [cited by applicant]
US 20180072323A1 · Gordon et al. · 2018 [cited by applicant]
US 20180089563A1 · Redding et al. · 2018 [cited by applicant]
US 20180235649A1 · Elkadi · 2018 [cited by applicant]
US 20180336018A1 · Lu et al. · 2018 [cited by applicant]
US 20190147331A1 · Arditi · 2019 [cited by applicant]
WO WO2017057528 · 2017 [cited by applicant]
WO WO2017189859 · 2017 [cited by applicant]
Abdi, “Cardinality Optimization Problems”, The University of Birmingham, PhD Dissertation, May 2013, 197 pages. [cited by applicant]
Aboulnaga, “Generating Synthetic Complex-structured XML Data”, Proceedings of the Fourth International Workshop on the Web and Databases, WebDB 2001, Santa Barbara, California, USA, May 24-25, 2001, 6 pages. [cited by applicant]
Abramson, “The Expected-Outcome Model of Two-Player Games”, PhD Thesis, Columbia University, New York, New York, 1987, 125 pages. [cited by applicant]
Agarwal et al., “Nearest-Neighbor Searching Under Uncertainty II”, ACM Transactions on Algorithms, vol. 13, Issue 1, Article 3, 2016, 25 pages. [cited by applicant]
Aggarwal et al., “On the Surprising Behavior of Distance Metrics in High Dimensional Space”, International Conference on Database Theory, London, United Kingdom, Jan. 4-6, 2001, pp. 420-434. [cited by applicant]
Akaike, “Information Theory and an Extension of the Maximum Likelihood Principle”, Proceedings of the 2nd International Symposium on Information Theory, Sep. 2-8, 1971, Tsahkadsor, Armenia, pp. 267-281. [cited by applicant]
Alhaija, “Augmented Reality Meets Computer Vision: Efficient Data Generation for Urban Driving Scenes”, arXiv:1708.01566v1, Aug. 4, 2017, 12 pages. [cited by applicant]
Alpaydin, “Machine Learning: The New AI”, MIT Press, Cambridge, Massachusetts, 2016, 225 pages. [cited by applicant]
Alpaydin, “Voting Over Multiple Condensed Nearest Neighbor”, Artificial Intelligence Review, vol. 11, 1997, pp. 115-132. [cited by applicant]
Altman, “An Introduction to Kernel and Nearest-Neighbor Nonparametric Regression”, The American Statistician, vol. 46, Issue 3, 1992, pp. 175-185. [cited by applicant]
Anderson, “Synthetic data generation for the internet of things,” 2014 IEEE International Conference on Big Data (Big Data), Washington, DC, 2014, pp. 171-176. [cited by applicant]
Archer et al., “Empirical Characterization of Random Forest Variable Importance Measures”, Computational Statistics & Data Analysis, vol. 52, 2008, pp. 2249-2260. [cited by applicant]
Beyer et al., “When is ‘Nearest Neighbor’ Meaningful?” International Conference on Database Theory, Springer, Jan. 10-12, 1999, Jerusalem, Israel, pp. 217-235. [cited by applicant]
Bull, Haploid-Diploid Evolutionary Algorithms: The Baldwin Effect and Recombination Nature's Way, Artificial Intelligence and Simulation of Behaviour Convention, Apr. 19-21, 2017, Bath, United Kingdom, pp. 91-94. [cited by applicant]
Cano et al., “Evolutionary Stratified Training Set Selection for Extracting Classification Rules with Tradeoff Precision-Interpretability” Data and Knowledge Engineering, vol. 60, 2007, pp. 90-108. [cited by applicant]
Chawla, “SMOTE: Synthetic Minority Over-sampling Technique”, arXiv:1106.1813v1, Jun. 9, 2011, 37 pages. [cited by applicant]
Chen, “DropoutSeer: Visualizing Learning Patterns in Massive Open Online Courses for Dropout Reasoning and Prediction”, 2016 IEEE Conference on Visual Analytics Science and Technology (VAST), Oct. 23-28, 2016, Baltimore… [cited by applicant]
Chomboon et al., An Empirical Study of Distance Metrics for k-Nearest Neighbor Algorithm, 3rd International Conference on Industrial Application Engineering, Kitakyushu, Japan, Mar. 28-31, 2015, pp. 280-285. [cited by applicant]
Colakoglu, “A Generalization of the Minkowski Distance and a New Definition of the Ellipse”, arXiv:1903.09657v1, Mar. 2, 2019, 18 pages. [cited by applicant]
Dernoncourt, “MoocViz: A Large Scale, Open Access, Collaborative, Data Analytics Platform for MOOCs”, NIPS 2013 Education Workshop, Nov. 1, 2013, Lake Tahoe, Utah, USA, 8 pages. [cited by applicant]
Ding, “Generating Synthetic Data for Neural Keyword-to-Question Models”, arXiv:1807.05324v1, Jul. 18, 2018, 12 pages. [cited by applicant]
Dwork et al., “The Algorithmic Foundations of Differential Privacy”, Foundations and Trends in Theoretical Computer Science, vol. 9, Nos. 3-4, 2014, pp. 211-407. [cited by applicant]
Efros et al., “Texture Synthesis by Non-Parametric Sampling”, International Conference on Computer Vision, Sep. 20-25, 1999, Corfu, Greece, 6 pages. [cited by applicant]
Fathony, “Discrete Wasserstein Generative Adversarial Networks (DWGAN)”, OpenReview.net, Feb. 18, 2018, 20 pages. [cited by applicant]
Ganegedara et al., “Self Organizing Map Based Region of Interest Labelling for Automated Defect Identification in Large Sewer Pipe Image Collections”, IEEE World Congress on Computational Intelligence, Jun. 10-15, 2012,… [cited by applicant]
Gao et al., “Efficient Estimation of Mutual Information for Strongly Dependent Variables”, 18th International Conference on Artificial Intelligence and Statistics, San Diego, California, May 9-12, 2015, pp. 277-286. [cited by applicant]
Gehr et al., “AI2: Safety and Robustness Certification of Neural Networks with Abstract Interpretation”, 39th IEEE Symposium on Security and Privacy, San Francisco, California, May 21-23, 2018, 16 pages. [cited by applicant]
Gemmeke et al., “Using Sparse Representations for Missing Data Imputation in Noise Robust Speech Recognition”, European Signal Processing Conference, Lausanne, Switzerland, Aug. 25-29, 2008, 5 pages. [cited by applicant]
Ghosh, “Inferential Privacy Guarantees for Differentially Private Mechanisms”, arXiv:1603:.01508v1, Mar. 4, 2016, 31 pages. [cited by applicant]
Goodfellow et al., “Deep Learning”, 2016, 800 pages. [cited by applicant]
Google AI Blog, “The What-If Tool: Code-Free Probing of Machine Learning Models”, Sep. 11, 2018, https://pair-code,github.io/what-if-tool, retrieved on Mar. 14, 2019, 5 pages. [cited by applicant]
Gottlieb et al., “Near-Optimal Sample Compression for Nearest Neighbors”, Advances in Neural Information Processing Systems, Montreal, Canada, Dec. 8-13, 2014, 9 pages. [cited by applicant]
Gray, “Quickly Generating Billion-Record Synthetic Databases”, SIGMOD '94: Proceedings of the 1994 ACM SIGMOD international conference on Management of data, May 1994, Minneapolis, Minnesota, USA, 29 pages. [cited by applicant]
Hastie et al., “The Elements of Statistical Learning”, 2001, 764 pages. [cited by applicant]
Hazard et al., “Natively Interpretable Machine Learning and Artificial Intelligence: Preliminary Results and Future Directions”, arXiv:1901v1, Jan. 2, 2019, 15 pages. [cited by applicant]
Hinneburg et al., “What is the Nearest Neighbor in High Dimensional Spaces?”, 26th International Conference on Very Large Databases, Cairo, Egypt, Sep. 10-14, 2000, pp. 506-515. [cited by applicant]
Hmeidi et al., “Performance of KNN and SVM Classifiers on Full Word Arabic Articles”, Advanced Engineering Informatics, vol. 22, Issue 1, 2008, pp. 106-111. [cited by applicant]
Hoag “A Parallel General-Purpose Synthetic Data Generator”, ACM SIGMOID Record, vol. 6, Issue 1, Mar. 2007, 6 pages. [cited by applicant]
Houle et al., “Can Shared-Neighbor Distances Defeat the Curse of Dimensionality?”, International Conference on Scientific and Statistical Database Management, Heidelberg, Germany, Jun. 31-Jul. 2, 2010, 18 pages. [cited by applicant]
Imbalanced-learn, “SMOTE”, 2016-2017, https://imbalanced-learn.readthedocs.io/en/stable/generated/imblearn.over_sampling.SMOTE.html, retrieved on Aug. 11, 2020, 6 pages. [cited by applicant]
Indyk et al., “Approximate Nearest Neighbors: Towards Removing the Curse of Dimensionality”, Procedures of the 30th ACM Symposium on Theory of Computing, Dallas, Texas, May 23-26, 1998, pp. 604-613. [cited by applicant]
International Search Report and Written Opinion for PCT/US2018/047118, mailed on Dec. 3, 2018, 9 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/US2019/026502, mailed on Jul. 24, 2019, 16 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/US2019/066321, mailed on Mar. 19, 2020, 17 pages. [cited by applicant]
Internet Archive, “System Verilog distribution Constraint—Verification Guide”, Aug. 6, 2018, http://web.archive.org/web/20180806225430/https://www.verificationguide.com/p/systemverilog-distribution-constraint.html, retr… [cited by applicant]
Internet Archive, “System Verilog Testbench Automation Tutorial”, Nov. 17, 2016, https://web.archive.org/web/20161117153225/http://www.doulos.com/knowhow/sysverilog/tutorial/constraints/, retrieved on Mar. 3, 2020, 6 pa… [cited by applicant]
Kittler, “Feature Selection and Extraction”, Handbook of Pattern Recognition and Image Processing, Jan. 1986, Chapter 3, pp. 115-132. [cited by applicant]
Kohavi et al., “Wrappers for Feature Subset Selection”, Artificial Intelligence, vol. 97, Issues 1-2, Dec. 1997, pp. 273-323. [cited by applicant]
Kontorovich et al., “Nearest-Neighbor Sample Compression: Efficiency, Consistency, Infinite Dimensions”, Advances in Neural Information Processing Systems, 2017, pp. 1573-1583. [cited by applicant]
Kulkarni et al., “Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation”, arXiv:1604.06057v2, May 31, 2016, 14 pages. [cited by applicant]
Kuramochi et al., “Gene Classification using Expression Profiles: A Feasibility Study”, Technical Report TR 01-029, Department of Computer Science and Engineering, University of Minnesota, Jul. 23, 2001, 18 pages. [cited by applicant]
Leinster et al., “Maximizing Diversity in Biology and Beyond”, Entropy, vol. 18, Issue 3, 2016, 23 pages. [cited by applicant]
Liao et al., “Similarity Measures for Retrieval in Case-Based Reasoning Systems”, Applied Artificial Intelligence, vol. 12, 1998, pp. 267-288. [cited by applicant]
Lin et al., “Why Does Deep and Cheap Learning Work So Well?” Journal of Statistical Physics, vol. 168, 2017, pp. 1223-1247. [cited by applicant]
Lin, “Development of a Synthetic Data Set Generator for Building and Testing Information Discovery Systems”, Proceedings of the Third International Conference on Information Technology: New Generations, Nevada, USA, Apr… [cited by applicant]
Lukaszyk, “A New Concept of Probability Metric and its Applications in Approximation of Scattered Data Sets”, Computational Mechanics, vol. 33, 2004, pp. 299-304. [cited by applicant]
Lukaszyk, “Probability Metric, Examples of Approximation Applications in Experimental Mechanics”, PhD Thesis, Cracow University of Technology, 2003, 149 pages. [cited by applicant]
Mann et al., “On a Test of Whether One or Two Random Variables is Stochastically Larger than the Other”, The Annals of Mathematical Statistics, 1947, pp. 50-60. [cited by applicant]
Martino et al., “A Fast Universal Self-Tuned Sampler within Gibbs Sampling”, Digital Signal Processing, vol. 47, 2015, pp. 68-83. [cited by applicant]
Mohri et al., Foundations of Machine Learning, 2012, 427 pages—uploaded as Part 1 and Part 2. [cited by applicant]
Montanez, “SDV: An Open Source Library for Synthetic Data Generation”, Massachusetts Institute of Technology, Master's Thesis, Sep. 2018, 105 pages. [cited by applicant]
Negra, “Model of a Synthetic Wind Speed Time Series Generator”, Wind Energy, Wiley Interscience, Sep. 6, 2007, 17 pages. [cited by applicant]
Nguyen et al., “NP-Hardness of ζ 0 Minimization Problems: Revision and Extension to the Non-Negative Setting”, 13th International Conference on Sampling Theory and Applications, Jul. 8-12, 2019, Bordeaux, France, 4 page… [cited by applicant]
Olson et al., “PMLB: A Large Benchmark Suite for Machine Learning Evaluation and Comparison”, arXiv:1703.00512v1, Mar. 1, 2017, 14 pages. [cited by applicant]
Patki, “The Synthetic Data Vault: Generative Modeling for Relational Databases”, Massachusetts Institute of Technology, Master's Thesis, Jun. 2016, 80 pages. [cited by applicant]
Patki, “The Synthetic Data Vault”, 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA), Oct. 17-19, 2016, Montreal, QC, pp. 399-410. [cited by applicant]
Pedregosa et al., “Machine Learning in Python”, Journal of Machine Learning Research, vol. 12, 2011, pp. 2825-2830. [cited by applicant]
Pei, “A Synthetic Data Generator for Clustering and Outlier Analysis”, The University of Alberta, 2006, 33 pages. [cited by applicant]
Phan et al., “Adaptive Laplace Mechanism: Differential Privacy Preservation in Deep Learning” 2017 IEEE International Conference on Data Mining, New Orleans, Louisiana, Nov. 18-21, 2017, 10 pages. [cited by applicant]
Poerner et al., “Evaluating Neural Network Explanation Methods Using Hybrid Documents and Morphosyntactic Agreement”, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Long Papers)… [cited by applicant]
Prakosa “Generation of Synthetic but Visually Realistic Time Series of Cardiac Images Combining a Biophysical Model and Clinical Images” IEEE Transactions on Medical Imaging, vol. 32, No. 1, Jan. 2013, pp. 99-109. [cited by applicant]
Priyardarshini, “Wedagen: A synthetic web database generator”, International Workshop of Internet Data Management (IDM'99), Sep. 2, 1999, Florence, IT, 24 pages. [cited by applicant]
Pudjijono, “Accurate Synthetic Generation of Realistic Personal Information” Advances in Knowledge Discovery and Data Mining, Pacific-Asia Conference on Knowledge Discovery and Data Mining, Apr. 27-30, 2009, Bangkok, Th… [cited by applicant]
Raikwal et al., “Performance Evaluation of SVM and K-Nearest Neighbor Algorithm Over Medical Data Set”, International Journal of Computer Applications, vol. 50, No. 14, Jul. 2012, pp. 975-985. [cited by applicant]
Rao et al., “Cumulative Residual Entropy: A New Measure of Information”, IEEE Transactions on Information Theory, vol. 50, Issue 6, 2004, pp. 1220-1228. [cited by applicant]
Reiter, “Using Cart to Generate Partially Synthetic Public Use Microdata” Journal of Official Statistics, vol. 21, No. 3, 2005, pp. 441-462. [cited by applicant]
Ribeiro et al., “‘Why Should I Trust You’: Explaining the Predictions of Any Classifier”, arXiv:1602.04938v3, Aug. 9, 2016, 10 pages. [cited by applicant]
Rosenberg et al., “Semi-Supervised Self-Training of Object Detection Models”, IEEE Workshop on Applications of Computer Vision, 2005, 9 pages. [cited by applicant]
Schaul et al., “Universal Value Function Approximators”, International Conference on Machine Learning, Lille, France, Jul. 6-11, 2015, 9 pages. [cited by applicant]
Schlabach et al., “Fox-GA: A Genetic Algorithm for Generating and ANalayzing Battlefield Courses of Action”, 1999 MIT. [cited by applicant]
Schreck, “Towards An Automatic Predictive Question Formulation”, Massachusetts Institute of Technology, Master's Thesis, Jun. 2016, 121 pages. [cited by applicant]
Schreck, “What would a data scientist ask? Automatically formulating and solving prediction problems”, 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA), Oct. 17-19, 2016, Montreal, QC, pp… [cited by applicant]
Schuh et al., “Improving the Performance of High-Dimensional KNN Retrieval Through Localized Dataspace Segmentation and Hybrid Indexing”, Proceedings of the 17th East European Conference, Advances in Databases and Infor… [cited by applicant]
Schuh et al., “Mitigating the Curse of Dimensionality for Exact KNN Retrieval”, Proceedings of the 26th International Florida Artificial Intelligence Research Society Conference, St. Pete Beach, Florida, May 22-24, 2014… [cited by applicant]
Schwarz et al., “Estimating the Dimension of a Model”, The Annals of Statistics, vol. 6, Issue 2, Mar. 1978, pp. 461-464. [cited by applicant]
Silver et al., “Mastering the Game of Go Without Human Knowledge”, Nature, vol. 550, Oct. 19, 2017, pp. 354-359. [cited by applicant]
Skapura, “Building Neural Networks”, 1996, p. 63. [cited by applicant]
Smith, “FeatureHub: Towards collaborative data science”, 2017 IEEE International Conference on Data Science and Advanced Analytics (DSAA), Oct. 19-21, 2017, Tokyo, pp. 590-600. [cited by applicant]
Stephenson et al., “A Continuous Information Gain Measure to Find the Most Discriminatory Problems for AI Benchmarking”, arxiv.org, arxiv.org/abs/1809.02904v2, retrieved on Aug. 21, 2019, 8 pages. [cited by applicant]
Stoppiglia et al., “Ranking a Random Feature for Variable and Feature Selection” Journal of Machine Learning Research, vol. 3, 2003, pp. 1399-1414. [cited by applicant]
Sun et al., “Fuzzy Modeling Employing Fuzzy Polyploidy Genetic Algorithms”, Journal of Information Science and Engineering, Mar. 2002, vol. 18, No. 2, pp. 163-186. [cited by applicant]
Sun, “Learning Vine Copula Models for Synthetic Data Generation”, arXiv:1812.01226v1, Dec. 4, 2018, 9 pages. [cited by applicant]
Surya et al., “Distance and Similarity Measures Effect on the Performance of K-Nearest Neighbor Classifier”, arXiv:1708.04321v1, Aug. 14, 2017, 50 pages. [cited by applicant]
Tan et al., “Incomplete Multi-View Weak-Label Learning”, 27th International Joint Conference on Artificial Intelligence, 2018, pp. 2703-2709. [cited by applicant]
Tao et al., “Quality and Efficiency in High Dimensional Nearest Neighbor Search”, Proceedings of the 2009 ACM SIGMOD International Conference on Management of Data, Providence, Rhode Island, Jun. 29-Jul. 2, 2009, pp. 56… [cited by applicant]
Tishby et al., “Deep Learning and the Information Bottleneck Principle”, arXiv:1503.02406v1, Mar. 9, 2015, 5 pages. [cited by applicant]
Tockar, “Differential Privacy: The Basics”, Sep. 8, 2014, https://research.neustar.biz/2014/09/08/differential-privacy-the-basics/ retrieved on Apr. 1, 2019, 3 pages. [cited by applicant]
Tomasev et al., “Hubness-aware Shared Neighbor Distances for High-Dimensional k-Nearest Neighbor Classification”, 7th international Conference on Hybrid Artificial Intelligent Systems, Salamanca, Spain, Mar. 28-30, 2012… [cited by applicant]
Tran, “Dist-GAN: An Improved GAN Using Distance Constraints”, arXiv:1803.08887V3, Dec. 15, 2018, 20 pages. [cited by applicant]
Trautmann et al., “On the Distribution of the Desirability Index using Harrington's Desirability Function”, Metrika, vol. 63, Issue 2, Apr. 2006, pp. 207-213. [cited by applicant]
Triguero et al., “Self-Labeled Techniques for Semi-Supervised Learning: Taxonomy, Software and Empirical Study”, Knowledge and Information Systems, vol. 42, Issue 2, 2015, pp. 245-284. [cited by applicant]
Tuomisto, “A Consistent Terminology for Quantifying Species Diversity? Yes, It Does Exist” Oecologia, vol. 164, 2010, pp. 853-860. [cited by applicant]
Vacek et al., “Using Case-Based Reasoning for Autonomous Vehicle Guidance”, International Conference on Intelligent Robots and Systems, San Diego, California, Oct. 29-Nov. 2, 2007, 5 pages. [cited by applicant]
Verleysen et al., “The Curse of Dimensionality in Data Mining and Time Series Prediction” International Work-Conference on Artificial Neural Networks, Barcelona, Spain, Jun. 8-10, 2005, pp. 758-770. [cited by applicant]
Viera, “Generating Synthetic Sequential Data using GANs”, Medium: Toward AI, Jun. 29, 2020, 31 pages. [cited by applicant]
Wachter et al., “Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR”, Harvard Journal of Law and Technology, vol. 31, No. 2, Spring 2018, 47 pages. [cited by applicant]
Wang et al., “Falling Rule Lists” 18th International Conference on Artificial Intelligence and Statistics, San Diego, California, May 9-12, 2015, 10 pages. [cited by applicant]
Wei et al., “An Operation-Time Simulation Framework for UAV Swarm Configuration and Mission Planning”, 2013. [cited by applicant]
Weselkowski et al., “TraDE: Training Device Selection via Multi-Objective Optimization”, IEEE 2014. [cited by applicant]
Wu, “The Synthetic Student: A Machine Learning Model to Simulate MOOC Data”, Massachusetts Institute of Technology, Master's Thesis, May 2015, 103 pages. [cited by applicant]
Xiao, “Towards Automatically Linking Data Elements”, Massachusetts Institute of Technology, Master's Thesis, Jun. 2017, 92 pages. [cited by applicant]
Xu et al., “An Algorithm for Remote Sensing Image Classification Based on Artificial Immune B-Cell Network”, The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. XXXVII… [cited by applicant]
Zhao et al., “Semi-Supervised Multi-Label Learning with Incomplete Labels”, 24th International Joint Conference on Artificial Intelligence, Buenos Aires, Argentina, Jul. 25-31, 2015, pp. 4062-4068. [cited by applicant]