IP Library › Granted Patent US 12,646,295
Granted Patent B2
US 12,646,295 · App. 18/199,825 · Granted Jun 2, 2026

Systems and methods for efficient dataset distillation using non-deterministic feature approximation

Inventors: Noel Loo (Cambridge, MA); Ramin Hasani (New York, NY); Alexander A. Amini (Brookline, MA); Daniela Rus (Weston, MA)
G06V10/774G06V10/776G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,295
App. No.
18/199,825
Granted
Jun 2, 2026
Kind
B2
Abstract

Dataset distillation compresses large datasets into smaller synthetic coresets that retain performance with the aim of reducing storage and computational burdens of processing an original, entire dataset. The present disclosure provides an improved algorithm that uses a non-deterministic feature approximation of neural network Gaussian process (NNGP) kernels, or other trained kernels, that reduces a kernel matrix computation to O(|S|). When combined with a modified Platt scaling loss, the disclosed algorithm can provide at least a 100-fold speedup over a Kernel-Inducing Points (KIP) algorithm and can run on a single graphics processing unit. The disclosed Random Feature Approximation Distillation (RFAD) algorithm can perform competitively with other dataset condensation algorithms in accuracy over a range of large-scale datasets, both in kernel regression and finite-width network training. The disclosed techniques can be effective on tasks such as model interpretability and data privacy preservation.

Claims (60)

1 . A method for performing dataset distillation:

sampling a batch of data from a dataset to form a coreset;

applying a non-deterministic feature neural network training kernel approximation to at least some portion of the coreset to define a modified coreset; and

generating a distilled dataset from the modified coreset to preserve privacy of the modified coreset, the distilled dataset comprising data that is synthetic and representative of the dataset

wherein the non-deterministic feature neural network training kernel is at least one of a neural network Gaussian process (NNGP) kernel, a neural tangent kernel (NTK), or other learned training kernels.

2 . The method of claim 1 , wherein the dataset comprises a large dataset exceeding approximately 10 4 samples, with each sample being of dimensionality of approximately 10 3 or larger.

3 . The method of claim 1 , further comprising:

applying Platt-scaling to the modified coreset to define a Platt-scaled coreset,

wherein generating the distilled dataset from the modified coreset further comprises generating the distilled dataset from the Platt-scaled coreset.

4 . The method of claim 3 ,

wherein applying Platt-scaling to the modified coreset further comprises applying a cross entropy loss to the Platt-scaled coreset, and

wherein generating the distilled dataset from the modified coreset further comprises generating the distilled dataset from the cross entropy loss applied Platt-scaled coreset.

5 . The method of claim 1 ,

wherein the dataset comprises a plurality of images, the images having labels associated therewith, and

wherein sampling a batch of data from a dataset comprises sampling a batch of images and labels from the dataset to form the coreset,

the method further comprising computing trained neural network predictions on the sampled batch of images.

6 . The method of claim 5 , further comprising:

after computing trained neural network predictions on the sampled batch of images, computing an accuracy of the trained network predictions on the sampled batch of images with respect to the labels associated with the respective images of the sampled batch of images; and

comparing at least one of:

the accuracy of the trained network predictions on the sampled batch of images to a threshold accuracy; or

a compute budget used to perform the action of applying a non-deterministic feature neural network training kernel approximation to a threshold compute budget,

wherein if at least one of the threshold accuracy or the threshold compute budget is not exceeded, performing the action of applying a non-deterministic feature neural network training kernel approximation to at least some portion of the coreset again, or

wherein if both the threshold accuracy and the threshold compute budget is exceeded, closing the distilled dataset.

7 . The method of claim 6 ,

wherein the dataset is an original dataset, and

wherein the threshold accuracy is about 70 percent of performance of learning the original dataset or better.

8 . The method of claim 6 , wherein a compute budget is approximately 14 GPU hours or less.

9 . The method of claim 1 , wherein an amount of time for performing the method is approximately in the range of about 1 hour to about 14 hours.

10 . The method of claim 1 , wherein the distilled dataset is minimized in size with respect to the dataset.

11 . The method of claim 1 , wherein the non-deterministic feature neural network training kernel is the NNGP.

12 . The method of claim 1 , wherein the distilled dataset is privacy-protected data.

13 . A system for performing dataset distillation, comprising:

a processor configured to perform a process comprising:

sampling a batch of data from a dataset to form a coreset;

applying a non-deterministic feature neural network training kernel approximation to at least some portion of the coreset to define a modified coreset; and

generating a distilled dataset from the modified coreset to preserve privacy of the modified coreset, the distilled dataset comprising data that is synthetic and representative of the dataset,

wherein the non-deterministic feature neural network training kernel is at least one of a neural network Gaussian process (NNGP) kernel, a neural tangent kernel (NTK), or other learned training kernels.

14 . The system of claim 13 ,

wherein the dataset comprises a plurality of images, the images having labels associated therewith, and

wherein sampling a batch of data from a dataset comprises sampling a batch of images and labels from the dataset to form the coreset,

the process the processor is configured to perform further comprises computing trained neural network predictions on the sampled batch of images.

15 . The system of claim 13 , wherein the process that the processor is configured to perform further comprises:

after computing trained neural network predictions on the sampled batch of images, computing an accuracy of the trained network predictions on the sampled batch of images with respect to the labels associated with the respective images of the sampled batch of images; and

comparing at least one of:

the accuracy of the trained network predictions on the sampled batch of images to a threshold accuracy; or

a compute budget used to perform the action of applying a non-deterministic feature neural network training kernel approximation to a threshold compute budget,

wherein if at least one of the threshold accuracy or the threshold compute budget is not exceeded, performing the action of applying a non-deterministic feature neural network training kernel approximation to at least some portion of the coreset again, or

wherein if both the threshold accuracy and the threshold compute budget is exceeded, closing the distilled dataset.

16 . The system of claim 13 , wherein the non-deterministic feature neural network training kernel is the NNGP.

17 . A method for performing dataset distillation for preserving data privacy:

sampling a batch of data from a dataset to form a coreset;

applying a non-deterministic feature neural network training kernel approximation to at least some portion of the coreset to define a modified coreset;

generating a distilled dataset from the modified coreset, the distilled dataset comprising data that is synthetic and representative of the dataset; and

returning the distilled dataset for preserving data privacy during run-time use in cloud infrastructures and local software as a service (SaaS) applications,

wherein the non-deterministic feature neural network training kernel is at least one of a neural network Gaussian process (NNGP) kernel, a neural tangent kernel (NIK), or other learned training kernels.

18 . The method of claim 17 , wherein the dataset comprises a large dataset exceeding approximately 10 4 samples, with each sample being of dimensionality of approximately 10 3 or larger.

19 . The method of claim 17 , further comprising:

applying Platt-scaling to the modified coreset to define a Platt-scaled coreset,

wherein generating the distilled dataset from the modified coreset further comprises generating the distilled dataset from the Platt-scaled coreset.

20 . The method of claim 17 , further comprising returning the distilled dataset for use in data privacy-preserving machine-learning-model training.

Assignments (2)
GOVERNMENT INTEREST AGREEMENT Recorded Aug 19, 2025
From: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
To: THE GOVERNMENT OF THE UNITED STATES OF AMERICA AS REPRESENTED BY THE SECRETARY OF THE NAVY
Reel/Frame 072476/0494 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2023
From: LOO, NOEL; AMINI, ALEXANDER A.; RUS, DANIELA; HASANI, RAMIN
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 065102/0805 →
Continuity (2)
Provisional Application 63390952 · Jul 20, 2022
Related Publication 20240212328A1 · Jun 27, 2024
References Cited (72)
US 20220398262A1 · Derakhshani · 2022 [cited by examiner]
CN 113569891A · 2021 [cited by examiner]
Refinetti, Maria, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová. “Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed.” In International Conference on Machine L… [cited by applicant]
Scholkopf B., et al., “Input space versus feature space in kernel-based methods,” in IEEE Transactions on Neural Networks, vol. 10, No. 5, pp. 1000-1017, Sep. 1999, doi: 10.1109/72.788641. [cited by applicant]
Shankar, Vaishaal, Alex Fang, Wenshuo Guo, Sara Fridovich-Keil, Jonathan Ragan-Kelley, Ludwig Schmidt, and Benjamin Recht. “Neural kernels without tangents.” In International conference on machine learning, pp. 8614-862… [cited by applicant]
Shleifer, Sam, and Eric Prokop. “Using small proxy datasets to accelerate hyperparameter search.” arXiv preprint arXiv:1906.04887 (2019). [cited by applicant]
Snell, Jake, Kevin Swersky, and Richard Zemel. “Prototypical networks for few-shot learning.” Advances in neural information processing systems 30 (2017). [cited by applicant]
Snelson, Edward, and Zoubin Ghahramani. “Sparse Gaussian processes using pseudo-inputs.” Advances in neural information processing systems 18 (2005). [cited by applicant]
Sucholutsky, Ilia, and Matthias Schonlau. “Soft-label dataset distillation and text dataset distillation.” In 2021 International Joint Conference on Neural Networks (IJCNN), pp. 1-8. IEEE, 2021. [cited by applicant]
Titsias, Michalis. “Variational learning of inducing variables in sparse Gaussian processes.” In Artificial intelligence and statistics, pp. 567-574. PMLR, 2009. [cited by applicant]
Toneva, Mariya, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J. Gordon. “An empirical study of example forgetting during deep neural network learning.” arXiv preprint arXiv: 18… [cited by applicant]
Vinyals, Oriol, Charles Blundell, Timothy Lillicrap, and Daan Wierstra. “Matching networks for one shot learning.” Advances in neural information processing systems 29 (2016). [cited by applicant]
Wang, Tongzhou, Jun-Yan Zhu, Antonio Torralba, and Alexei A. Efros. “Dataset distillation.” arXiv preprint arXiv:1811.10959 (2018). [cited by applicant]
Wang, Tsun-Hsuan, Wei Xiao, Tim Seyde, Ramin Hasani, and Daniela Rus. “Interpreting neural policies with disentangled tree representations.” arXiv preprint arXiv:2210.06650 (2022). [cited by applicant]
Williams, Christopher KI, and David Barber. “Bayesian classification with Gaussian processes.” IEEE Transactions on pattern analysis and machine intelligence 20, No. 12 (1998): 1342-1351. [cited by applicant]
Xiao, Han, Kashif Rasul, and Roland Vollgraf. “Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms.” arXiv preprint arXiv:1708.07747 (2017). [cited by applicant]
Yaida, Sho. “Non-Gaussian processes and neural networks at finite widths.” In Mathematical and Scientific Machine Learning, pp. 165-192. PMLR, 2020. [cited by applicant]
Yang, Greg. “Wide feedforward or recurrent neural networks of any architecture are gaussian processes.” Advances in Neural Information Processing Systems 32 (2019). [cited by applicant]
Zandieh, Amir, Insu Han, Haim Avron, Neta Shoham, Chaewon Kim, and Jinwoo Shin. “Scaling neural tangent kernels via sketching and random features.” Advances in Neural Information Processing Systems 34 (2021): 1062-1073. [cited by applicant]
Zhao, Bo, and Hakan Bilen. “Dataset condensation with distribution matching.” In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 6514-6523. 2023. [cited by applicant]
Zhao, Bo, and Hakan Bilen. “Dataset condensation with differentiable siamese augmentation.” In International Conference on Machine Learning, pp. 12674-12685. PMLR, 2021. [cited by applicant]
Zhao, Shengyu, Zhijian Liu, Ji Lin, Jun-Yan Zhu, and Song Han. “Differentiable augmentation for data-efficient gan training.” Advances in neural information processing systems 33 (2020): 7559-7570. [cited by applicant]
Zhuang, Juntang, Tommy Tang, Yifan Ding, Sekhar C. Tatikonda, Nicha Dvornek, Xenophon Papademetris, and James Duncan. “Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.” Advances in neural inf… [cited by applicant]
Abadi, Martin, et al. “Deep learning with differential privacy.” Proceedings of the 2016 ACM SIGSAC conference on computer and communications security (2016). [cited by applicant]
Aljundi, Rahaf, et al. “Gradient based sample selection for online continual learning.” Advances in neural information processing systems 32 (2019). [cited by applicant]
Belouadah, Eden, and Adrian Popescu. “Scail: Classifier weights scaling for class incremental learning.” Proceedings of the IEEE/CVF winter conference on applications of computer vision. 2020. [cited by applicant]
Bien, Jacob, and Robert Tibshirani. “Prototype selection for interpretable classification.” (2011): 2403-2424. [cited by applicant]
Bohdal, Ondrej, Yongxin Yang, and Timothy Hospedales. “Flexible dataset distillation: Learn labels instead of images.” arXiv preprint arXiv:2006.08572 (2020). [cited by applicant]
Borsos, Zalán, Mojmir Mutny, and Andreas Krause. “Coresets via bilevel optimization for continual learning and streaming.” Advances in neural information processing systems 33 (2020): 14879-14890. [cited by applicant]
Castro, Francisco M., et al. “End-to-end incremental learning.” Proceedings of the European conference on computer vision (ECCV). 2018. [cited by applicant]
Chen, Yutian, Max Welling, and Alex Smola. “Super-samples from kernel herding.” arXiv preprint arXiv:1203.3472 (2012). [cited by applicant]
Chizat, Lenaic, Edouard Oyallon, and Francis Bach. “On lazy training in differentiable programming.” Advances in neural information processing systems 32 (2019). [cited by applicant]
Cohen, M. B., Musco, C., & Musco, C. (2017). Input sparsity time low-rank approximation via ridge leverage score sampling. In Proceedings of the Twenty-Eighth Annual ACM-SIAM Symposium on Discrete Algorithms (pp. 1758-1… [cited by applicant]
Daniely, Amit, Roy Frostig, and Yoram Singer. “Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity.” Advances in neural information processing systems 29 (2016). [cited by applicant]
Garriga-Alonso, Adrià, Carl Edward Rasmussen, and Laurence Aitchison. “Deep convolutional networks as shallow gaussian processes.” arXiv preprint arXiv:1808.05587 (2018). [cited by applicant]
Ghorbani, Behrooz, et al. “When do neural networks outperform kernel methods?.” Advances in Neural Information Processing Systems 33 (2020): 14820-14830. [cited by applicant]
Hampel, Frank R. “The influence curve and its role in robust estimation.” Journal of the american statistical association 69.346 (1974): 383-393. [cited by applicant]
Har-Peled, S., & Mazumdar, S. (Jun. 2004). On coresets for k-means and k-median clustering. In Proceedings of the thirty-sixth annual ACM symposium on Theory of computing (pp. 291-300). [cited by applicant]
Hasani, Ramin, et al. “Response characterization for auditing cell dynamics in long short-term memory networks.” 2019 International Joint Conference on Neural Networks (IJCNN). IEEE, 2019. [cited by applicant]
Hron, Jiri, et al. “Infinite attention: NNGP and NTK for deep attention networks.” International Conference on Machine Learning. PMLR, 2020. [cited by applicant]
Hu, Wei, Zhiyuan Li, and Dingli Yu. “Simple and effective regularization methods for training on noisily labeled data with generalization guarantee.” arXiv preprint arXiv:1905.11368 (2019). [cited by applicant]
Hutter, Frank, Lars Kotthoff, and Joaquin Vanschoren. Automated machine learning: methods, systems, challenges. Springer Nature, 2019. [cited by applicant]
Jacot, Arthur, Franck Gabriel, and Clément Hongler. “Neural tangent kernel: Convergence and generalization in neural networks.” Advances in neural information processing systems 31 (2018). [cited by applicant]
Jubran, Ibrahim, Alaa Maalouf, and Dan Feldman. “Introduction to coresets: Accurate coresets.” arXiv preprint arXiv:1910.08707 (2019). [cited by applicant]
Kabra, Mayank, Alice Robie, and Kristin Branson. “Understanding classifier errors by examining influential neighbors.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3917-3925. 201… [cited by applicant]
Kaya, Mahmut, and Hasan Şakir Bilge. “Deep metric learning: A survey.” Symmetry 11.9 (2019): 1066. [cited by applicant]
Kim, Been, Rajiv Khanna, and Oluwasanmi O. Koyejo. “Examples are not enough, learn to criticize! criticism for interpretability.” Advances in neural information processing systems 29 (2016). [cited by applicant]
Koh, Pang Wei, and Percy Liang. “Understanding black-box predictions via influence functions.” In International conference on machine learning, pp. 1885-1894. PMLR, 2017. [cited by applicant]
Krizhevsky, Alex, and Geoffrey Hinton. “Learning multiple layers of features from tiny images.” (2009): 7. [cited by applicant]
Lechner, Mathias, Ramin Hasani, Alexander Amini, Thomas A. Henzinger, Daniela Rus, and Radu Grosu. “Neural circuit policies enabling auditable autonomy.” Nature Machine Intelligence 2, No. 10 (2020): 642-652. [cited by applicant]
LeCun, C. Cortes, and C. Burges. Mnist handwritten digit database. ATT Labs [Online]. Available: http://yann.lecun.com/exdb/mnist, 2, 2010. [cited by applicant]
Lee, Jaehoon, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein. “Deep neural networks as gaussian processes.” arXiv preprint arXiv:1711.00165 (2017). [cited by applicant]
Lee, Jaehoon, Samuel Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein. “Finite versus infinite neural networks: an empirical study.” Advances in Neural Information Processin… [cited by applicant]
Loo, Noel, et al. “Efficient dataset distillation using random feature approximation.” Advances in Neural Information Processing Systems 35 (2022): 13877-13891. [cited by applicant]
Lorraine, Jonathan, Paul Vicol, and David Duvenaud. “Optimizing millions of hyperparameters by implicit differentiation.” In International conference on artificial intelligence and statistics, pp. 1540-1552. PMLR, 2020. [cited by applicant]
Lucic, Mario, Matthew Faulkner, Andreas Krause, and Dan Feldman. “Training gaussian mixture models at scale via coresets.” The Journal of Machine Learning Research 18, No. 1 (2017): 5885-5909. [cited by applicant]
Maclaurin, Dougal, David Duvenaud, and Ryan Adams. “Gradient-based hyperparameter optimization through reversible learning.” In International conference on machine learning, pp. 2113-2122. PMLR, 2015. [cited by applicant]
Matthews, Alexander G. de G., et al. “Gaussian process behaviour in wide deep neural networks.” arXiv preprint arXiv:1804.11271 (2018). [cited by applicant]
Milios, Dimitrios, Raffaello Camoriano, Pietro Michiardi, Lorenzo Rosasco, and Maurizio Filippone. “Dirichlet-based gaussian processes for large-scale calibrated classification.” Advances in Neural Information Processin… [cited by applicant]
Mirzasoleiman, Baharan, Jeff Bilmes, and Jure Leskovec. “Coresets for data-efficient training of machine learning models.” In International Conference on Machine Learning, pp. 6950-6960. PMLR, 2020. [cited by applicant]
Netzer, Yuval, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng. “Reading digits in natural images with unsupervised feature learning.” (2011). [cited by applicant]
Nguyen, Timothy, Zhourong Chen, and Jaehoon Lee. “Dataset meta-learning from kernel ridge-regression.” arXiv preprint arXiv:2011.00050 (2020). [cited by applicant]
Nguyen, Timothy, Roman Novak, Lechao Xiao, and Jaehoon Lee. “Dataset distillation with infinitely wide convolutional networks.” Advances in Neural Information Processing Systems 34 (2021): 5186-5198. [cited by applicant]
Novak, Roman, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Jiri Hron, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein. “Bayesian deep convolutional networks with many channels are gaussian proce… [cited by applicant]
Novak, Roman, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz. “Neural tangents: Fast and easy infinite neural networks in python.” arXiv preprint arXiv:1912.0280… [cited by applicant]
Novak, Roman, Jascha Sohl-Dickstein, and Samuel S. Schoenholz. “Fast finite width neural tangent kernel.” In International Conference on Machine Learning, pp. 17018-17044. PMLR, 2022. [cited by applicant]
Park, Daniel S., Jaehoon Lee, Daiyi Peng, Yuan Cao, and Jascha Sohl-Dickstein. “Towards nngp-guided neural architecture search.” arXiv preprint arXiv:2011.06006 (2020). [cited by applicant]
Phillips, Jeff M. “Coresets and sketches.” arXiv preprint arXiv:1601.00617 (2016). [cited by applicant]
Platt, John. “Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods.” Advances in large margin classifiers 10, No. 3 (1999): 61-74. [cited by applicant]
Rahimi, Ali, and Benjamin Recht. “Random features for large-scale kernel machines.” Advances in neural information processing systems 20 (2007). [cited by applicant]
Rasmussen, Carl Edward, and Christopher KI Williams. Gaussian processes for machine learning. vol. 1. Cambridge, MA: MIT press, 2006. [cited by applicant]
Rebuffi, Sylvestre-Alvise, Alexander Kolesnikov, Georg Sperl, and Christoph H. Lampert. “icarl: Incremental classifier and representation learning.” In Proceedings of the IEEE conference on Computer Vision and Pattern R… [cited by applicant]