IP Library › Granted Patent US 12,346,812
Granted Patent B2
US 12,346,812 · App. 17/784,620 · Granted Jul 1, 2025

Query optimization for deep convolutional neural network inferences

Inventors: Arun Kumar (San Diego, CA); Supun Nakandala (San Diego, CA)
Assignee: The Regents of the University of California
G06N3/08G06F16/24535G06V10/751G06V10/759G06V10/82G06V20/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,346,812
App. No.
17/784,620
Granted
Jul 1, 2025
Kind
B2
Abstract

A method may include generating views materializing tensors generated by a convolutional neural network operating on an image. Determining the outputs of the convolutional neural network operating on the image with a patch occluding various portions of the image. The outputs being determined by generating queries on the views that performs, based at least on the changes associated with occluding different portions of the image, partial re-computations of the views. A heatmap may be generated based on the outputs of the convolutional neural network. The heatmap may indicate the quantities to which the different portions of the image contribute to the output of the convolutional neural network operating on the image. Related systems and articles of manufacture, including computer program products, are also provided.

Claims (48)

1. A system, comprising:

at least one processor; and

at least one memory including program code which when executed by the at least one processor provides operations comprising:

generating one or more views materializing one or more tensors generated by a convolutional neural network operating on an image;

determining a first output of the convolutional neural network operating on the image with a patch occluding a first portion of the image, the first output being determined by generating a first query on the one or more views, the first query performing, based at least on a first change associated with occluding the first portion of the image, a first partial re-computation of the one or more views;

generating, based at least on the first output, a first heatmap indicating a first quantity to which the first portion of the image contributes to an output of the convolutional neural network operating on the image;

generating, at a first stride size, a second heatmap;

identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and

determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.

2. The system of claim 1 , further comprising:

determining a second output of the convolutional neural network operating on the image with the patch occluding a second portion of the image, the second output being determined by generating a second query on the one or more views, the second query performing, based at least on a second change associated with occluding the second portion of the image, a second partial re-computation of the one or more views; and

generating, based at least on the second output, the first heatmap to further indicate a second quantity to which the second portion of the image contributes to the output of the convolutional neural network operating on the image.

3. The system of claim 2 , wherein the performing of the first query and the second query is batched.

4. The system of claim 1 , wherein the first change corresponds to a size of the patch occluding the first portion of the image, a size of a filter kernel associated with the convolutional neural network, and a size of a stride associated with the filter kernel.

5. The system of claim 1 , wherein the first partial re-computation is performed by at least propagating the first change through successive layers of the convolutional neural network.

6. The system of claim 5 , further comprising:

limiting, to a threshold value, a quantity of elements in each layer of the convolutional neural network affected by the propagation of the first change, the limiting generating an approximation of an output at each layer of the convolutional neural network.

7. The system of claim 6 , further comprising:

generating, based on one or more sample images, an approximate heatmap and an exact heatmap at a plurality of different threshold values; and

determining, based at least on an index measuring a difference between the approximate heatmap and the exact heatmap, the threshold value.

8. The system of claim 1 , wherein the threshold fraction is specified by one or more user inputs, and wherein the first stride size is determined based on a target speedup specified by the one or more user inputs.

9. The system of claim 1 , wherein the first partial re-computation of the one or more views is limited to the first change associated with occluding the first portion of the image.

10. A computer-implemented method, comprising:

generating one or more views materializing one or more tensors generated by a convolutional neural network operating on an image;

determining a first output of the convolutional neural network operating on the image with a patch occluding a first portion of the image, the first output being determined by generating a first query on the one or more views, the first query performing, based at least on a first change associated with occluding the first portion of the image, a first partial re-computation of the one or more views;

generating, based at least on the first output, a first heatmap indicating a first quantity to which the first portion of the image contributes to an output of the convolutional neural network operating on the image;

generating, at a first stride size, a second heatmap;

identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and

determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.

11. The method of claim 10 , further comprising:

determining a second output of the convolutional neural network operating on the image with the patch occluding a second portion of the image, the second output being determined by generating a second query on the one or more views, the second query performing, based at least on a second change associated with occluding the second portion of the image, a second partial re-computation of the one or more views; and

generating, based at least on the second output, the first heatmap to further indicate a second quantity to which the second portion of the image contributes to the output of the convolutional neural network operating on the image.

12. The method of claim 11 , wherein the performing of the first query and the second query is batched.

13. The method of claim 10 , wherein the first change corresponds to a size of the patch occluding the first portion of the image, a size of a filter kernel associated with the convolutional neural network, and a size of a stride associated with the filter kernel.

14. The method of claim 10 , wherein the first partial re-computation is performed by at least propagating the first change through successive layers of the convolutional neural network.

15. The method of claim 14 , further comprising:

limiting, to a threshold value, a quantity of elements in each layer of the convolutional neural network affected by the propagation of the first change, the limiting generating an approximation of an output at each layer of the convolutional neural network.

16. The method of claim 15 , further comprising:

generating, based on one or more sample images, an approximate heatmap and an exact heatmap at a plurality of different threshold values; and

determining, based at least on an index measuring a difference between the approximate heatmap and the exact heatmap, the threshold value.

17. The method of claim 10 , wherein the first partial re-computation of the one or more views is limited to the first change associated with occluding the first portion of the image.

18. A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising:

generating one or more views materializing one or more tensors generated by a convolutional neural network operating on an image;

determining an output of the convolutional neural network operating on the image with a patch occluding a portion of the image, the output being determined by generating a query on the one or more views, the query performing, based at least on a change associated with occluding the portion of the image, a partial re-computation of the one or more views;

generating, based at least on the output, a heatmap indicating a quantity to which the portion of the image contributes to an output of the convolutional neural network operating on the image;

generating, at a first stride size, a second heatmap;

identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and

determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 10, 2022
From: KUMAR, ARUN; NAKANDALA, SUPUN
To: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
Reel/Frame 060173/0370 →
Continuity (2)
Provisional Application 62971862 · Feb 7, 2020
Related Publication 20230042004A1 · Feb 9, 2023
References Cited (54)
US 9721203B1 · Young · 2017 [cited by examiner]
US 11816185B1 · Roth · 2023 [cited by examiner]
US 20180025257A1 · van den Oord · 2018 [cited by examiner]
US 20190156274A1 · Fisher et al. · 2019 [cited by applicant]
US 20210135625A1 · Deng · 2021 [cited by examiner]
US 20210158096A1 · Sinha · 2021 [cited by examiner]
US 20210397170A1 · Zhou · 2021 [cited by examiner]
US 20220391621A1 · Chen · 2022 [cited by examiner]
US 20230048386A1 · Wang · 2023 [cited by examiner]
Ordookhanians et al. “Demonstration of Krypton: optimized CNN inference for occlusion-based deep CNN explanations”, Proceedings of the VLDB Endowment, vol. 12, Issue 12 pp. 1894-1897, Aug. 1, 2019 (Year: 2019). [cited by examiner]
Muneeb ul Hassan, “VGG16—Convolutional Network for Classification and Detection”, Nov. 20, 2018, Retrieved from Internet at < https://neurohive.io/en/popular-networks/vgg16/ > (Year: 2018). [cited by examiner]
“Cafee Model Zoo.” (Available at http://caffe.berkeleyvision.org/model_zoo.html). [cited by applicant]
“Torch Vison Models.” (Available at http://web.archive.org/web/20211026122306/https://github.com/pytorch/vision/tree/master/torchvision/models). [cited by applicant]
Arbabzadah, F. et al. “Identifying individual facial expressions by deconstructing a neural network.” In German Conference on Pattern Recognition, arXiv preprint arXiv:1606.07285, 2016. [cited by applicant]
Buckler, M. et al. “EVA2: Exploiting Temporal Redundancy in Live Computer Vision.” arXiv preprint arXiv:1803.06312, 2018. [cited by applicant]
Cavigelli, L. et al. “CBinfer: Change-Based Inference for Convolutional Neural Networks on Video Data.” In Proceedings of the 11th International Conference on Distributed Smart Cameras, pp. 1-8. ACM, 2017. [cited by applicant]
Chetlur, S. et al. “cuDNN: Efficient Primitives for Deep Learning.” arXiv preprint arXiv:1410.0759, 2014. [cited by applicant]
Chirkova, R. et al., “Materialized views,” Foundations and Trends in Databases, 4(4):295-405, 2012. [cited by applicant]
Deng, J. et al. “ImageNet: A Large-Scale Hierarchical Image Database.” In Computer Vision and Pattern Recognition, 2009. CVPR 2009. WE Conference on, pp. 248-255. Ieee, 2009. [cited by applicant]
De Vries, S.E.J. et al. “The projective field of a retinal amacrine cell.” Journal of Neuroscience, 31(23):8595-8604, 2011. [cited by applicant]
Garofalakis, M.N. et al. “Approximate query processing: Taming the terabytes.” In VLDB, pp. 343-352, 2001. [cited by applicant]
Goodfellow, I. et al. “Deep learning, vol. 1.” MIT press Cambridge, 2016. [cited by applicant]
Gupta, A. et al. “Maintenance of materialized views: Problems, techniques, and applications.” IEEE Data Eng. Bull., 18(2):3-18, 1995. [cited by applicant]
He, K. et al., “Deep residual learning for image recognition,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770-778, 2016. [cited by applicant]
He, Y. “Channel Pruning for Accelerating Very Deep Neural Networks.” Proceedings of the IEEE international conference on computer vision. 2017. [cited by applicant]
Islam, M.T. et al. “Abnormality detection and localization in chest x-rays using deep convolutional neural networks.” arXiv preprint arXiv:1705.09850, 2017. [cited by applicant]
Jung, K.-H. et al. “Deep learning for medical image analysis: Applications to computed tomography and magnetic resonance imaging.” Hanyang Medical Reviews, 37(2):61-70, 2017. [cited by applicant]
Kang, D. et al. “NoScope: Optimizing Neural Network Queries over Video at Scale.” Proceedings of the VLDB Endowment, 10(11):1586-1597, 2017. [cited by applicant]
Kermany, D.S. et al. “Identifying medical diagnoses and treatable diseases by image-based deep learning.” Cell, 172(5):1122-1131, 2018. [cited by applicant]
Kingma, D. P. et al., “Adam: A method for stochastic optimization.” arXiv preprint arXiv:1412.6980, 2014. [cited by applicant]
Le, W. et al. “Scalable Multi-Query Optimization for SPARQL.” In Data Engineering (ICDE), 2012 IEEE 28th International Conference on, pp. 666-677. IEEE, 2012. [cited by applicant]
Lee, K.J. “AI device for detechting diabetic retinopathy earns swift FDA approval.” American Academy of Opthamology (Available at https://www.aao.org/headline/first-ai-screen-diabetic-retinopathy-approved-by-f) Apr. 12,… [cited by applicant]
Levy, A.L. et al. “Answering queries using views.” In Proceedings of the fourteenth ACM SIGACT-SIGMOD-SIGART symposium on Principles of database systems, pp. 95-104. ACM, 1995. [cited by applicant]
Luo, W. et al. “Understanding the effective receptive field in deep convolutional neural networks.” In Advances in neural information processing systems, pp. 4898-4906, 2016. [cited by applicant]
Marko, K. “Radiologists are often in short supply and overworked—deep learning to the rescue.” (Available at https://diginomica.com/radiologists-often-short-supply-overworked-deep-learning-rescue). Dec. 19, 2017. [cited by applicant]
Miller, T. “Explanation in artificial intelligence: Insights from the social sciences.” arXiv preprint arXiv:1706.07269, 2017. [cited by applicant]
Mohanty, S.P. et al. “Using deep learning for image-based plant disease detection.” Frontiers in plant science, 7:1419, 2016. [cited by applicant]
Moons, B. et al. “A 0.3-2.6 tops/w precision-scalable processor for real-time large-scale convnets.” In VLSI Circuits (VLSI-Circuits), 2016 IEEE Symposium on, pp. 1-2. IEEE, 2016. [cited by applicant]
Motamedi, M. et al. “Resource-Scalable CNN Synthesis for IoT Applications.” arXiv preprint arXiv:1901.00738, 2018. [cited by applicant]
Nikolic, M. et al. “Linview: incremental view maintenance for complex analytical queries.” In Proceedings of the 2014 ACM SIGMOD international conference on Management of data, pp. 253-264. ACM, 2014. [cited by applicant]
Ribeiro, M.T. et al. “Why Should I Trust You? Explaining the Predictions of Any Classifier.” NAACL HLT 2016 (2016): 97. [cited by applicant]
Russakovsky, O. et al. “Imagenet large scale visual recognition challenge.” International Journal of Computer Vision, 115(3):211-252, 2015. [cited by applicant]
Sellis, T.K. “Multiple-Query Optimization.” ACM Transactions on Database Systems (TODS), 13(1):23-52, 1988. [cited by applicant]
Selvaraju, R.R. et al. “Grad-cam: Visual explanations from deep networks via gradient-based localization.” In 2017 IEEE International Conference on Computer Vision (ICCV), pp. 618-626. IEEE, 2017. [cited by applicant]
Simonyan, K. et al. “Deep inside convolutional networks: Visualising image classification models and saliency maps.” arXiv preprint arXiv:1312.6034, 2013. [cited by applicant]
Simonyan, K. et al., “Very deep convolutional networks for large-scale image recognition.” arXiv preprint arXiv:1409.1556, 2014. [cited by applicant]
Sundararajan, M. et al. “Axiomatic attribution for deep networks.” arXiv preprint arXiv:1703.01365, 2017. [cited by applicant]
Szegedy, C. et al. “Rethinking the inception architecture for computer vision.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2818-2826, 2016. [cited by applicant]
Voigt, P. et al. “The EU General Data Protection Regulation (GDPR), vol. 18.” Springer, 2017. [cited by applicant]
Wang, Y. et al. “Deep neural networks are more accurate than humans at detecting sexual orientation from facial images.” Journal of personality and social psychology, 114(2):246, 2018. [cited by applicant]
Wang, Z. et al. “Image quality assessment: from error visibility to structural similarity.” IEEE transactions on image processing, 13(4):600-612, 2004. [cited by applicant]
Zeiler, M.D. et al. “Visualizing and understanding convolutional networks.” In European conference on computer vision, pp. 818-833. Springer, 2014. [cited by applicant]
Zhao, W. et al. “Incremental view maintenance over array data.” In Proceedings of the 2017 ACM International Conference on Management of Data, pp. 139-154. ACM, 2017. [cited by applicant]
Zintgraf, L. M. et al. “Visualizing deep neural network decisions: Prediction difference analysis.” arXiv preprint arXiv:1702.04595, 2017. [cited by applicant]