IP Library › Granted Patent US 12,417,261
Granted Patent B2
US 12,417,261 · App. 17/479,655 · Granted Sep 16, 2025

Visualization system and method for interpretation and diagnosis of deep neural networks

Inventors: Panpan Xu (Santa Clara, CA); Liu Ren (Saratoga, CA); Zhenge Zhao (Tuscon, AZ)
Assignee: Robert Bosch GmbH
G06F18/41G06F3/0482G06F3/0484G06T7/10G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,261
App. No.
17/479,655
Granted
Sep 16, 2025
Kind
B2
Abstract

A computer-implemented method includes receiving one or more images from one or more sensors, creating one or more image patches utilizing the one or more images, creating one or more latent representations from the one or more image patches via a neural network, outputting, to a concept extractor network, the one or more latent representations utilizing the one or more image patches, defining one or more scores associated with the one or more latent representations, and outputting one or more scores associated with the one or more image patches utilizing at least the concept extractor network.

Claims (40)

1. A computer-implemented method, comprising:

receiving one or more images from one or more sensors;

creating one or more image patches utilizing the one or more images, wherein the patches are segmented utilizing fixed window sizes or super-pixel segmentation algorithms;

creating one or more latent representations from the one or more image patches via a neural network;

outputting, to a concept extractor network, the one or more latent representations utilizing the one or more image patches, wherein the concept extractor network is in communication with the neural network and configured to identify patches that contain a common concept, and the concept extractor network is further trainable via a user interface configured to manually label one or more informative examples associated with the one or more images configured to prioritize impact training on the neural network;

defining one or more scores associated with the one or more latent representations, wherein the scores are defined by the concept extractor network and the scores indicate one or more concepts retrieved from the one or more image patches, wherein scores of the one or more informative examples that are manually labeled are prioritized; and

outputting, to the user interface, one or more scores associated with the one or more image patches utilizing at least the concept extractor network, wherein the user interface includes scores associated with a first model and a second model, wherein the second model is different from the first model based on at least weights or architecture.

2. The method of claim 1 , wherein the method includes outputting the user interface at a display, wherein the user interface includes at least an image patch viewer.

3. The method of claim 1 , wherein the method includes outputting the user interface at a display, wherein the user interface includes at least a portion outputting positive samples and negative samples associated with the one or more images.

4. The method of claim 1 , wherein the method includes outputting the user interface at a display, wherein the user interface includes a training sample list that includes one or more savable attributes.

5. The method of claim 1 , wherein the concept extractor network is configured to output a confidence score associated with the one or more latent representations.

6. The method of claim 1 , wherein the concept extractor network includes exactly two layers.

7. The method of claim 1 , wherein the neural network is configured to classify the one or more image patches.

8. The method of claim 1 , wherein the concept extractor network is configured to be trained utilizing at least the one or more images or image patches.

9. A computer-implemented method, comprising:

receiving one or more images from one or more sensors;

creating one or more image patches utilizing the one or more images, wherein the patches are segmented utilizing fixed window sizes or super-pixel segmentation algorithms;

resizing the one or more images to an input scale to create re-sized images;

sending the one or more image patches to a neural network to define one or more latent representations;

outputting, to a concept extractor network, one or more latent representations utilizing the one or more image patches and a plurality of models associated with the concept extractor network, wherein the concept extractor network is configured to identify patches that contain a common concept and the concept extractor network is trainable via a user interface configured to manually label one or more informative examples associated with the one or more images; and

outputting, to a user interface, one or more scores associated with the plurality of models associated one or more attributes of the image patches utilizing the concept extractor network, wherein the user interface includes scores associated with a first model and a second model, wherein the second model is different from the first model based on at least weights or architecture.

10. The computer-implemented method of claim 9 , wherein the method includes outputting one or more attributes associated with the one or more latent representations.

11. The computer-implemented method of claim 10 , wherein the one or more attributes are associated with the plurality of models.

12. The computer-implemented method of claim 10 , wherein the one or more attributes are associated with representative attribute images retrieved from the one or more images.

13. The computer-implemented method of claim 10 , wherein the method includes receiving, from a user, input associated with one or more attributes associated with the image patches.

14. A system, comprising:

a processor in communication with a display, the processor programmed to:

receive one or more images from one or more sensors;

create one or more image patches utilizing the one or more images a wherein the patches are segmented utilizing fixed window sizes or super-pixel segmentation algorithms;

resizing the one or more images to an input scale to create re-sized images;

send the one or more resized images to a neural network to define one or more latent representations utilizing the one or more image patches;

output, to a concept extractor network, one or more latent representations utilizing the one or more image patches and a plurality of models associated with the concept extractor network, wherein the concept extractor network is configured to identify patches that contain a common concept, and the concept extractor network is further trainable via a user interface configured to manually label one or more informative examples associated with the one or more images; and

output, to a user interface on the display, one or attributes associated with the one or more image patches and the plurality of models; and

output one or more scores associated with the plurality of models associated one or more attributes of the image patches utilizing the concept extractor network, wherein the user interface includes scores associated with a first model and a second model, wherein the second model is different from the first model based on at least weights or architecture.

15. The system of claim 14 , wherein the processor is further programmed to output one or more confidence scores associated with the one or more attributes and further programmed to prioritizing scoring of the one or more informative examples that are manually labeled.

16. The system of claim 14 , wherein the user interface is configured to activate one or more filters associated with the one or more image patches.

17. The system of claim 14 , wherein the one or more image patches include a first section with a first set of attributes and a second section with a second set of attributes.

18. The system of claim 14 , wherein the one more sensors includes one or more cameras.

19. The system of claim 14 , wherein the user interface outputs one or filters associated with the one or more attributes.

20. The system of claim 14 , wherein the user interface includes one or more training sampling lists that includes one or more editable attributes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2022
From: XU, PANPAN; REN, LIU; ZHAO, ZHENGE
To: ROBERT BOSCH GMBH
Reel/Frame 060370/0726 →
Continuity (1)
Related Publication 20230085927A1 · Mar 23, 2023
References Cited (10)
US 11017019B1 · Hohwald · 2021 [cited by examiner]
Graziani M. et al., “Concept attribution: Explaining CNN decisions to physician” Computers in Biology and Medicine 123 (2020), pub. Jun. 17, 2020 (Year: 2020). [cited by examiner]
Graziani et al., , “Concept attribution: Explaining CNN decisions to physician”, pub. Jun. 17, 2020, (Year: 2020). [cited by examiner]
Zhengqing, et al., “Concept-based Explanation for Fine-grained Images and Its Application in Infectious Keratitis Classification”, pub. Dec. 16, 2020, (Year: 2020). [cited by examiner]
Gonda et al., , “An Interactive Approach to Train Deep Neural Networks for Segmentation of Neuronal Structures”, Pub. 2017. (Year: 2017). [cited by examiner]
Graziani M. et al., “Concept attribution: Explaining CNN decisions to physician”, pub. Jun. 17, 2020 (Year: 2020). [cited by examiner]
Zhengqing Fang et al., “Concept-based Explanation for Fine-grained Images and Its Application in Infectious Keratitis Classification”, pub. Dec. 16, 2020 (Year: 2020). [cited by examiner]
Gonda et al., “An Interactive Approach to Train Deep Neural Networks for Segmentation of Neuronal Structures”, Pub. Pub. 2017, (Year: 2017). [cited by examiner]
B. Kim, M. Wattenberg, J. Gilmer, C. J. Cai, J. Wexler, F. B. Viegas, and R. Sayres. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). In ICML, 2017. [cited by applicant]
L. van der Maaten and G. Hinton. Visualizing high-dimensional data using t-sne. Journal of Machine Learning Research, 9:2579-2605, 2008. [cited by applicant]