IP Library › Granted Patent US 12,694,661
Granted Patent B2
US 12,694,661 · App. 18/482,085 · Granted Jul 28, 2026

Explainable visual attention for deep learning

Inventors: Utkarsh Agrawal (Bengaluru, IN); Bipul Das (Chennai, IN); Prasad Sudhakara Murthy (Bengaluru, IN)
Assignee: GE PRECISION HEALTHCARE LLC
G06V10/82G06V10/7715G16H30/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,694,661
App. No.
18/482,085
Filed
Oct 6, 2023
Granted
Jul 28, 2026
Kind
B2
Art Unit
2667
USPC
382/155
Abstract

Systems or techniques that facilitate explainable visual attention for deep learning are provided. In various embodiments, a system can access a medical image generated by a medical imaging scanner. In various aspects, the system can perform, via execution of a deep learning neural network, an inferencing task on the medical image. In various instances, the deep learning neural network can receive as input the medical image and can produce as output both an inferencing task result and an attention map indicating on which pixels or voxels of the medical image the deep learning neural network focused in generating the inferencing task result.

Claims (48)

1 . A system, comprising:

a processor that executes computer-executable instructions stored in a non-transitory computer-readable memory, wherein execution of the computer-executable instructions causes the processor to:

access a deep learning neural network that comprises a first processing channel and a second processing channel;

access an annotated training dataset that comprises a training medical image, a ground-truth inferencing task result corresponding to the training medical image, and a ground-truth attention map indicating which pixels or voxels of the training medical image should be focused on to produce the ground-truth inferencing task result;

randomly initialize first trainable internal parameters of the first processing channel and second trainable internal parameters of the second processing channel;

execute the deep learning neural network on the training medical image, thereby causing the first processing channel to produce a first output and causing the second processing channel to produce a second output;

compute a first loss between the ground-truth inferencing task result and the first output;

weight the first loss via point-wise multiplication with the ground-truth attention map;

update the first trainable internal parameters of the first processing channel, via backpropagation driven by the weighted first loss;

compute a second loss between the ground-truth attention map and the second output; and

update the second trainable internal parameters of the second processing channel, via backpropagation driven by the second loss.

2 . The system of claim 1 , wherein execution of the computer-executable instructions further causes the processor to:

access a medical image generated by a medical imaging scanner;

execute, after training, the deep learning neural network on the medical image, wherein the deep learning neural network receives the medical image, wherein the first processing channel produces an inferencing task result, and wherein the second processing channel produces an attention map indicating on which pixels or voxels of the medical image the deep learning neural network focused in generating the inferencing task result; and

visually render the inferencing task result and the attention map on an electronic display.

3 . The system of claim 1 wherein the ground-truth attention map is generated based on computing a difference array between the training medical image and the ground-truth inferencing task result.

4 . The system of claim 1 , wherein the ground-truth attention map is generated via execution of an edge detector on the training medical image.

5 . The system of claim 1 , wherein the ground-truth attention map is generated based on computing a difference array between the ground-truth inferencing task result and an output array produced by the first processing channel during a previous training epoch.

6 . A computer-implemented method, comprising:

accessing, by a device operatively coupled to a processor, a deep learning neural network that comprises a first processing channel and a second processing channel;

accessing, by the device, an annotated training dataset that comprises a training medical image, a ground-truth inferencing task result corresponding to the training medical image, and a ground-truth attention map indicating which pixels or voxels of the training medical image should be focused on to produce the ground-truth inferencing task result;

randomly initializing, by the device, first trainable internal parameters of the first processing channel and second trainable internal parameters of the second processing channel;

executing, by the device, the deep learning neural network on the training medical image, thereby causing the first processing channel to produce a first output and causing the second processing channel to produce a second output;

computing, by the device, a first loss between the ground-truth inferencing task result and the first output;

weighting, by the device, the first loss via point-wise multiplication with the ground-truth attention map;

updating, by the device, the first trainable internal parameters of the first processing channel, via backpropagation driven by the weighted first loss;

computing, by the device, a second loss between the ground-truth attention map and the second output; and

updating, by the device, the second trainable internal parameters of the second processing channel, via backpropagation driven by the second loss.

7 . The computer-implemented method of claim 6 , further comprising:

accessing, by the device, a medical image generated by a medical imaging scanner;

executing, by the device and after training, the deep learning neural network on the medical image, wherein the deep learning neural network receives the medical image, wherein the first processing channel produces an inferencing task result, and wherein the second processing channel produces an attention map indicating on which pixels or voxels of the medical image the deep learning neural network focused in generating the inferencing task result; and

visually rendering, by the device, the inferencing task result and the attention map on an electronic display.

8 . The computer-implemented method of claim 6 , wherein the ground-truth attention map is generated based on computing a difference array between the training medical image and the ground-truth inferencing task result.

9 . The computer-implemented method of claim 6 , wherein the ground-truth attention map is generated via execution of an edge detector on the training medical image.

10 . The computer-implemented method of claim 6 , wherein the ground-truth attention map is generated based on computing a difference array between the ground-truth inferencing task result and an output array produced by the first processing channel during a previous training epoch.

11 . A computer program product for facilitating explainable visual attention for deep learning, the computer program product comprising a non-transitory computer-readable memory having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to:

access a deep learning neural network that comprises a first processing channel and a second processing channel;

access an annotated training dataset that comprises a training medical image, a ground-truth inferencing task result corresponding to the training medical image, and a ground-truth attention map indicating which pixels or voxels of the training medical image should be focused on to produce the ground-truth inferencing task result;

randomly initialize first trainable internal parameters of the first processing channel and second trainable internal parameters of the second processing channel;

execute the deep learning neural network on the training medical image, thereby causing the first processing channel to produce a first output and causing the second processing channel to produce a second output;

compute a first loss between the ground-truth inferencing task result and the first output;

weight the first loss via point-wise multiplication with the ground-truth attention map;

update the first trainable internal parameters of the first processing channel, via backpropagation driven by the weighted first loss;

compute a second loss between the ground-truth attention map and the second output; and

update the second trainable internal parameters of the second processing channel, via backpropagation driven by the second loss.

12 . The computer program product of claim 11 , wherein the ground-truth attention map is generated based on computing a difference array between the training medical image and the ground-truth inferencing task result.

13 . The computer program product of claim 11 , wherein the ground-truth attention map is generated via execution of an edge detector on the training medical image.

14 . The computer program product of claim 11 , wherein the ground-truth attention map is generated based on computing a difference array between the ground-truth inferencing task result and an output array produced by the first processing channel during a previous training epoch.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: AGRAWAL, UTKARSH; DAS, BIPUL; MURTHY, PRASAD SUDHAKARA
To: GE PRECISION HEALTHCARE LLC
Reel/Frame 065143/0674 →
Continuity (1)
Related Publication 20250118062A1 · Apr 10, 2025
References Cited (24)
US 10997433B2 · Xu · 2021 [cited by examiner]
US 20170236057A1 · Lane · 2017 [cited by examiner]
US 20170270664A1 · Hoogi · 2017 [cited by examiner]
US 20190332932A1 · Sivaraman · 2019 [cited by examiner]
US 20190365341A1 · Chan · 2019 [cited by examiner]
US 20220058803A1 · Bhattacharya · 2022 [cited by examiner]
US 20220101635A1 · Koivisto · 2022 [cited by examiner]
US 20220284570A1 · Tan · 2022 [cited by examiner]
US 20220327810A1 · Nagori · 2022 [cited by examiner]
US 20220383489A1 · Shi · 2022 [cited by examiner]
US 20230137369A1 · Balicki · 2023 [cited by examiner]
CN 115393246A · 2022 [cited by applicant]
JP 2021022368A · 2021 [cited by applicant]
Watanabe et al., “Improving Disease Classification Performance and Explainability of Deep Learning Models in Radiology with Heatmap Generators,” Jun. 2022, https://arxiv.org/pdf/2207.00157 (Year: 2022). [cited by examiner]
Qurri et al., “Improved UNet with Attention for Medical Image Segmentation,” Sensors 2023, 23, 8589. https://doi.org/10.3390/s23208589 (Year: 2023). [cited by examiner]
Huang, Z. et al. | “A novel tongue segmentation method based on improved U-Net.” Neurocomputing, vol. 500, Aug. 21, 2022, pp. 73-89, 17 pages. [cited by applicant]
Rengasamy, D. et al. | “Deep Learning with Dynamically Weighted Loss Function for Sensor-Based Prognostics and Health Management.” Sensors (Basel). Jan. 28, 2020;20(3):723. doi: 10.3390/s20030723. PMID: 32012944; PMCID:… [cited by applicant]
Akino Watanabe et al: “Improving Disease Classification Performance and Explainability of Deep Learning Models in Radiology with Heatmap Generators”, arxiv.org, Cornell University Library, 201OLIN Library Cornell Univer… [cited by applicant]
CN 115393246 English Abstract; Espacenet.com; 1 page. [cited by applicant]
Cristiano Patr\'icio et al: “Explainable Deep Learning Methods in Medical Imaging Diagnosis: A Survey”, arxiv.org, Cornell University Library, 201, Olin Library Cornell University Ithaca, NY, XP091245323, * the whole do… [cited by applicant]
EP application 242007391 filed Sep. 17, 2024—extended Search Report issued Feb. 21, 2025; 25 pages. [cited by applicant]
Xiaozheng Xie et al: “A Survey on Domain Knowledge Powered Deep Learning for Medical Image Analysis”, arxiv.org, Cornell University Library, Olin Library Cornell University Ithaca, NY, 14853, XP081652902 Sec. 2.3.3; * a… [cited by applicant]
JP application 2024-161572 filed Sep. 19, 2024—Office Action issued on Dec. 3, 2025; Machine Translation; 4 pages. [cited by applicant]
JP2021-022368 English Abstract, Espacenet.com Feb. 26, 2026; 1 page. [cited by applicant]