IP Library Granted Patent US 12,347,161
Granted Patent B2
US 12,347,161 · App. 17/868,660 · Granted Jul 1, 2025

System and method for dual-value attention and instance boundary aware regression in computer vision system

Inventors: Qingfeng Liu (San Diego, CA); Mostafa El-Khamy (San Diego, CA)
Assignee: Samsung Electronics Co., Ltd.
G06V10/52G06T7/11G06V10/26G06V10/7715
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,161
App. No.
17/868,660
Granted
Jul 1, 2025
Kind
B2
Abstract

A computer vision system including: one or more processors; and memory including instructions that, when executed by the one or more processors, cause the one or more processors to: determine a semantic multi-scale context feature and an instance multi-scale context feature of an input scene; generate a joint attention map based on the semantic multi-scale context feature and the instance multi-scale context feature; refine the semantic multi-scale context feature and instance multi-scale context feature based on the joint attention map; and generate a panoptic segmentation image based on the refined semantic multi-scale context feature and the refined instance multi-scale context feature.

Claims (71)

1. A computer vision system comprising:

one or more processors; and

memory comprising instructions that, when executed by the one or more processors, cause the one or more processors to:

determine a semantic multi-scale context feature and an instance multi-scale context feature of an input scene;

generate a joint attention map based on the semantic multi-scale context feature and the instance multi-scale context feature;

refine the semantic multi-scale context feature and instance multi-scale context feature based on the joint attention map; and

generate a panoptic segmentation image based on the refined semantic multi-scale context feature and the refined instance multi-scale context feature.

2. The computer vision system of claim 1 , wherein to generate the joint attention map, the instructions further cause the one or more processors to calculate normalized correlations between the semantic multi-scale context feature and the instance multi-scale context feature.

3. The computer vision system of claim 2 , wherein to calculate the normalized correlations, the instructions further cause the one or more processors to:

convolve and reshape the semantic multi-scale context feature;

convolve, reshape, and transpose the instance multi-scale context feature;

apply a matrix multiplication operation between the reshaped semantic multi-scale context feature and the transposed instance multi-scale context feature; and

apply a softmax function to the output of the matrix multiplication operation.

4. The computer vision system of claim 1 , wherein to refine the semantic multi-scale context feature, the instructions further cause the one or more processors to:

convolve and reshape the semantic multi-scale context feature;

apply a matrix multiplication operation between the reshaped semantic multi-scale context feature and the joint attention map; and

reshape the output of the matrix multiplication operation.

5. The computer vision system of claim 1 , wherein to refine the instance multi-scale context feature, the instructions further cause the one or more processors to:

convolve and reshape the instance multi-scale context feature;

apply a matrix multiplication operation between the reshaped instance multi-scale context feature and the joint attention map; and

reshape the output of the matrix multiplication operation.

6. The computer vision system of claim 1 , wherein to generate the panoptic segmentation image, the instructions further cause the one or more processors to:

generate a semantic segmentation prediction based on the refined semantic multi-scale context feature;

generate an instance segmentation prediction based on the refined instance multi-scale context feature; and

fuse the semantic segmentation prediction with the instance segmentation prediction.

7. The computer vision system of claim 6 , wherein the instance segmentation prediction comprises:

an instance center prediction; and

an instance center offset prediction.

8. The computer vision system of claim 7 , wherein the instance center offset prediction is trained accordingly to an instance boundary aware regression loss function that applies more weights and penalties to boundary pixels defined in a ground truth boundary map than those applied to other pixels.

9. The computer vision system of claim 8 , wherein the ground truth boundary map is generated by comparing a Euclidean distance between a pixel and a boundary with a threshold distance.

10. A panoptic segmentation method, comprising:

determining a semantic multi-scale context feature and an instance multi-scale context feature of an input scene;

generating a joint attention map based on the semantic multi-scale context feature and the instance multi-scale context feature;

refining the semantic multi-scale context feature and instance multi-scale context feature based on the joint attention map; and

generating a panoptic segmentation image based on the refined semantic multi-scale context feature and the refined instance multi-scale context feature.

11. The method of claim 10 , wherein the generating of the joint attention map comprises calculating normalized correlations between the semantic multi-scale context feature and the instance multi-scale context feature.

12. The method of claim 11 , wherein the calculating of the normalized correlations comprises:

convolving and reshaping the semantic multi-scale context feature;

convolving, reshaping, and transposing the instance multi-scale context feature;

applying a matrix multiplication operation between the reshaped semantic multi-scale context feature and the transposed instance multi-scale context feature; and

applying a softmax function to the output of the matrix multiplication operation.

13. The method of claim 10 , wherein the refining of the semantic multi-scale context feature comprises:

convolving and reshaping the semantic multi-scale context feature;

applying a matrix multiplication operation between the reshaped semantic multi-scale context feature and the joint attention map; and

reshaping the output of the matrix multiplication operation.

14. The method of claim 10 , wherein the refining of the instance multi-scale context feature comprises:

convolving and reshaping the instance multi-scale context feature;

applying a matrix multiplication operation between the reshaped instance multi-scale context feature and the joint attention map; and

reshaping the output of the matrix multiplication operation.

15. The method of claim 10 , wherein the generating of the panoptic segmentation image comprises:

generating a semantic segmentation prediction based on the refined semantic multi-scale context feature;

generating an instance segmentation prediction based on the refined instance multi-scale context feature; and

fusing the semantic segmentation prediction with the instance segmentation prediction.

16. The method of claim 15 , wherein the instance segmentation prediction comprises:

an instance center prediction; and

an instance center offset prediction.

17. The method of claim 16 , wherein the instance center offset prediction is trained accordingly to an instance boundary aware regression loss function that applies more weights and penalties to boundary pixels defined in a ground truth boundary map than those applied to other pixels.

18. The method of claim 17 , wherein the ground truth boundary map is generated by comparing a Euclidean distance between a pixel and a boundary with a threshold distance.

19. A panoptic segmentation system comprising:

one or more processors; and

memory comprising instructions that, when executed by the one or more processors, cause the one or more processors to:

extract a semantic multi-scale context feature of an input scene;

extract an instance multi-scale context feature of the input scene;

generate a joint attention map based on the semantic multi-scale context feature and the instance multi-scale context feature;

refine the semantic multi-scale context feature based on the joint attention map;

refine the instance multi-scale context feature based on the joint attention map;

predict a semantic class label based on the refined semantic multi-scale context feature;

predict an instance center and an instance center offset based on the refined instance multi-scale context feature; and

generate a panoptic segmentation image based on the semantic class label, the instance center, and the instance center offset.

20. The panoptic segmentation system of claim 19 , wherein the predicting of the instance center and the instance center offset is based on an instance boundary aware regression loss function that applies more weights and penalties to boundary pixels defined in a ground truth boundary map than those applied to other pixels, and

wherein the ground truth boundary map is generated by comparing a Euclidean distance between a pixel and a boundary with a threshold distance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2022
From: LIU, QINGFENG; EL-KHAMY, MOSTAFA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061226/0488 →
Continuity (2)
Provisional Application 63311246 · Feb 17, 2022
Related Publication 20230260247A1 · Aug 17, 2023
References Cited (28)
US 8107726B2 · Xu et al. · 2012 [cited by applicant]
US 9443314B1 · Huang et al. · 2016 [cited by applicant]
US 9697443B2 · Chen et al. · 2017 [cited by applicant]
US 10424064B2 · Price et al. · 2019 [cited by applicant]
US 10803328B1 · Bai et al. · 2020 [cited by applicant]
US 10997433B2 · Xu et al. · 2021 [cited by applicant]
US 11200667B2 · Lay et al. · 2021 [cited by applicant]
US 20180108138A1 · Kluckner et al. · 2018 [cited by applicant]
US 20180174311A1 · Kluckner et al. · 2018 [cited by applicant]
US 20190286153A1 · Rankawat et al. · 2019 [cited by applicant]
US 20200278408A1 · Sung et al. · 2020 [cited by applicant]
US 20200302224A1 · Jaganathan et al. · 2020 [cited by applicant]
US 20200349711A1 · Duke et al. · 2020 [cited by applicant]
US 20210027098A1 · Ge et al. · 2021 [cited by applicant]
US 20210150722A1 · Homayounfar et al. · 2021 [cited by applicant]
US 20210279950A1 · Phalak · 2021 [cited by applicant]
US 20210326656A1 · Lee · 2021 [cited by examiner]
US 20210374968A1 · Martinez et al. · 2021 [cited by applicant]
US 20230104262A1 · Lin · 2023 [cited by examiner]
US 20240212164A1 · Li · 2024 [cited by examiner]
WO WO2018187622A1 · 2018 [cited by applicant]
WO WO2021068182A1 · 2021 [cited by applicant]
Li, Yanwei, et al. “Attention-guided unified network for panoptic segmentation.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019. (Year: 2019). [cited by examiner]
Cheng, Bowen, et al., “Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, 16 pages. [cited by applicant]
Gao, Naiyu, et al., “Learning Category- and Instance-Aware Pixel Embedding for Fast Panoptic Segmentation,” IEEE Transactions on Image Processing, vol. 30, 2021, pp. 6013-6023. [cited by applicant]
Liu, Xiaolong, et al., “CASNet: Common Attribute Support Network for image instance and panoptic segmentation,” 2020 25th International Conference on Pattern Recognition (ICPR), 2020, pp. 8469-8475. [cited by applicant]
Wang, Haochen, et al. “Pixel Consensus Voting for Panoptic Segmentation,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 9464-9473. [cited by applicant]
Zhang, Weitong, et al., “End-To-End Panoptic Segmentation With Pixel-Level Non-Overlapping Embedding,” 2019 IEEE International Conference on Multimedia and Expo (ICME), 2019, pp. 976-981. [cited by applicant]