IP Library › Granted Patent US 12,444,055
Granted Patent B2
US 12,444,055 · App. 18/316,823 · Granted Oct 14, 2025

Convolution and transformer-based image segmentation

Inventors: Xin Li (San Diego, CA); Jiancheng Lyu (San Diego, CA); Yingyong Qi (San Diego, CA)
Assignee: QUALCOMM Incorporated
G06T7/11G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,055
App. No.
18/316,823
Granted
Oct 14, 2025
Kind
B2
Abstract

Techniques are provided for image processing. For instance, a process can include obtaining an image; extracting a first set of features at a first scale resolution; extracting a second set of features at a second scale resolution (lower than the first scale resolution); performing a self-attention transform to generate similarity scores for the second set of features; adding the similarity scores to the second set of features to generate a first feature extractor output; up-sampling the first feature extractor output to generate a second feature extractor output; adding the second feature extractor output to the first set of features to generate a third feature extractor output; receiving an instance query; performing a cross-attention transform on the instance query and the first feature extractor output to generate a set of weights; and matrix multiplying the set of weights and the third feature extractor output to generate instance masks.

Claims (83)

1. A method for image processing, comprising:

obtaining an image of an environment;

extracting a first set of features at a first scale resolution of the image;

extracting a second set of features at a second scale resolution of the image, wherein the second scale resolution is lower than the first scale resolution;

performing a self-attention transform to generate similarity scores for the second set of features;

adding the similarity scores to the second set of features to generate a first feature extractor output;

up-sampling the first feature extractor output to generate a second feature extractor output;

adding the second feature extractor output to the first set of features to generate a third feature extractor output;

receiving an instance query for instances of a feature;

performing a cross-attention transform on the instance query and the first feature extractor output to generate a first set of weights;

performing a matrix multiplication on the first set of weights and the third feature extractor output to generate instance masks for the image; and

outputting the instance masks.

2. The method of claim 1 , wherein the feature received with the instance query is associated with a feature of the second set of features.

3. The method of claim 1 , wherein generating the first feature extractor output comprises performing a convolution operation on the added similarity scores and the second set of features.

4. The method of claim 1 , wherein the first set of features and the second set of features are extracted using a convolutional neural network based backbone.

5. The method of claim 1 , further comprising:

obtaining a set of probability scores based on the cross-attention transform;

associating probability scores of the set of probability scores with feature instances; and

outputting the probability scores.

6. The method of claim 1 , wherein the second set of features are extracted at a lowest scale resolution of all of the extracted sets of features for the image.

7. The method of claim 1 , further comprising performing a convolution operation and batch normalization operation on the second set of features.

8. The method of claim 1 , further comprising:

receiving a semantic query for a set of feature categories;

performing a cross-attention transform on the semantic query and the second feature extractor output to generate a second set of weights; and

performing a matrix multiplication on the second set of weights and the third feature extractor output to generate semantic masks for the image.

9. The method of claim 8 , further comprising outputting the semantic masks for the image.

10. The method of claim 1 , wherein the instance query indicates a number of instances of a feature to output.

11. An apparatus for image processing, comprising:

a memory; and

a processor coupled to the memory and configured to:

obtain an image of an environment;

extract a first set of features at a first scale resolution of the image;

extract a second set of features at a second scale resolution of the image, wherein the second scale resolution is lower than the first scale resolution;

perform a self-attention transform to generate similarity scores for the second set of features;

add the similarity scores to the second set of features to generate a first feature extractor output;

up-sample the first feature extractor output to generate a second feature extractor output;

add the second feature extractor output to the first set of features to generate a third feature extractor output;

receive an instance query for instances of a feature;

perform a cross-attention transform on the instance query and the first feature extractor output to generate a first set of weights;

perform a matrix multiplication on the first set of weights and the third feature extractor output to generate instance masks for the image; and

output the instance masks.

12. The apparatus of claim 11 , wherein the feature received with the instance query is associated with a feature of the second set of features.

13. The apparatus of claim 11 , wherein, to generate the first feature extractor output, the processor is configured to perform a convolution operation on the added similarity scores and the second set of features.

14. The apparatus of claim 11 , wherein the first set of features and the second set of features are extracted using a convolutional neural network based backbone.

15. The apparatus of claim 11 , wherein the processor is further configured to:

obtain a set of probability scores based on the cross-attention transform;

associate probability scores of the set of probability scores with feature instances; and

output the probability scores.

16. The apparatus of claim 11 , wherein the processor is further configured to extract the second set of features at a lowest scale resolution of all of the extracted sets of features for the image.

17. The apparatus of claim 11 , wherein the processor is further configured to perform a convolution operation and batch normalization operation on the second set of features.

18. The apparatus of claim 11 , wherein the processor is further configured to:

receive a semantic query for a set of feature categories;

perform a cross-attention transform on the semantic query and the second feature extractor output to generate a second set of weights; and

perform a matrix multiplication on the second set of weights and the third feature extractor output to generate semantic masks for the image.

19. The apparatus of claim 18 , wherein the processor is further configured to output the semantic masks for the image.

20. The apparatus of claim 11 , wherein the instance query indicates a number of instances of a feature to output.

21. A non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor, cause the processor to:

obtain an image of an environment;

extract a first set of features at a first scale resolution of the image;

extract a second set of features at a second scale resolution of the image, wherein the second scale resolution is lower than the first scale resolution;

perform a self-attention transform to generate similarity scores for the second set of features;

add the similarity scores to the second set of features to generate a first feature extractor output;

up-sample the first feature extractor output to generate a second feature extractor output;

add the second feature extractor output to the first set of features to generate a third feature extractor output;

receive an instance query for instances of a feature;

perform a cross-attention transform on the instance query and the first feature extractor output to generate a first set of weights;

perform a matrix multiplication on the first set of weights and the third feature extractor output to generate instance masks for the image; and

output the instance masks.

22. The non-transitory computer-readable medium of claim 21 , wherein the feature received with the instance query is associated with a feature of the second set of features.

23. The non-transitory computer-readable medium of claim 21 , wherein, to generate the first feature extractor output, the instructions cause the processor to perform a convolution operation on the added similarity scores and the second set of features.

24. The non-transitory computer-readable medium of claim 21 , wherein the first set of features and the second set of features are extracted using a convolutional neural network based backbone.

25. The non-transitory computer-readable medium of claim 21 , wherein the instructions cause the processor to:

obtain a set of probability scores based on the cross-attention transform;

associate probability scores of the set of probability scores with feature instances; and

output the probability scores.

26. The non-transitory computer-readable medium of claim 21 , wherein the instructions cause the processor to extract the second set of features at a lowest scale resolution of all of the extracted sets of features for the image.

27. The non-transitory computer-readable medium of claim 21 , wherein the instructions cause the processor to perform a convolution operation and batch normalization operation on the second set of features.

28. The non-transitory computer-readable medium of claim 21 , wherein the instructions cause the processor to:

receive a semantic query for a set of feature categories;

perform a cross-attention transform on the semantic query and the second feature extractor output to generate a second set of weights; and

perform a matrix multiplication on the second set of weights and the third feature extractor output to generate semantic masks for the image.

29. The non-transitory computer-readable medium of claim 28 , wherein the instructions cause the processor to output the semantic masks for the image.

30. The non-transitory computer-readable medium of claim 21 , wherein the instance query indicates a number of instances of a feature to output.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2023
From: LI, XIN; LYU, JIANCHENG; QI, YINGYONG
To: QUALCOMM INCORPORATED
Reel/Frame 063872/0158 →
Continuity (1)
Related Publication 20240378727A1 · Nov 14, 2024
References Cited (27)
US 11971955B1 · Chakraborty · 2024 [cited by examiner]
US 12061094B2 · Iqbal · 2024 [cited by examiner]
US 12094159B1 · Akbas · 2024 [cited by examiner]
US 12347068B2 · Liu · 2025 [cited by examiner]
US 20150279113A1 · Knorr · 2015 [cited by examiner]
US 20200302225A1 · Dutta · 2020 [cited by examiner]
US 20210012576A1 · Riegler · 2021 [cited by examiner]
US 20210209837A1 · Chen · 2021 [cited by examiner]
US 20230101653A1 · Matsumura · 2023 [cited by examiner]
US 20230245495A1 · Ninh · 2023 [cited by examiner]
US 20230298272A1 · Ezhov · 2023 [cited by examiner]
US 20230306600A1 · Zhang · 2023 [cited by examiner]
US 20230377093A1 · Djelouah · 2023 [cited by examiner]
US 20230386052A1 · Lyu · 2023 [cited by examiner]
US 20230410339A1 · Sawarkar · 2023 [cited by examiner]
US 20240029203A1 · Xu · 2024 [cited by examiner]
US 20240062365A1 · Shen · 2024 [cited by examiner]
US 20240070809A1 · Xu · 2024 [cited by examiner]
US 20240089580A1 · Nomura · 2024 [cited by examiner]
US 20240249434A1 · Hong · 2024 [cited by examiner]
US 20240265676A1 · Janousková · 2024 [cited by examiner]
US 20240412319A1 · Singh · 2024 [cited by examiner]
Li, Lingling et al. “Scale-Insensitive Object Detection via Attention Feature Pyramid Transformer Network” Neural Processing Letters (2022) 54:581-595 (Year: 2021). [cited by examiner]
International Search Report and Written Opinion—PCT/US2024/018663—ISA/EPO—Jun. 27, 2024. [cited by applicant]
Karim R., et al., “MED-VT: Multiscale Encoder-Decoder Video Transformer with Application to Object Segmentation”, arXiv:2304.05930v1 [cs.CV], arxiv.org, Cornell University Library, 201 Olin Library Cornell University It… [cited by applicant]
Li L., et al., “Scale-Insensitive Object Detection via Attention Feature Pyramid Transformer Network”, Neural Processing Letters, Kluwer Academic Publishers, Norwell, MA, US, vol. 54, No. 1, Oct. 19, 2021, pp. 581-595, … [cited by applicant]
Zhu X., et al., “Deformable Detr: Deformable Transformers for End-to-End Object Detection”, arXiv:2010.04159v4 [cs.CV], Mar. 18, 2021, pp. 1-16, XP093174967, Section 4, Figures 1, 2. [cited by applicant]