IP Library › Granted Patent US 12,536,689
Granted Patent B2
US 12,536,689 · App. 18/171,868 · Granted Jan 27, 2026

Mining unlabeled images with vision and language models for improving object detection

Inventors: Samuel Schulter (New York, NY); Vijay Kumar Baikampady Gopalkrishna (Santa Clara, CA)
Assignee: NEC Corporation
G06T7/70G06T3/40G06V10/25G06V20/70G08G1/16G06T2207/10024G06T2207/20132G06T2207/30252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,689
App. No.
18/171,868
Granted
Jan 27, 2026
Kind
B2
Abstract

A method for object detection obtains, from a set of RGB images lacking annotations, a set of regions that include potential objects, a bounding box, and an objectness score indicating a region prediction confidence. The method obtains, by a region scorer for each region in the set, a category from a fixed set of categories and a confidence for the category responsive to the objectness score. The method duplicates each region in the set to obtain a first and a second patch. The method encodes the patches to obtain an image vector. The method encodes a template sentence using the category to obtain a text vector for each category. The method compares the image vector to the text vector via a similarity function to obtain a similarity probability based on the confidence. The method defines a final set of pseudo labels based on the similarity probability being above a threshold.

Claims (103)

1 . A computer-implemented method for object detection, comprising:

obtaining, from a set of RGB images lacking annotations, a set of regions in an RGB image that include one or more potential objects, a bounding box for the one or more potential objects, and an objectness score indicating a region prediction confidence;

obtaining, by a region scorer for each region in the set, a category from a fixed set of categories and a confidence for the category name calculated as:

S

i

u

=

S

RPN

(

R

i

)

+

max

⁡

(

p

i

u

)

2

where S RPN (R i ) is the objectness score for a region R i and p i u is a probability distribution for the category name;

duplicating each region in the set to obtain a first and a second image patch;

encoding, by a visional and language (V & L) image encoder, the first and the second image patches to obtain an image vector;

encoding, by a V & L text encoder, a template sentence using the category to obtain a text vector for each category having a same dimensionality as the image vector;

comparing the image vector to the text vector via a similarity function to obtain a similarity probability based on the confidence; and

defining a final set of pseudo labels based on the similarity probability being above a user-defined similarity probability threshold.

2 . The computer-implemented method of claim 1 , further comprising discarding any regions equal to and below the user-defined threshold probability.

3 . The computer-implemented method of claim 1 , further comprising duplicating each region using into an original scale region and a larger scale region having same bounding box centers to obtain the first and the second image patch.

4 . The computer-implemented method of claim 3 , wherein the larger scale region is 1.5× larger than the original scale region.

5 . The computer-implemented method of claim 1 , further comprising rescaling the first and the second image patch to obtain a rescaled first and second image patches fitting an input of the vision and language encoder.

6 . The computer-implemented method of claim 1 , further comprising applying a softmax function over similarities between the image vector and the text vector to obtain a probability between 0 and 1 for a region to belong to a particular category c.

7 . The computer-implemented method of claim 1 , wherein a dimensionality of the image vector is defined by a V & L model comprising the V & L image and text encoders, the V & L model trained on pairs of images and corresponding captions.

8 . The computer-implemented method of claim 1 , wherein the image vector and the text vector have a same dimensionality.

9 . The computer-implemented method of claim 1 , wherein a maximum similarity probability is averaged with the confidence score, the confidence score having a value between 0 and 1.

10 . The computer-implemented method of claim 1 , further comprising training an object detection system with a combination of labeled data and the pseudo labels for open-vocabulary detection.

11 . The computer-implemented method of claim 1 , further comprising cropping the first and the second image patch from the RGB image prior to an encoding.

12 . The computer-implementing method of claim 1 , further comprising automatically controlling a vehicle system for collision avoidance responsive to a pseudo label predicting an impending collision.

13 . A computer program product for object detection, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

obtaining, by a hardware processor of the computer from a set of ROB images lacking annotations, a set of regions in an RGB image that include one or more potential objects, a bounding box for the one or more potential objects, and an objectness score indicating a region prediction confidence;

obtaining, by a region scorer implemented by the hardware processor for each region in the set, a category from a fixed set of categories and a confidence for the category name calculated as:

S

i

u

=

S

RPN

(

R

i

)

+

max

⁡

(

p

i

u

)

2

where S RPN (R i ) is the objectness score for a region R i and p i u is a probability distribution for the category name;

duplicating, by the hardware processor, each region in the set to obtain a first and a second image patch;

encoding, by a visional and language (V& L) image encoder implemented by the hardware processor, the first and the second image patches to obtain an image vector;

encoding, by a V&L text encoder implemented by the hardware processor, a template sentence using the category to obtain a text vector for each category having a same dimensionality as the image vector;

comparing, by the hardware processor, the image vector to the text vector via a similarity function to obtain a similarity probability based on the confidence; and

defining, by the hardware processor, a final set of pseudo labels based on the similarity probability being above a user-defined similarity probability threshold.

14 . The computer program product of claim 13 , wherein the method further comprises discarding any regions equal to and below the user-defined threshold probability.

15 . The computer program product of claim 13 , wherein the method further comprises duplicating each region using into an original scale region and a larger scale region having same bounding box centers to obtain the first and the second image patch.

16 . The computer program product of claim 15 , wherein the larger scale region is 1.5× larger than the original scale region.

17 . The computer program product of claim 13 , wherein the method further comprises rescaling the first and the second image patch to obtain a rescaled first and second image patches fitting an input of the vision and language encoder.

18 . The computer program product of claim 13 , wherein the method further comprises applying a softmax function over similarities between the image vector and the text vector to obtain a probability between 0 and 1 for a region to belong to a particular category c.

19 . The computer program product of claim 13 , wherein a dimensionality of the image vector is defined by a V & L model comprising the V & L image and text encoders, the V & L model trained on pairs of images and corresponding captions.

20 . A computer processing system fir object detection, comprising:

a memory device for storing program code; and

a hardware processor operatively coupled to the memory device for storing program code to:

obtain, from a set of RGB images lacking annotations, a set of regions in an RGB image that include one or more potential objects, a bounding box for the one or more potential objects, and an objectness score indicating a region prediction confidence

obtain, for each region in the set, a category from a fixed set of categories and a confidence for the category name calculated as:

S

i

u

=

S

RPN

(

R

i

)

+

max

⁡

(

p

i

u

)

2

where S RPN (R i ) is the objectness score for a region R i and p i u is a probability distribution for the category name;

duplicate each region in the set to obtain a first and a second image patch;

encode, by a visional and language (V& L) image encoder implemented by the hardware processor, the first and the second image patches to obtain an image vector;

encode, by a V&L text encoder implemented by the hardware processor, a template sentence using the category to obtain a text vector for each category having a same dimensionality as the image vector;

compare the image vector to the text vector via a similarity function to obtain a similarity probability based on the confidence; and

define a final set of pseudo labels based on the similarity probability being above a user-defined similarity probability threshold.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 073203/0407 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2023
From: SCHULTER, SAMUEL; GOPALKRISHNA, VIJAY KUMAR BAIKAMPADY
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 062753/0939 →
Continuity (3)
Provisional Application 63400766 · Aug 25, 2022
Provisional Application 63317496 · Mar 7, 2022
Related Publication 20230281858A1 · Sep 7, 2023
References Cited (25)
US 20210158096A1 · Sinha · 2021 [cited by examiner]
US 20220004771A1 · Grancharov · 2022 [cited by examiner]
US 20220391766A1 · Acuna Marrero · 2022 [cited by examiner]
US 20240119725A1 · Zeng · 2024 [cited by examiner]
US 20240338850A1 · Jing · 2024 [cited by examiner]
Zhong, Yiwu, et al. “Regionclip: Region-based language-image pretraining.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022. (Year: 2022). [cited by examiner]
Ren, Shaoqing, et al. “Faster R-CNN: Towards real-time object detection with region proposal networks.” IEEE transactions on pattern analysis and machine intelligence 39.6 (2016): 1137-1149. (Year: 2016). [cited by examiner]
Changpinyo, Soravit, et al. “Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2021. (Y… [cited by examiner]
Radford, Alec, et al. “Learning transferable visual models from natural language supervision.” International conference on machine learning. PmLR, 2021. (Year: 2021). [cited by examiner]
Lin, Chuang, et al. “Learning object-language alignments for open-vocabulary object detection.” arXiv preprint arXiv:2211.14843 (2022). (Year: 2022). [cited by examiner]
Jia, Chao, et al. “Scaling up visual and vision-language representation learning with noisy text supervision.” International conference on machine learning. PMLR, 2021. (Year: 2021). [cited by examiner]
Bansal, A., Sikka, K., Sharma, G., Chellappa, R., & Divakaran, A. (Sep. 8, 2018). Zero-shot object detection. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 384-400). [cited by applicant]
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. ;End-to-end object detection with transformers. In Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, Aug. 23-28, 2020, Proc… [cited by applicant]
Gao, M., Xing, C., Niebles, J. C., Li, J., Xu, R., Liu, W., & Xiong, C. (Nov. 18, 2021). Towards open vocabulary object detection without human-provided bounding boxes. arXiv preprint arXiv:2111.09452. [cited by applicant]
Gu, X., Lin, T. Y., Kuo, W., & Cui, Y. (Apr. 28, 2021). Open-vocabulary object detection via vision and language knowledge distillation. arXiv preprint arXiv:2104.13921. [cited by applicant]
Hu, R., & Singh, A. (2021, Oct. 11). Unit: Multimodal multitask learning with a unified transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 1439-1449). [cited by applicant]
Huynh, D., Kuen, J., Lin, Z., Gu, J., & Elhamifar, E. (Jun. 19, 2022). Open-vocabulary instance segmentation via robust cross-modal pseudo-labeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patte… [cited by applicant]
Kamath, A., Singh, M., LeCun, Y., Synnaeve, G., Misra, I., & Carion, N. (Oct. 11, 2021). Mdetr-modulated detection for end-to-end multi-modal understanding. In Proceedings of the IEEE/CVF International Conference on Com… [cited by applicant]
Li, B., Weinberger, K. Q., Belongie, S., Koltun, V., & Ranftl, R. (Apr. 25, 2022). Language-driven semantic segmentation. arXiv preprint arXiv:2201.03546. [cited by applicant]
Rahman, S., Khan, S., & Barnes, N. (Apr. 3, 2020). Improved visual-semantic alignment for zero-shot object detection. In Proceedings of the AAAI Conference on Artificial Intelligence (vol. 34, No. 07, pp. 11932-11939). [cited by applicant]
Rao, Y., Zhao, W., Chen, G., Tang, Y., Zhu, Z., Huang, G., . . . & Lu, J. (2022, Jun. 19). Denseclip: Language-guided dense prediction with context-aware prompting. In Proceedings of the IEEE/CVF Conference on Computer … [cited by applicant]
Shi, H., Hayat, M., Wu, Y., & Cai, J. (Jun. 19, 2022). ProposalCLIP: unsupervised open-category object proposal generation via exploiting clip cues. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patte… [cited by applicant]
Siméoni, O., Puy, G., Vo, H. V., Roburin, S., Gidaris, S., Bursuc, A., . . . & Ponce, J. (Nov. 25, 2021). Localizing objects with self-supervised transformers and No. labels. arXiv preprint arXiv:2109.14279. [cited by applicant]
Zhong, Y., Yang, J., Zhang, P., Li, C., Codella, N., Li, L. H., . . . & Gao, J. (Jun. 19, 2022). Regionclip: Region-based language-image pretraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patt… [cited by applicant]
Zhou, C., Loy, C. C., & Dai, B. (Dec. 2, 2021). Denseclip: Extract free dense labels from clip. arXiv preprint arXiv:2112.01071. [cited by applicant]
Cited By (1)
US 12,743,872