IP Library › Granted Patent US 12,411,887
Granted Patent B2
US 12,411,887 · App. 18/057,008 · Granted Sep 9, 2025

Transformer-based object detection

Inventors: Linjie Yang (Los Angeles, CA); Yiming Cui (Los Angeles, CA); Haichao Yu (Los Angeles, CA)
Assignee: Lemon Inc.
G06F16/538G06V10/42G06V10/7715G06V10/774G06V10/82G06V20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,411,887
App. No.
18/057,008
Granted
Sep 9, 2025
Kind
B2
Abstract

Object detection using a transformer-based object detection model includes randomly initializing basic queries for the model, modulating the basic queries based on semantics of input images, and training the model basic on features extracted from input images and the modulated queries.

Claims (36)

1. A method of detecting objects using a transformer-based object detection model, comprising:

generating modulated queries based on basic queries of a transformer-based object detection model and extracted image features of an input image;

training the transformer-based object detection model using the basic queries and the modulated queries;

replacing the basic queries with the modulated queries as input for a transformer decoder that is to execute the trained transformer-based object detection model; and

inputting the modulated queries and the extracted image features to the transformer decoder.

2. The method of claim 1 , wherein the generating of the modulated queries includes generating convex combinations of the basic queries to produce the modulated queries.

3. The method of claim 2 , wherein the generating of the convex combinations is based on combination coefficients generated by inputting global features of the extracted image features to a multi-layer perceptron (MLP).

4. The method of claim 1 , wherein the extracted image features include image features extracted from input images.

5. A non-volatile computer-readable medium that, when executed, causes at least one processor to perform operations related to object detection comprising:

receiving an input image;

generating dynamic detection queries based on basic queries and semantics of the input image;

training a transformer-based object detection model using the basic queries and the dynamic detection queries; and

performing object detection based on the dynamic detection queries using the trained transformer-based object detection model.

6. The non-volatile computer-readable medium of claim 5 , wherein the generating of the dynamic detection queries comprises:

calculating a global feature as a global average pooling of a feature map of the semantics of the input image;

generating combination coefficients by inputting the global feature to a multi-layer perceptron (MLP); and

performing a convex calculation of the combination coefficients and the basic queries to generate the dynamic detection queries.

7. The non-volatile computer-readable medium of claim 6 , wherein the calculating is based on features extracted from input images.

8. The non-volatile computer-readable medium of claim 5 , wherein the performing of object detection using the trained transformer-based object detection model includes replacing the basic queries with modulated queries.

9. The non-volatile computer-readable medium of claim 5 , wherein the performing of the object detection is by a DETR transformer decoder.

10. The non-volatile computer-readable medium of claim 5 , wherein the at least one processor executes on a server corresponding to a social media platform.

11. The non-volatile computer-readable medium of claim 5 , wherein the at least one processor executes upon a smart device.

12. An apparatus comprising:

at least one processor; and

at least one computer-readable medium that, when executed, causes the at least one processor to perform operations comprising:

extracting features from an input image;

calculating combination coefficients using output from the features;

generating dynamic detection queries as a function of the combination coefficients and basic queries for a transformer-based object detection model; and

performing object detection based on the features and the dynamic detection queries.

13. The apparatus of claim 12 , wherein the apparatus executes a DETR object detection model.

14. The apparatus of claim 12 , wherein calculating the combination coefficients comprise:

receiving the features extracted from the input image;

calculating a global feature of the extracted images; and

inputting the global feature to a multi-layer perceptron (MLP) to generate the combination coefficients.

15. The apparatus of claim 14 , wherein the dynamic detection queries are generated by performing a convex calculation of the combination coefficients and the basic queries.

16. The apparatus of claim 12 , wherein inputs for performing the object detection excludes the basic queries.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2023
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 064351/0164 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2023
From: YANG, LINJIE; CUI, YIMING; YU, HAICHAO
To: BYTEDANCE INC.
Reel/Frame 064351/0172 →
Continuity (1)
Related Publication 20240168991A1 · May 23, 2024
References Cited (28)
US 11886488B2 · Juan · 2024 [cited by examiner]
US 20220019889A1 · Aum · 2022 [cited by examiner]
US 20230206456A1 · Lee · 2023 [cited by examiner]
Wang, Y., et al. “Anchor detr: Query design for transformer-based object detection. arxiv 2021.” arXiv preprint arXiv:2109.07107 3. (Year: 2022). [cited by examiner]
Liu, Zhengyi, et al. “TriTransNet: RGB-D salient object detection with a triplet transformer embedding network.” Proceedings of the 29th ACM international conference on multimedia. 2021. (Year: 2021). [cited by examiner]
Carion, Nicolas, et al. “End-to-end object detection with transformers.” European conference on computer vision. Cham: Springer International Publishing, 2020. (Year: 2020). [cited by examiner]
Meinhardt, Tim, et al. “TrackFormer: Multi-Object Tracking with Transformers.” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2022. (Year: 2022). [cited by examiner]
Wang, Y., et al. “Anchor detr: Query design for transformer-based object detection. arxiv 2021.” arXiv preprint arXiv:2109.07107 3. (Year: 2022) (Year: 2022). [cited by examiner]
Liu, Zhengyi, et al. “TriTransNet: RGB-D salient object detection with a triplet transformer embedding network.” Proceedings of the 29th ACM international conference on multimedia. 2021. (Year: 2021) (Year: 2021). [cited by examiner]
Carion, Nicolas, et al. “End-to-end object detection with transformers.” European conference on computer vision. Cham: Springer International Publishing, 2020. (Year: 2020) (Year: 2020). [cited by examiner]
Meinhardt, Tim, et al. “TrackFormer: Multi-Object Tracking with Transformers.” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2022. (Year: 2022) (Year: 2022). [cited by examiner]
Zhi Tian et al. Conditional convolutions for instance segmentation. In Proc. Eur. Conf. Computer Vision (ECCV), Jul. 2020. (18 pages). [cited by applicant]
Yingming Wang et al. Anchor DETR: Query design for transformer-based object detection. arXiv preprint arXiv:2109.07107, Jan. 2022. (8 pages). [cited by applicant]
Tsung-Yi Lin et al. Microsoft COCO: Common objects in context. In European conference on computer vision, pp. 740-755. Springer, Feb. 2015. [cited by applicant]
Xizhou Zhu et al. Deformable DETR: Deformable transformers for end-to-end object detection. In International Conference on Learning Representations, Mar. 2021. (16 pages). [cited by applicant]
Bowen Cheng et al. Mask2former for video instance segmentation. Dec. 2021. (3 pages). [cited by applicant]
Shilong Liu et al.. DAB-DETR: Dynamic anchor boxes are better queries for DETR. In International Conference on Learning Representations, Mar. 2022. (19 pages). [cited by applicant]
Yinpeng Chen et al. Dynamic convolution: Attention over convolution kernels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11030-11039, Mar. 2020. [cited by applicant]
Linjie Yang et al. Video instance segmentation. In Proceedings of the IEEE International Conference on Computer Vision, pp. 5188-5197, Aug. 2019. [cited by applicant]
Marius Cordts et al. The cityscapes dataset for semantic urban scene understanding. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Apr. 2016. (29 pages). [cited by applicant]
Alexander Kirillov et al. Panoptic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9404-9413, Apr. 2019. [cited by applicant]
Peng Gao et al. Fast convergence of DETR with spatially modulated co-attention, 2021. (10 pages). [cited by applicant]
Kaiming He et al. Mask R-CNN. In Proceedings of the IEEE international conference on computer vision, pp. 2961-2969, Jan. 2018. [cited by applicant]
Yuxin Fang et al. Tracking Instances as Queries. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 6910-6919, Jun. 2021. [cited by applicant]
Yuwen Xiong et al. UPSNet: A unified panoptic segmentation network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8818-8826, 2019. [cited by applicant]
Nicolas Carion et al. End-to-end object detection with transformers. In ECCV, 2020. (26 pages). [cited by applicant]
Sukjun Hwang et al. Video instance segmentation using inter-frame communication transformers. Advances in Neural Information Processing Systems, 34:13352-13363, 2021. [cited by applicant]
Junfeng Wu et al. Seqformer: a frustratingly simple model for video instance segmentation. arXiv preprint arXiv:2112.08275, 2021. (11 pages). [cited by applicant]