IP Library › Granted Patent US 12,189,714
Granted Patent B2
US 12,189,714 · App. 18/266,744 · Granted Jan 7, 2025

System and method for improved few-shot object detection using a dynamic semantic network

Inventors: Marios Savvides (Pittsburgh, PA); Chenchen Zhu (Pittsburgh, PA); Fangyi Chen (Pittsburgh, PA); Uzair Ahmed (Pittsburgh, PA); Ran Tao (Pittsburgh, PA)
Assignee: Carnegie Mellon University
G06F18/2136G06N3/04G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,189,714
App. No.
18/266,744
Granted
Jan 7, 2025
Kind
B2
Abstract

Disclosed herein is an improved few-shot detector which utilizes a dynamic semantic network which takes as input a language feature and generates trainable parameters for a visual network. The visual network takes a visual feature as input and generates a classification and localization of an object.

Claims (37)

1. A few-shot object detector comprising:

a visual network with trainable parameters producing an output based on an input of a visual feature; and

a dynamic semantic network which accepts as input a language feature representing a semantic representation of a class and outputs a class-specific parameter for the visual network;

wherein the dynamic semantic network includes a dynamic relation graph for building direct connections between base classes and novel classes using non-visual knowledge of the base classes.

2. The few-shot detector of claim 1 wherein the language feature input to the dynamic semantic network represents a language representation of a class for which the visual network is trained to detect.

3. The few-shot detector of claim 2 further comprising a loss function generating a gradient for backpropagating to the visual network.

4. The few-shot detector of claim 3 wherein the visual network backpropagates the gradients to the dynamic semantic network.

5. The few-shot detector of claim 4 wherein the visual network computes partial derivatives of the gradients using chain rules before backpropagating the gradients to the dynamic semantic network.

6. The few-shot detector of claim 5 wherein the dynamic semantic network uses the gradient received from the visual network to update the trainable parameters of the dynamic semantic network.

7. The few-shot detector of claim 6 wherein the visual network comprises a classification sub-network and a localization sub-network.

8. The few-shot detector of claim 7 wherein the dynamic semantic network updates trainable parameters of both the classification sub-network and the localization sub-network.

9. The few-shot detector of claim 1 wherein the visual network is trained on a dataset comprising many instances of base class objects and few instances of novel class objects.

10. A system comprising:

a processor; and

memory, storing software that, when executed by the processor, implements the few-shot detector of claim 1 .

11. A method comprising:

training a visual network with trainable parameters to produce an output based on an input of a visual feature; and

training a dynamic semantic network to accept as input a language feature representing a semantic representation of a class and to output a class-specific parameter for the visual network;

wherein the dynamic semantic network includes a dynamic relation graph for building direct connections between base classes and novel classes using non-visual knowledge of the base classes.

12. The method of claim 11 wherein the language feature input to the dynamic semantic network represents a language representation of a class for which the visual network is trained to detect.

13. The method of claim 12 wherein the visual network:

receives backpropagated gradients from a loss function.

14. The method of claim 13 wherein the visual network:

backpropagates the gradients to the dynamic semantic network.

15. The method of claim 14 wherein the visual network:

computes partial derivatives of the gradients using chain rules before backpropagating the gradients to the dynamic semantic network.

16. The method of claim 15 wherein the dynamic semantic network:

uses the gradients received from the visual network to update the trainable parameters of the dynamic semantic network.

17. The method of claim 16 wherein the visual network comprises a classification sub-network and a localization sub-network.

18. The method of claim 7 wherein the dynamic semantic network:

generates the trainable parameters of both the classification sub-network and the localization sub-network.

19. The method of claim 11 wherein the visual network is trained on a dataset comprising many instances of base class objects and few instances of novel class objects.

20. A system comprising:

a processor; and

memory, storing software that, when executed by the processor, performs the method of claim 11 .

21. The method of claim 1 wherein the dynamic relation graph adapts word embeddings for a vision domain.

22. The method of claim 11 wherein the dynamic relation graph adapts word embeddings for a vision domain.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2024
From: SAVVIDES, MARIOS; ZHU, CHENCHEN; CHEN, FANGYI; AHMED, UZAIR; TAO, RAN
To: CARNEGIE MELLON UNIVERSITY
Reel/Frame 069049/0171 →
Continuity (4)
Continuation 17408674 · Aug 23, 2021
Provisional Application 63068871 · Aug 21, 2020
Provisional Application 63147782 · Feb 10, 2021
Related Publication 20240045925A1 · Feb 8, 2024
References Cited (11)
US 20140015855A1 · Denney · 2014 [cited by applicant]
US 20180089540A1 · Merler et al. · 2018 [cited by applicant]
US 20180212985A1 · Zadeh · 2018 [cited by applicant]
US 20190095716A1 · Shrestha et al. · 2019 [cited by applicant]
US 20190272451A1 · Lin · 2019 [cited by examiner]
US 20190325243A1 · Sikka · 2019 [cited by applicant]
US 20200364499A1 · Bagherinezhad · 2020 [cited by applicant]
US 20210019572A1 · Munoz Delgado · 2021 [cited by examiner]
US 20210117949A1 · Guo · 2021 [cited by applicant]
International Search Report and Written Opinion for the International Application No. PCT/US22/14833, mailed Apr. 26, 2022, 6 pages. [cited by applicant]
International Preliminary Report on Patentability for International Application No. PCT/US2022/014833, mailed May 21, 2024, 8 pages. [cited by applicant]