IP Library Granted Patent US 12,205,357
Granted Patent B2
US 12,205,357 · App. 17/715,901 · Granted Jan 21, 2025

Learning ordinal representations for deep reinforcement learning based object localization

Inventors: Shaobo Han (Princeton, NJ); Renqiang Min (Princeton, NJ); Tingfeng Li (Plainsboro, NJ)
Assignee: NEC Corporation
G06V10/7788G06V10/82G06V30/19167
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,357
App. No.
17/715,901
Granted
Jan 21, 2025
Kind
B2
Abstract

A reinforcement learning based approach to the problem of query object localization, where an agent is trained to localize objects of interest specified by a small exemplary set. We learn a transferable reward signal formulated using the exemplary set by ordinal metric learning. It enables test-time policy adaptation to new environments where the reward signals are not readily available, and thus outperforms fine-tuning approaches that are limited to annotated images. In addition, the transferable reward allows repurposing of the trained agent for new tasks, such as annotation refinement, or selective localization from multiple common objects across a set of images. Experiments on corrupted MNIST dataset and CU-Birds dataset demonstrate the effectiveness of our approach.

Claims (6)

1. A deep reinforcement learning (RL) method for object localization comprising:

acquiring a seed dataset including a set of seed images each with ground truth bounding box annotation;

pretrain ordinal embedding by randomly perturbing the ground truth bounding box at different levels denoted by parameter p, said ordinal embedding satisfying an ordinal constraint locally for each pair of perturbed data augmented from the same image, wherein the pretraining is performed through the effect of a backbone network, a region of interest (RoI) head, and a triplet loss; and

using an embedding function, configuring RL agents to start from a whole image and recursively sample actions from a discrete action space such that rewards are produced, the rewards of a sample action determined from embedding distances and updating a policy network based on the rewards so determined; and

outputting an annotation policy and embedding function.

2. The method of claim 1 wherein the seed image bounding box annotation is initially provided by a human action.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2024
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 069540/0269 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2022
From: HAN, SHAOBO; MIN, RENQIANG; LI, TINGFENG
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 059537/0430 →
Continuity (4)
Provisional Application 63193916 · Jul 28, 2021
Provisional Application 63191403 · May 21, 2021
Provisional Application 63172171 · Apr 8, 2021
Related Publication 20220327814A1 · Oct 13, 2022
References Cited (5)
US 11995522B2 · Lourentzou · 2024 [cited by examiner]
US 20110134128A1 · Hu · 2011 [cited by examiner]
US 20190147621A1 · Alesiani · 2019 [cited by examiner]
US 20200410228A1 · Wang · 2020 [cited by examiner]
US 20220292107A1 · Novotny · 2022 [cited by examiner]