IP Library › Granted Patent US 12,118,062
Granted Patent B2
US 12,118,062 · App. 17/246,803 · Granted Oct 15, 2024

Method and apparatus with adaptive object tracking

Inventors: Seohyung Lee (Seoul, KR); Dongwook Lee (Suwon-si, KR); Changyong Son (Anyang-si, KR); SeungWook Kim (Seoul, KR); Byung In Yoo (Seoul, KR)
Assignee: Samsung Electronics Co., Ltd.
G06F18/214G06T7/75G06V10/25G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,118,062
App. No.
17/246,803
Granted
Oct 15, 2024
Kind
B2
Abstract

Disclosed is a method and apparatus for adaptive tracking of a target object. The method includes method of tracking an object, the method including estimating a dynamic characteristic of an object in an input image based on frames of the input image, determining a size of a crop region for a current frame of the input image based on the dynamic characteristic of the object, generating a cropped image by cropping the current frame based on the size of the crop region, and generating a result of tracking the object for the current frame using the cropped image.

Claims (64)

1. A processor-implemented method, the method comprising:

estimating a dynamic characteristic of an object in an input image based on frames of the input image, wherein the dynamic characteristic of the object includes at least one of a change of size, a change of shape, or a change of location of the object;

determining a size of a crop region for a current frame of the input image based on the dynamic characteristic of the object;

generating a cropped image by cropping the current frame based on the size of the crop region; and

generating a result of tracking the object for the current frame using the cropped image.

2. The method of claim 1 , wherein the dynamic characteristic comprises a movement of the object, and

the determining of the size of the crop region for the current frame comprises increasing the size of the crop region, in response to the movement meeting a threshold, and decreasing the size of the crop region, in response to the movement failing to meet the threshold.

3. The method of claim 1 , wherein the generating of the result of tracking the object comprises:

selecting a neural network model corresponding to the size of the crop region from among neural network models configured to perform object tracking; and

generating the result of tracking the object using the cropped image and the selected neural network model.

4. The method of claim 3 , wherein the neural network models comprise a first neural network model for a first size of the crop region and a second neural network model for a second size of the crop region.

5. The method of claim 4 , wherein the selecting of the neural network model comprises:

selecting the first neural network model, in response to the size of the crop region being the first size; and

selecting the second neural network model, in response to the size of the crop region being the second size.

6. The method of claim 4 , wherein the first size is smaller than the second size, and

the first neural network model is configured to amplify input feature information more than the second neural network.

7. The method of claim 4 , wherein the first size is smaller than the second size, and

the first neural network model is configured to amplify input feature information more than the second neural network model by increasing a channel size using more weight kernels than the second neural network.

8. The method of claim 4 , wherein the first size is smaller than the second size, and

the first neural network model is configured to amplify input feature information more than the second neural network model by using a smaller pooling window than a pooling window in the second neural network.

9. The method of claim 3 , wherein the neural network models share at least one weight with each other.

10. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1 .

11. An apparatus, the apparatus comprising:

a processor configured to:

estimate a dynamic characteristic of an object in an input image based on frames of the input image, wherein the dynamic characteristic of the object includes at least one of a change of size, a change of shape, or a change of location of the object;

determine a size of a crop region for a current frame of the input image based on the dynamic characteristic of the object;

generate a cropped image by cropping the current frame based on the size of the crop region; and

generate a result of tracking the object for the current frame using the cropped image.

12. The apparatus of claim 11 , wherein the dynamic characteristic comprises a movement of the object, and

the processor is further configured to increase the size of the crop region, in response to the movement meeting a threshold, and to decrease the size of the crop region, in response to the movement failing to meet the threshold.

13. The apparatus of claim 11 , wherein the processor is further configured to:

select a neural network model corresponding to the size of the crop region from among neural network models configured to track the object; and

generate the result of tracking the object using the cropped image and the selected neural network model.

14. An apparatus, the apparatus comprising:

a processor configured to:

estimate a dynamic characteristic of an object in an input image based on frames of the input image;

determine a size of a crop region for a current frame of the input image based on the dynamic characteristic of the object;

generate a cropped image by cropping the current frame based on the size of the crop region;

select a meural network model corresponding to the size of the crop region from among neural network models configured to track the object; and

generate a result of tracking the object for the current frame using the cropped image and the selected neural network model,

wherein the neural network models comprise a first neural network model for a first size of the crop region and a second neural network model for a second size of the crop region.

15. The apparatus of claim 14 , wherein the first size is smaller than the second size, and

the first neural network model is configured to amplify input feature information more than the second neural network.

16. The apparatus of claim 14 , wherein the first size is smaller than the second size, and

the first neural network model is configured to amplify input feature information more than the second neural network model by increasing a channel size using more weight kernels than the second neural network.

17. The apparatus of claim 14 , wherein the first size is smaller than the second size, and

the first neural network model is configured to amplify input feature information more than the second neural network model by using a smaller pooling window than the second neural network.

18. AR-A.device, comprising:

a camera configured to generate an input image based on sensed visual information; and

a processor configured to:

estimate a dynamic characteristic of an object in the input image based on frames of the input image;

determine a size of a crop region for a current frame of the input image based on the dynamic characteristic of the object;

generate a cropped image by cropping the current frame based on the size of the crop region;

select a neural network model corresponding to the size of the crop region from among neural network models configured to track the object, and

generate a result of tracking the object for the current frame using the cropped image and the selected neural network model,

wherein the neural network models share at least one weight with each other.

19. The device of claim 18 , wherein the dynamic characteristic comprises a movement of the object, and

the processor is further configured to increase the size of the crop region, in response to the movement meeting a threshold, and to decrease the size of the crop region, in response to the movement failing to meet the threshold.

20. The apparatus of claim 13 , wherein the neural network models comprise a first neural network model for a first size of the crop region and a second neural network model for a smaller second size of the crop region.

21. The apparatus of claim 20 , wherein the first neural network model is configured to amplify input feature information more than the second neural network.

22. The apparatus of claim 20 , wherein the first neural network model is configured to amplify input feature information more than the second neural network model by increasing a channel size using more weight kernels than the second neural network.

23. The apparatus of claim 20 , wherein the first neural network model is configured to amplify input feature information more than the second neural network model by using a smaller pooling window than the second neural network.

24. The device of claim 18 , wherein the dynamic characteristic of the object includes at least one of a change of size, a change of shape, or a change of location of the object.

25. The device of claim 18 , wherein the neural network models comprise a first neural network model for a first size of the crop region and a second neural network model for a smaller second size of the crop region.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2021
From: LEE, SEOHYUNG; LEE, DONGWOOK; SON, CHANGYONG; KIM, SEUNGWOOK; YOO, BYUNG IN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 056111/0104 →
Priority Claims (1)
KR 10-2020-0144491 · Nov 2, 2020 · national
Continuity (1)
Related Publication 20220138493A1 · May 5, 2022
Cited By (1)
US 12,452,524