IP Library › Granted Patent US 12,067,733
Granted Patent B2
US 12,067,733 · App. 17/461,978 · Granted Aug 20, 2024

Video target tracking method and apparatus, computer device, and storage medium

Inventors: Zhen Cui (Shenzhen, CN); Zequn Jie (Shenzhen, CN); Li Wei (Shenzhen, CN); Chunyan Xu (Shenzhen, CN); Tong Zhang (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G06T7/215G06F18/214G06N20/00G06T7/11G06V10/462G06V20/46G06T2207/10016G06T2207/20081G06T2207/20084G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,067,733
App. No.
17/461,978
Granted
Aug 20, 2024
Kind
B2
Abstract

A video target tracking method is provided to a computing device, the method includes: obtaining a partial detection map corresponding to a target image frame in a to-be-detected video; obtaining a relative motion saliency map corresponding to the target image frame; determining constraint information corresponding to the target image frame according to the partial detection map and the relative motion saliency map; adjusting a parameter of an image segmentation model by using the constraint information, to obtain an adjusted image segmentation model; and extracting a target object from the target image frame by using the adjusted image segmentation model.

Claims (74)

1. A video target tracking method, performed by a computer device, the method comprising:

obtaining a partial detection map corresponding to a target image frame in a to-be-detected video, the partial detection map being generated based on appearance information of a target object that is in the to-be-detected video and is to be tracked by an image segmentation model;

obtaining a relative motion saliency map corresponding to the target image frame, the relative motion saliency map being generated based on motion information of the target object, comprising:

calculating an optical flow between the target image frame and a neighboring image frame;

determining a background optical flow according to an optical flow of a background region in the partial detection map; and

generating the relative motion saliency map according to the background optical flow and the optical flow corresponding to the target image frame;

determining constraint information corresponding to the target image frame according to the partial detection map and the relative motion saliency map, the constraint information including an absolute positive sample pixel, an absolute negative sample pixel, and an undetermined sample pixel in the target image frame;

adjusting a parameter of the image segmentation model by using the constraint information, to obtain an adjusted image segmentation model; and

extracting the target object from the target image frame by using the adjusted image segmentation model.

2. The method according to claim 1 , wherein obtaining the partial detection map comprises:

selecting at least one training sample from a marked image frame of the to-be-detected video, the training sample including the marked image frame and a detection target box corresponding to the marked image frame, and the detection target box referring to an image region in which a proportion of the target object in the detection target box is greater than a preset threshold;

adjusting a parameter of a target detection model by using the training sample, to obtain an adjusted target detection model; and

processing the target image frame by using the adjusted target detection model, to obtain the partial detection map.

3. The method according to claim 2 , wherein selecting the at least one training sample comprises:

scattering a box in the marked image frame to generate a scattered box;

calculating a proportion of the target object in the randomly scattered box; and

determining the box as the detection target box corresponding to the marked image frame in response to determining that the proportion of the target object in the scattered box is greater than the preset threshold, and selecting the marked image frame and the detection target box as the training sample.

4. The method according to claim 1 , wherein the background region in the partial detection map refers to a remaining region of the partial detection map except a region in which the target object is detected.

5. The method according to claim 1 , wherein determining the constraint information comprises:

determining, for a target pixel in the target image frame, the target pixel as the absolute positive sample pixel in response to determining that a value of the target pixel in the partial detection map satisfies a first preset condition, and a value of the target pixel in the relative motion saliency map satisfies a second preset condition;

determining the target pixel as the absolute negative sample pixel in response to determining that the value of the target pixel in the partial detection map does not satisfy the first preset condition, and the value of the target pixel in the relative motion saliency map does not satisfy the second preset condition;

determining the target pixel as the undetermined sample pixel in response to determining that the value of the target pixel in the partial detection map satisfies the first preset condition, and the value of the target pixel in the relative motion saliency map does not satisfy the second preset condition; or

determining the target pixel as the undetermined sample pixel in response to determining that the value of the target pixel in the partial detection map does not satisfy the first preset condition, and the value of the target pixel in the relative motion saliency map satisfies the second preset condition.

6. The method according to claim 1 , wherein adjusting the parameter of the image segmentation model comprises:

adjusting the parameter of the image segmentation model by using the absolute positive sample pixel and the absolute negative sample pixel, to obtain the adjusted image segmentation model.

7. The method according to claim 1 , wherein the image segmentation model is trained by:

constructing an initial image segmentation model;

performing preliminary training on the initial image segmentation model by using a first sample set, to obtain a preliminarily trained image segmentation model, the first sample set including at least one marked picture; and

retraining, by using a second sample set, the preliminarily trained image segmentation model, to obtain a pre-trained image segmentation model, the second sample set including at least one marked video.

8. A video target tracking apparatus, comprising: a memory storing computer program instructions; and a processor coupled to the memory and configured to execute the computer program instructions and perform:

obtaining a partial detection map corresponding to a target image frame in a to-be-detected video, the partial detection map being generated based on appearance information of a target object that is in the to-be-detected video and is to be tracked by an image segmentation model, and the image segmentation model being a neural network model configured to segment and extract the target object from an image frame of the to-be-detected video;

obtaining a relative motion saliency map corresponding to the target image frame, the relative motion saliency map being generated based on motion information of the target object, comprising:

calculating an optical flow between the target image frame and a neighboring image frame;

determining a background optical flow according to an optical flow of a background region in the partial detection map; and

generating the relative motion saliency map according to the background optical flow and the optical flow corresponding to the target image frame;

determining constraint information corresponding to the target image frame according to the partial detection map and the relative motion saliency map, the constraint information including an absolute positive sample pixel, an absolute negative sample pixel, and an undetermined sample pixel in the target image frame;

adjusting a parameter of the image segmentation model by using the constraint information, to obtain an adjusted image segmentation model; and

extracting the target object from the target image frame by using the adjusted image segmentation model.

9. The apparatus according to claim 8 , wherein the processor is configured to execute the computer program instructions and further perform:

selecting at least one training sample from a marked image frame of the to-be-detected video, the training sample including the marked image frame and a detection target box corresponding to the marked image frame, and the detection target box referring to an image region in which a proportion of the target object is greater than a preset threshold;

adjusting a parameter of a target detection model by using the training sample, to obtain an adjusted target detection model; and

processing the target image frame by using the adjusted target detection model, to obtain the partial detection map.

10. The apparatus according to claim 8 , wherein the background region in the partial detection map refers to a remaining region of the partial detection map except a region in which the target object is detected.

11. The apparatus according to claim 8 , wherein the processor is configured to execute the computer program instructions and further perform:

determining, for a target pixel in the target image frame, the target pixel as the absolute positive sample pixel in response to determining that a value of the target pixel in the partial detection map satisfies a first preset condition, and a value of the target pixel in the relative motion saliency map satisfies a second preset condition;

determining the target pixel as the absolute negative sample pixel in response to determining that the value of the target pixel in the partial detection map does not satisfy the first preset condition, and the value of the target pixel in the relative motion saliency map does not satisfy the second preset condition; and

determining the target pixel as the undetermined sample pixel in response to determining that the value of the target pixel in the partial detection map satisfies the first preset condition, and the value of the target pixel in the relative motion saliency map does not satisfy the second preset condition, or in response to determining that the value of the target pixel in the partial detection map does not satisfy the first preset condition, and the value of the target pixel in the relative motion saliency map satisfies the second preset condition.

12. The apparatus according to claim 8 , wherein the processor is configured to execute the computer program instructions and further perform:

adjusting the image segmentation model by using the absolute positive sample pixel and the absolute negative sample pixel, to obtain the adjusted image segmentation model.

13. The apparatus according to claim 8 , wherein the processor is configured to execute the computer program instructions and further perform:

scattering a box in the marked image frame to generate a scattered box;

calculating a proportion of the target object in the randomly scattered box; and

determining the box as the detection target box corresponding to the marked image frame in response to determining that the proportion of the target object in the scattered box is greater than the preset threshold, and selecting the marked image frame and the detection target box as the training sample.

14. The apparatus according to claim 8 , wherein the image segmentation model is trained by:

constructing an initial image segmentation model;

performing preliminary training on the initial image segmentation model by using a first sample set, to obtain a preliminarily trained image segmentation model, the first sample set including at least one marked picture; and

retraining, by using a second sample set, the preliminarily trained image segmentation model, to obtain a pre-trained image segmentation model, the second sample set including at least one marked video.

15. A non-transitory computer-readable storage medium, storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set, when being loaded and executed by a processor, causes the processor to perform the following operations:

obtaining a partial detection map corresponding to a target image frame in a to-be-detected video, the partial detection map being generated based on appearance information of a target object that is in the to-be-detected video and is to be tracked by an image segmentation model;

obtaining a relative motion saliency map corresponding to the target image frame, the relative motion saliency map being generated based on motion information of the target object, comprising:

calculating an optical flow between the target image frame and a neighboring image frame;

determining a background optical flow according to an optical flow of a background region in the partial detection map; and

generating the relative motion saliency map according to the background optical flow and the optical flow corresponding to the target image frame;

determining constraint information corresponding to the target image frame according to the partial detection map and the relative motion saliency map, the constraint information including an absolute positive sample pixel, an absolute negative sample pixel, and an undetermined sample pixel in the target image frame;

adjusting a parameter of the image segmentation model by using the constraint information, to obtain an adjusted image segmentation model; and

extracting the target object from the target image frame by using the adjusted image segmentation model.

16. The non-transitory storage medium according to claim 15 , wherein the at least one instruction, the at least one program, the code set or the instruction set, causes the processor to perform:

selecting at least one training sample from a marked image frame of the to-be-detected video, the training sample including the marked image frame and a detection target box corresponding to the marked image frame, and the detection target box referring to an image region in which a proportion of the target object in the detection target box is greater than a preset threshold;

adjusting a parameter of a target detection model by using the training sample, to obtain an adjusted target detection model; and

processing the target image frame by using the adjusted target detection model, to obtain the partial detection map.

17. The non-transitory storage medium according to claim 16 , wherein the at least one instruction, the at least one program, the code set or the instruction set, causes the processor to perform:

scattering a box in the marked image frame to generate a scattered box;

calculating a proportion of the target object in the scattered box; and

determining the box as the detection target box corresponding to the marked image frame in response to determining that the proportion of the target object in the scattered box is greater than the preset threshold, and selecting the marked image frame and the detection target box as the training sample.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2021
From: CUI, ZHEN; JIE, ZEQUN; WEI, LI; XU, CHUNYAN; ZHANG, TONG
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 057333/0552 →
Priority Claims (1)
CN 201910447379.3 · May 27, 2019 · national
Continuity (2)
Continuation PCTCN2020088286 · Apr 30, 2020
Related Publication 20210398294A1 · Dec 23, 2021
Cited By (2)
US 12,511,906 US 12,737,896