IP Library Granted Patent US 11,836,931
Granted Patent B2
US 11,836,931 · App. 17/273,320 · Granted Dec 5, 2023

Target detection method, apparatus and device for continuous images, and storage medium

Inventors: Xuchen Liu (Henan, CN); Xing Fang (Henan, CN); Hongbin Yang (Henan, CN); Yun Cheng (Henan, CN); Gang Dong (Henan, CN)
Assignee: ZHENGZHOU YUNHAI INFORMATION TECHNOLOGY CO., LTD.
G06T7/215G06N3/04G06T7/223G06T2207/10016G06T2207/20021G06T2207/20081G06T2207/20084G06T2207/20172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,836,931
App. No.
17/273,320
Granted
Dec 5, 2023
Kind
B2
Abstract

A method, an apparatus, and a device for target detection in consecutive images, and a computer-readable storage medium. A second frame is divided into multiple sub-images, before a target in the second frame in a video sequence is detected through a target-detecting network model. A first frame is searched, according to a preset rule for motion estimation, for a corresponding image block matched with each sub-image. Pixels of a sub-image, of which the matched image block is found in the first frame, are replaced with preset background pixels. Hence, a target repeating in both frames is replaced. Finally, the second frame subject to the replacement is inputted in to the target-detecting network model, to obtain a bounding box of a target object of the second frame and a category of such target object. An algorithm for target detection in consecutive images is optimized.

Claims (39)

1. A method for target detection in consecutive images, comprising:

inputting a first frame in a video sequence into a target-detecting network model, to obtain a bounding box of a target object of the first frame and a category of the target object of the first frame;

dividing a second frame in the video sequence into a plurality of first sub-images;

searching, according to a preset rule for motion estimation, the first frame for an image block matching with each of the plurality of first sub-images, to determine a position of the target object in the second frame;

replacing pixels of a first sub-image in the plurality of first sub-images with preset background pixels, wherein the image block matched with the first sub-image is found in the first frame; and

inputting the second frame, in which the pixels of the first sub-image are replaced, into the target-detecting network model, to obtain a bounding box of a target object of the second frame and a category of the target object of the second frame;

wherein the second frame is subsequent and adjacent to the first frame.

2. The method according to claim 1 , wherein the target-detecting network model is a YOLO-v3 network model or a SSD (Single Shot Multibox Detector) network model.

3. The method according to claim 2 , wherein after inputting the second frame in which the pixels of the first sub-image are replaced into the target-detecting network model, the method further comprises:

dividing a third frame in the video sequence into a plurality of second sub-images;

searching, according to the preset rule for motion estimation, the first frame and the second frame for image blocks matching with each of the plurality of second sub-images, to determine a position of the target object in the third frame;

replacing pixels of a second sub-image in the plurality of second sub-images with the preset background pixels, wherein the image blocks matched with the second sub-image are found in the first frame and the second frame; and

inputting the third frame, in which the pixels of the second sub-image are replaced, into the target-detecting network model, to obtain a bounding box of a target object of the third frame and a category of the target object of the third frame;

wherein the third frame is subsequent and adjacent to the second frame.

4. The method according to claim 1 , wherein before dividing the second frame in the video sequence into the plurality of first sub-images, the method further comprises:

denoising the second frame in the video sequence that is acquired, to remove noise interference in the second frame.

5. A device for target detection in consecutive images, comprising:

a memory, storing a computer program;

a processor, configured to implement the method according to claim 1 , when executing the computer program.

6. A computer-readable storage medium, storing a program for target detection in consecutive images, wherein:

the program when executed by a processor implements the method according to claim 1 .

7. An apparatus for target detection in consecutive images, comprising:

a first-frame inputting module, configured to input a first frame in a video sequence into a target-detecting network model, to obtain a bounding box of a target object of the first frame and a category of the target object of the first frame;

an image matching module, configured to:

divide a second frame in the video sequence into a plurality of first sub-images, and

search, according to a preset rule for motion estimation, the first frame for an image block matching with each of the plurality of first sub-images, to determine a position of the target object in the second frame,

wherein the second frame is subsequent and adjacent to the first frame;

a background replacing module, configured to replace pixels of a first sub-image in the plurality of first sub-images with preset background pixels, wherein the image block matched with the first sub-image is found in the first frame; and

a second-frame inputting module, configured to input the second frame, in which the pixels of the first sub-image are replaced, into the target-detecting network model, to obtain a bounding box of a target object of the second frame and a category of the target object of the second frame.

8. The apparatus according to claim 7 , wherein the target-detecting network model is a YOLO-v3 network model or a SSD (Single Shot Multibox Detector) network model.

9. The apparatus according to claim 7 , further comprising a third-frame processing module, wherein the third-frame processing module comprises:

a previous-frame matching sub-module, configured to

divide a third frame in the video sequence into a plurality of second sub-images, and

search, according to the preset rule for motion estimation, the first frame and the second frame for image blocks matching with each of the plurality of second sub-images, to determine a position of the target object in the third frame,

wherein the third frame is subsequent and adjacent to the second frame;

a repeated-target replacing sub-module, configured to replace pixels of a second sub-image in the plurality of second sub-images with the preset background pixels, wherein the image blocks matched with the second sub-image are found in the first frame and the second frame; and

a third-frame inputting sub-module, configured to input the third frame, in which the pixels of the second sub-image are replaced, into the target-detecting network model, to obtain a bounding box of a target object of the third frame and a category of the target object of the third frame.

10. The apparatus according to claim 7 , further comprising:

a denoising module, configured to denoise the second frame and the third frame in the video sequence, to remove noise interference in the second frame and the third frame.

Assignments (2)
LICENSE Recorded Jun 30, 2026
From: IEIT SYSTEMS CO., LTD
To: AIVRES SYSTEMS INC.
Reel/Frame 075857/0939 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2021
From: LIU, XUCHEN; FANG, XING; YANG, HONGBIN; CHENG, YUN; DONG, GANG
To: ZHENGZHOU YUNHAI INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 055486/0807 →
Priority Claims (1)
CN 201811038286.7 · Sep 6, 2018 · national
Continuity (1)
Related Publication 20210319565A1 · Oct 14, 2021
Cited By (1)
US 12,711,574