IP Library Granted Patent US 12,073,621
Granted Patent B2
US 12,073,621 · App. 17/370,764 · Granted Aug 27, 2024

Method and apparatus for detecting information insertion region, electronic device, and storage medium

Inventors: Hui Sheng (Shenzhen, CN); Dongbo Huang (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G06V20/41G06F18/23G06V10/255G06V10/462G06V20/49H04N21/44008H04N21/812H04N21/8456
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,073,621
App. No.
17/370,764
Granted
Aug 27, 2024
Kind
B2
Abstract

A method for detecting an information insertion region is provided. In the method, a video is obtained. The video is segmented to obtain video fragments, each of the video fragments including a subset of image frames in the video. A target frame is obtained in the video fragments. Objects in the target frame are identified and segmented, to obtain labeling information corresponding to the objects. A target object is determined according to the labeling information. Clustering is performed on the target object, to obtain a plurality of candidate to-be-inserted regions. A target candidate to-be-inserted region is determined from the candidate to-be-inserted regions. Further, maximum rectangle searching is performed in the target candidate to-be-inserted region to obtain a target to-be-inserted region in which an image is to be inserted.

Claims (108)

1. A method for detecting an information insertion region, the method comprising:

obtaining a video;

segmenting the video to obtain video fragments, each of the video fragments including a subset of image frames in the video;

obtaining a target frame in the video fragments;

identifying and segmenting objects in the target frame, to obtain labeling information corresponding to the objects;

determining a target object according to the labeling information;

performing clustering on the target object, to obtain a plurality of candidate to-be-inserted regions in the target object;

determining whether each of the plurality of candidate to-be-inserted regions includes at least one non-target object;

obtaining a subset of the plurality of candidate to-be-inserted regions that is determined not to include the at least one non-target object from the plurality of candidate to-be-inserted regions;

calculating areas of the subset of the plurality of candidate to-be-inserted regions;

determining, by processing circuitry, a target candidate to-be-inserted region in the target object from the subset of the plurality of candidate to-be-inserted regions in the target object based on the calculated areas of the subset of the plurality of candidate to-be-inserted regions; and

performing maximum rectangle searching in the target candidate to-be-inserted region in the target object to obtain a target to-be-inserted sub-region of the target object in which an image is to be inserted.

2. The method according to claim 1 , wherein the obtaining the video comprises:

transmitting a video obtaining request to a server, the video obtaining request indicating an identifier of the video; and

receiving the video corresponding to the identifier from the server.

3. The method according to claim 1 , wherein the segmenting the video comprises:

extracting target features of the image frames from the video;

performing similarity identification on the target features of adjacent image frames; and

segmenting the video according to an identification result of the similarity identification to obtain the video fragments.

4. The method according to claim 1 , wherein the identifying and the segmenting the objects comprises:

inputting the target frame into an instance segmentation model; and

identifying and segmenting the objects in the target frame by using the instance segmentation model, to obtain the labeling information corresponding to the objects.

5. The method according to claim 4 , wherein the identifying and the segmenting the objects comprises:

preprocessing the target frame by the instance segmentation model,

performing feature extraction on the target frame after the preprocessing, to obtain a feature image;

determining a plurality of candidate regions of interest on the feature image;

performing classification and regression on the plurality of candidate regions of interest, to obtain a target region of interest;

performing an alignment operation on the target region of interest, to align pixels in the target frame and pixels in the target region of interest; and

performing classification, bounding box regression, and mask generation on the target region of interest, to obtain the labeling information corresponding to the objects.

6. The method according to claim 1 , wherein the labeling information includes at least one of classification information, confidence levels, masks, or calibration boxes of the objects.

7. The method according to claim 6 , wherein

the determining the target object according to the labeling information includes determining the target object according to the classification information and areas of the masks in the labeling information; and

the performing clustering includes performing mean shift processing on the target object, to obtain the plurality of candidate to-be-inserted regions.

8. The method according to claim 7 , wherein

the video fragments include a plurality of target frames; and

before the determining the target object, the method further includes:

comparing confidence levels of objects included in each of the target frames with a preset confidence level threshold respectively;

reserving the objects of which the confidence levels are greater than the preset confidence level threshold in the respective target frames; and

deleting one or more of the target frames not including a to-be-inserted object, classification information of the to-be-inserted object and classification information of the target object being the same.

9. The method according to claim 7 , wherein the performing mean shift processing comprises:

using a pixel point in the target object as a target point;

determining a target range by using the target point as a center of a circle according to a preset radius;

determining a mean shift vector according to a distance vector between the target point and the pixel point within the target range;

moving the target point to an endpoint of the mean shift vector according to the mean shift vector;

determining the endpoint as the target point;

repeating the using the pixel point, the determining the target range, the determining the mean shift vector, the moving the target point, and the determining the endpoint until a position of the target point no longer changes;

determining pixel sets according to the pixel point corresponding to the target point of which the position no longer changes and pixel points within a range of the preset radius;

obtaining a distance between the pixel sets; and

comparing the distance with a preset distance threshold, to determine the candidate to-be- inserted region according to the comparison of the distance.

10. The method according to claim 9 , wherein the comparing the distance comprises:

when the distance is less than or equal to the preset distance threshold, merging two pixel sets corresponding to the distance, to form the candidate to-be-inserted region; and

when the distance is greater than the preset distance threshold, using the two pixel sets corresponding to the distance as the candidate to-be-inserted region respectively.

11. The method according to claim 1 , wherein the determining the target candidate to-be-inserted region comprises:

determining the candidate to-be-inserted region with a greatest area as the target candidate to-be-inserted region.

12. The method according to claim 1 , wherein before the performing the maximum rectangle searching, the method further comprises:

performing mean filtering on the target candidate to-be-inserted region, to obtain a uniform and smooth target candidate to-be-inserted region.

13. The method according to claim 12 , wherein the performing the maximum rectangle searching comprises:

for each of a plurality of adjacent pixel points with a same pixel value,

using the respective pixel point within the target candidate to-be-inserted region as a reference point,

searching for the adjacent pixel point with the same pixel value according to a pixel value of the reference point, and

when the adjacent pixel point exists, determining the adjacent pixel point as the reference point;

determining an initial pixel point of the plurality of adjacent pixel points as a vertex;

forming rectangles according to the vertex and the adjacent pixel points;

calculating areas of the rectangles;

selecting a target rectangle of the rectangles with a greatest area; and

determining a region corresponding to the target rectangle as the target to-be-inserted sub-region.

14. An apparatus for detecting an information insertion region, comprising:

processing circuitry configured to:

obtain a video;

segment the video to obtain video fragments, each of the video fragments including a subset of image frames in the video;

obtain a target frame in the video fragments;

identify and segment objects in the target frame, to obtain labeling information corresponding to the objects;

determine a target object according to the labeling information;

perform clustering on the target object, to obtain a plurality of candidate to-be- inserted regions in the target object;

determine whether each of the plurality of candidate to-be-inserted regions includes at least one non-target object;

obtain a subset of the plurality of candidate to-be-inserted regions that is determined not to include the at least one non-target object from the plurality of candidate to-be- inserted regions;

calculate areas of the subset of the plurality of candidate to-be-inserted regions;

determine a target candidate to-be-inserted region in the target object from the subset of the plurality of candidate to-be-inserted regions in the target object based on the calculated areas of the subset of the plurality of candidate to-be-inserted regions; and

perform maximum rectangle searching in the target candidate to-be-inserted region in the target object to obtain a target to-be-inserted sub-region of the target object in which an image is to be inserted.

15. The apparatus according to claim 14 , wherein the processing circuitry is configured to:

transmit a video obtaining request to a server, the video obtaining request indicating an identifier of the video; and

receive the video corresponding to the identifier from the server.

16. The apparatus according to claim 14 , wherein the processing circuitry is configured to:

extract target features of the image frames from the video;

perform similarity identification on the target features of adjacent image frames; and

segment the video according to an identification result of the similarity identification to obtain the video fragments.

17. The apparatus according to claim 14 , wherein the processing circuitry is configured to:

input the target frame into an instance segmentation model; and

identify and segment the objects in the target frame by using the instance segmentation model, to obtain the labeling information corresponding to the objects.

18. The apparatus according to claim 17 , wherein the processing circuitry is configured to:

preprocess the target frame by the instance segmentation model, perform feature extraction on the target frame after the preprocessing, to obtain a feature image;

determine a plurality of candidate regions of interest on the feature image;

perform classification and regression on the plurality of candidate regions of interest, to obtain a target region of interest;

perform an alignment operation on the target region of interest, to align pixels in the target frame and pixels in the target region of interest; and

perform classification, bounding box regression, and mask generation on the target region of interest, to obtain the labeling information corresponding to the objects.

19. The apparatus according to claim 14 , wherein the labeling information includes at least one of classification information, confidence levels, masks, or calibration boxes of the objects.

20. A non-transitory computer-readable storage medium, storing instructions which when executed by at least one processor cause the at least one processor to perform:

obtaining a video;

segmenting the video to obtain video fragments, each of the video fragments including a subset of image frames in the video;

obtaining a target frame in the video fragments;

identifying and segmenting objects in the target frame, to obtain labeling information corresponding to the objects;

determining a target object according to the labeling information;

performing clustering on the target object, to obtain a plurality of candidate to-be-inserted regions in the target object;

determining whether each of the plurality of candidate to-be-inserted regions includes at least one non-target object;

obtaining a subset of the plurality of candidate to-be-inserted regions that is determined not to include the at least one non-target object from the plurality of candidate to-be-inserted regions;

calculating areas of the subset of the plurality of candidate to-be-inserted regions;

determining a target candidate to-be-inserted region in the target object from the subset of the plurality of candidate to-be-inserted regions in the target object based on the calculated areas of the subset of the plurality of candidate to-be-inserted regions; and

performing maximum rectangle searching in the target candidate to-be-inserted region in the target object to obtain a target to-be-inserted sub-region of the target object in which an image is to be inserted.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE'S ADDRESS PREVIOUSLY RECORDED ON REEL 056795 FRAME 0697. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 31, 2021
From: SHENG, HUI; HUANG, DONGBO
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 057389/0436 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2021
From: HUANG, DONGBO; SHENG, HUI
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 056795/0697 →
Priority Claims (1)
CN 201910578322.7 · Jun 28, 2019 · national
Continuity (2)
Continuation PCTCN2020097782 · Jun 23, 2020
Related Publication 20210406549A1 · Dec 30, 2021