IP Library › Granted Patent US 11,495,264
Granted Patent B2
US 11,495,264 · App. 17/079,662 · Granted Nov 8, 2022

Method and system of clipping a video, computing device, and computer storage medium

Inventors: Heming Cai (Shanghai, CN); Long Qian (Shanghai, CN)
Assignee: Shanghai Bilibili Technology Co., Ltd.
G11B27/02G06K9/6215G06N20/00G06T7/248G06V20/41G06V20/46G06V40/10G06N3/08G06T2207/10016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,495,264
App. No.
17/079,662
Granted
Nov 8, 2022
Kind
B2
Abstract

Embodiments of the present disclosure describes techniques for clipping a video. The disclosed techniques comprise obtaining a video including a plurality of frames performing object detection on each frame; identifying objects contained in each frame, wherein a region where each object is located is selected through a detection box; classifying and recognizing the objects identified in each frame using a pre-trained classification model; selecting human body region images; determining a similarity between each human body region image selected from the plurality of frames and a target character image; in response to determining that a similarity between a human body region image and the target character image is greater than a predetermined threshold, identifying the human body region image as a clipping image; and synthesizing clipping images identified in the plurality of frames in order of time to obtain a clipping video.

Claims (74)

1. A method of clipping a video, comprising:

obtaining a video, the video comprising a plurality of frames;

performing object detection on each of the plurality of frames;

identifying objects contained in each of the plurality of frames, wherein a region where each object is located is selected through a detection box;

classifying and recognizing the identified objects in each of the plurality of frames using a classification model, the classification model being pre-trained;

selecting human body region images based on the classifying and recognizing the objects;

determining a similarity between each of the human body region images selected from the plurality of frames and a target character image;

in response to determining that a similarity between a human body region image among the human body region images and the target character image is greater than a predetermined threshold, identifying the human body region image as a clipping image, wherein the identifying the human body region image as a clipping image further comprises setting a clipping box based on a detection box corresponding to the human body region image, wherein the setting a clipping box based on a detection box corresponding to the human body region image further comprises determining a moving speed of the detection box, and wherein the determining a moving speed of the detection box further comprises:

determining whether a distance between center points of the detection box in adjacent frames among the plurality of frames is greater than a predetermined distance value, and

identifying an average speed of the detection box in a unit frame as the moving speed of the detection box in response to determining that the distance between the center points of the detection box in the adjacent frames is greater than the predetermined distance value; and

synthesizing clipping images identified in the plurality of frames in order of time to obtain a clipping video.

2. The method of 1 , further comprising:

performing the object detection on each of the plurality of frames using a pre-trained object detection model to identify the objects contained in each of the plurality of frames.

3. The method of claim 1 , wherein training the classification model comprises:

classifying images to be processed using a sample character image as a reference object;

identifying an image among the images with a same category as the sample character image as positive sample data, and identifying another image among the images with a different category from the sample character image as negative sample data; and

adjusting an inter-class distance between the positive sample data and the negative sample data based on Triplet loss to enlarge a difference between the positive sample data and the negative sample data.

4. The method of claim 1 , wherein the determining a similarity between each of the human body region images and a target character image further comprises:

extracting multiple first feature vectors of each of the human body region images to obtain an n-dimensional first feature vector;

extracting multiple second feature vectors of the target character image to obtain an m-dimensional second feature vector; wherein n≤m, and both n and m are positive integers; and

determining a Euclidean distance between the first feature vector and the second feature vector, the Euclidean distance being indicative of the similarity between each of the human body region images and a target character image.

5. The method of claim 1 , wherein the clipping box includes the human body region image and the similarity correspond to the human body region image, and wherein the identifying the human body region image as a clipping image in response to determining that a similarity between a human body region image and the target character image is greater than a predetermined threshold further comprises:

identifying the clipping box including the human body region image and the corresponding similarity, and selecting the human body region image with the similarity being greater than the predetermined threshold in the clipping box as the clipping image.

6. The method of claim 5 , wherein the setting a clipping box based on a detection box corresponding to the human body region image further comprises:

identifying the moving speed of the detection box as a moving speed of the clipping box.

7. A system of clipping a video, comprising:

at least one processor; and

at least one memory communicatively coupled to the at least one processor and storing instructions that upon execution by the at least one processor cause the system to perform operations, the operations comprising:

obtaining a video, the video comprising a plurality of frames;

performing object detection on each of the plurality of frames;

identifying objects contained in each of the plurality of frames, wherein a region where each object is selected through a detection box;

classifying and recognizing the objects identified in each of the plurality of frames using a classification model, the classification model being pre-trained;

selecting human body region images based on the classifying and recognizing the objects;

determining a similarity between each of the human body region images selected from the plurality of frames and a target character image;

in response to determining that a similarity between a human body region image among the human body region images and the target character image is greater than a predetermined threshold, identifying the human body region image as a clipping image, wherein the identifying the human body region image as a clipping image further comprises setting a clipping box based on a detection box corresponding to the human body region image, wherein the setting a clipping box based on a detection box corresponding to the human body region image further comprises determining a moving speed of the detection box, and wherein the determining a moving speed of the detection box further comprises:

determining whether a distance between center points of the detection box in adjacent frames among the plurality of frames is greater than a predetermined distance value, and

identifying an average speed of the detection box in a unit frame as the moving speed of the detection box in response to determining that the distance between the center points of the detection box in the adjacent frames is greater than the predetermined distance value; and

synthesizing clipping images identified in the plurality of frames in order of time to obtain a clipping video.

8. The system of claim 7 , the operations further comprising:

performing the object detection on each of the plurality of frames using a pre-trained object detection model to identify the objects contained in each of the plurality of frames.

9. The system of claim 7 , wherein training the classification model comprises:

classifying images to be processed using a sample character image as a reference object;

identifying an image among the images with a same category as the sample character image as positive sample data, and identifying another image among the images with a different category from the sample character image as negative sample data: and adjusting an inter-class distance between the positive sample data and the negative sample data based on Triplet loss to enlarge a difference between the positive sample data and the negative sample data.

10. The system of claim 7 , wherein the determining a similarity between each of the human body region images and a target character image further comprises:

extracting multiple first feature vectors of each of the human body region images to obtain an n-dimensional first feature vector;

extracting multiple second feature vectors of the target character image to obtain an m-dimensional second feature vector; wherein n≤m, and both n and m are positive integers;

determining a Euclidean distance between the first feature vector and the second feature vector, the Euclidean distance being indicative of the similarity between each of the human body region images and a target character image.

11. The system of claim 7 , wherein the clipping box includes the human body region image and the similarity correspond to the human body region image, and wherein the identifying the human body region image as a clipping image in response to determining that a similarity between a human body region image and the target character image is greater than a predetermined threshold further comprises:

identifying the clipping box including the human body region image and the corresponding similarity, and selecting the human body region image with the similarity being greater than the predetermined threshold in the clipping box as the clipping image.

12. The system of claim 11 , wherein the setting a clipping box based on a detection box corresponding to the human body region image further comprises:

identifying the moving speed the detection box as a moving speed of the clipping box.

13. A non-transitory computer-readable storage medium, storing computer-readable instructions that upon execution by a processor cause the processor to implement operations comprising:

obtaining a video, the video comprising a plurality of frames;

performing object detection on each of the plurality of frames;

identifying objects contained in each of the plurality of frames, wherein a region where each object is located is selected through a detection box;

classifying and recognizing the objects identified in each of the plurality of frames using a classification model, the classification model pre-trained;

selecting human body region images based on the classifying and recognizing the objects;

determining a similarity between each of the human body region images selected from the plurality of frames and a target character image;

in response to determining that a similarity between a human body region image among the human body region images and the target character image is greater than a predetermined threshold, identifying the human body region image as a clipping image, wherein the identifying the human body region image as a clipping image further comprises setting a clipping box based on a detection box corresponding to the human body region image, wherein the setting a clipping box based on a detection box corresponding to the human body region image further comprises

determining a moving speed of the detection box, and wherein the determining a moving speed of the detection box further comprises:

determining whether a distance between center points of the detection box in adjacent frames among the plurality of frames is greater than a predetermined distance value, and identifying an average speed of the detection box in a unit frame as the moving speed of the detection box in response to determining that the distance between the center points of the detection box in the adjacent frames is greater than the predetermined distance value; and

synthesizing clipping images identified in the plurality of frames in order of time to obtain a clipping video.

14. The non-transitory computer-readable storage medium of claim 13 , wherein training the classification model comprises:

classifying images to be processed using a sample character image as a reference object;

identifying an image among the images with a same category as the sample character image as positive sample data, and identifying another image among the images with a different category from the sample character image as negative sample data; and

adjusting an inter-class distance between the positive sample data and the negative sample data based on Triplet loss to enlarge a difference between the positive sample data and the negative sample data.

15. The non-transitory computer-readable storage medium of claim 13 , wherein the determining a similarity between each of the human region images and a target character image further comprises:

extracting multiple first feature vectors of each of the human body region images to obtain an n-dimensional first feature vector;

extracting multiple second feature vectors of the target character image to obtain an m-dimensional second feature vector; wherein n≤m, and both n and m are positive integers; and

determining a Euclidean distance between the first feature vector and the second feature vector, the Euclidean distance being indicative of the similarity between each of the human body region images and a target character image.

16. The non-transitory computer-readable storage medium of claim 13 , wherein the clipping box includes the human body region image and the similarity correspond to the human body region image, and wherein the identifying the human body region image as a clipping image in response to determining that a similarity between a human body region image and the target character image is greater than a predetermined threshold further comprises:

identifying the clipping box including the human body region image and the corresponding similarity, and selecting the human body region image with the similarity being greater than the predetermined threshold in the clipping box as the clipping image.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the setting a clipping box based on a detection box corresponding to the human body region image further comprises:

identifying the moving speed of the detection box as a moving speed of the clipping box.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2020
From: CAI, HEMING; QIAN, LONG
To: SHANGHAI BILIBILI TECHNOLOGY CO. LTD.
Reel/Frame 054162/0430 →
Continuity (1)
Related Publication 20210125639A1 · Apr 29, 2021