IP Library › Granted Patent US 11,727,688
Granted Patent B2
US 11,727,688 · App. 17/473,940 · Granted Aug 15, 2023

Method and apparatus for labelling information of video frame, device, and storage medium

Inventors: Ruizheng Wu (Shenzhen, CN); Jiaya Jia (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G06V20/48G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,727,688
App. No.
17/473,940
Granted
Aug 15, 2023
Kind
B2
Abstract

A method and apparatus for labelling information of a video frame, includes: obtaining a video; performing feature extraction on a target video frame in the video, to obtain a target image feature of the target video frame; determining, according to image feature matching degrees between the target video frame and labelled video frames, a guide video frame of the target video frame from the labelled video frames, the guide video frame being used for guiding the target video frame for information labelling, and the image feature matching degrees being matching degrees between the target image feature and image features corresponding to the labelled video frames; and generating target label information corresponding to the target video frame according to label information corresponding to the guide video frame.

Claims (81)

1. A method for labelling information of a video frame, applied to a computer device, the method comprising:

obtaining a video;

performing feature extraction on a target video frame in the video, to obtain a target image feature of the target video frame;

determining, according to image feature matching degrees between the target video frame and labelled video frames, a guide video frame of the target video frame from the labelled video frames, the labelled video frames belonging to the video, the guide video frame being used for guiding the target video frame for information labelling, the image feature matching degrees being matching degrees between the target image feature and image features corresponding to the labelled video frames, and an image feature matching degree between the guide video frame and the target video frame being higher than image feature matching degrees between other labelled video frames and the target video frame; and

generating target label information corresponding to the target video frame according to label information corresponding to the guide video frame,

wherein determining the guide video frame of the target video frame from the labelled video frames comprises:

obtaining a labelled object image feature of a labelled object in an initially labelled video frame, the initially labelled video frame being a video frame with preset label information in the video, and the labelled object being an object comprising label information in the initially labelled video frame;

obtaining a candidate image feature from a memory pool of a memory selection network (MSN), the MSN comprising the memory pool and a selection network, and the memory pool storing the image features of the labelled video frames;

inputting the candidate image feature, the labelled object image feature, and the target image feature into the selection network to obtain an image feature score outputted by the selection network, the image feature score indicating an image feature matching degree between the candidate image feature and the target image feature; and

determining a labelled video frame corresponding to a highest image feature score as the guide video frame; and

the method further comprises: storing the target image feature of the target video frame into the memory pool.

2. The method according to claim 1 , wherein the selection network comprises a first selection branch and a second selection branch; and

the inputting the candidate image feature, the target image feature, and the labelled object image feature into the selection network to obtain the image feature score outputted by the selection network comprises:

performing an association operation on any two image features among the candidate image feature, the target image feature, and the labelled object image feature to obtain three associated image features, each associated image feature indicating a similarity between the two corresponding image features;

concatenating the associated image features, and inputting the concatenated associated image features into the first selection branch, to obtain a first feature vector outputted by the first selection branch;

concatenating the candidate image feature, the target image feature, and the labelled object image feature to obtain a concatenated result, and inputting the concatenated result into the second selection branch to obtain a second feature vector outputted by the second selection branch; and

determining the image feature score according to the first feature vector and the second feature vector.

3. The method according to claim 1 , wherein the obtaining a candidate image feature from a memory pool of an MSN comprises:

obtaining, when a frame rate of the video is greater than a frame rate threshold, candidate image features corresponding to labelled video frames from the memory pool every predetermined quantity of frames, or obtaining candidate image features of n adjacent labelled video frames corresponding to the target video frame from the memory pool, n being a positive integer.

4. The method according to claim 1 , wherein the generating target label information corresponding to the target video frame according to label information corresponding to the guide video frame comprises:

inputting the guide video frame, the label information corresponding to the guide video frame, and the target video frame into a temporal propagation network (TPN), to obtain the target label information outputted by the TPN.

5. The method according to claim 4 , wherein the TPN comprises an appearance branch and a motion branch; and

the inputting the guide video frame, the label information corresponding to the guide video frame, and the target video frame into a TPN, to obtain the target label information outputted by the TPN comprises:

inputting the label information corresponding to the guide video frame and the target video frame into the appearance branch, to obtain an image information feature outputted by the appearance branch;

determining a video frame optical flow between the guide video frame and the target video frame, and inputting the video frame optical flow and the label information corresponding to the guide video frame into the motion branch to obtain a motion feature outputted by the motion branch; and

determining the target label information according to the image information feature and the motion feature.

6. The method according to claim 4 , wherein before the obtaining a video, the method further comprises:

training the TPN according to a sample video, sample video frames in the sample video comprising label information;

inputting a target sample video frame in the sample video and other sample video frames in the sample video into the TPN to obtain predicted sample label information outputted by the TPN;

determining a sample guide video frame in the sample video frames according to the predicted sample label information and sample label information corresponding to the target sample video frame; and

training the MSN according to the target sample video frame and the sample guide video frame.

7. The method according to claim 6 , wherein the determining a sample guide video frame in the sample video frames according to the predicted sample label information and sample label information corresponding to the target sample video frame comprises:

calculating information accuracy between the predicted sample label information and the sample label information; and

determining a positive sample guide video frame and a negative sample guide video frame in the sample video frames according to the information accuracy,

wherein first information accuracy of performing information labelling on the target sample video frame according to the positive sample guide video frame is higher than second information accuracy of performing information labelling on the target sample video frame according to the negative sample guide video frame.

8. An apparatus for labelling information of a video frame, comprising a processor and a memory, the memory storing computer instructions that, when being loaded and executed by the processor, cause the processor to:

obtain a video;

perform feature extraction on a target video frame in the video, to obtain a target image feature of the target video frame;

determine, according to image feature matching degrees between the target video frame and labelled video frames, a guide video frame of the target video frame from the labelled video frames, the labelled video frames belonging to the video, the guide video frame being used for guiding the target video frame for information labelling, the image feature matching degrees being matching degrees between the target image feature and image features corresponding to the labelled video frames, and an image feature matching degree between the guide video frame and the target video frame being higher than image feature matching degrees between other labelled video frames and the target video frame; and

generate target label information corresponding to the target video frame according to label information corresponding to the guide video frame,

wherein the computer instructions further causes the processor to:

obtain a labelled object image feature of a labelled object in an initially labelled video frame, the initially labelled video frame being a video frame with preset label information in the video, and the labelled object being an object comprising label information in the initially labelled video frame;

obtain a candidate image feature from a memory pool of a memory selection network (MSN), the MSN comprising the memory pool and a selection network, and the memory pool storing the image features of the labelled video frames;

input the candidate image feature, the labelled object image feature, and the target image feature into the selection network to obtain an image feature score outputted by the selection network, the image feature score indicating an image feature matching degree between the candidate image feature and the target image feature;

determine a labelled video frame corresponding to a highest image feature score as the guide video frame; and

store the target image feature of the target video frame into the memory pool.

9. The apparatus according to claim 8 , wherein the selection network comprises a first selection branch and a second selection branch; and

the computer instructions further causes the processor to:

perform an association operation on any two image features among the candidate image feature, the target image feature, and the labelled object image feature to obtain three associated image features, each associated image feature indicating a similarity between the two corresponding image features;

concatenate associated image features, and input the concatenated associated image features into the first selection branch, to obtain a first feature vector outputted by the first selection branch;

concatenate the candidate image feature, the target image feature, and the labelled object image feature to obtain a concatenated result, and input the concatenated result into the second selection branch to obtain a second feature vector outputted by the second selection branch; and

determine the image feature score according to the first feature vector and the second feature vector.

10. The apparatus according to claim 8 , wherein the computer instructions further causes the processor to:

obtaining, when a frame rate of the video is greater than a frame rate threshold, candidate image features corresponding to labelled video frames from the memory pool every predetermined quantity of frames, or obtaining candidate image features of n adjacent labelled video frames corresponding to the target video frame from the memory pool, n being a positive integer.

11. The apparatus according to claim 8 , wherein the computer instructions further causes the processor to:

input the guide video frame, the label information corresponding to the guide video frame, and the target video frame into a temporal propagation network (TPN), to obtain the target label information outputted by the TPN.

12. The apparatus according to claim 11 , wherein the TPN comprises an appearance branch and a motion branch; and

the inputting the guide video frame, the label information corresponding to the guide video frame, and the target video frame into a TPN, to obtain the target label information outputted by the TPN comprises:

inputting the label information corresponding to the guide video frame and the target video frame into the appearance branch, to obtain an image information feature outputted by the appearance branch;

determining a video frame optical flow between the guide video frame and the target video frame, and inputting the video frame optical flow and the label information corresponding to the guide video frame into the motion branch to obtain a motion feature outputted by the motion branch; and

determining the target label information according to the image information feature and the motion feature.

13. The apparatus according to claim 11 , wherein before the obtaining a video, the computer instructions further causes the processor to perform:

training the TPN according to a sample video, sample video frames in the sample video comprising label information;

inputting a target sample video frame in the sample video and other sample video frames in the sample video into the TPN to obtain predicted sample label information outputted by the TPN;

determining a sample guide video frame in the sample video frames according to the predicted sample label information and sample label information corresponding to the target sample video frame; and

training the MSN according to the target sample video frame and the sample guide video frame.

14. The apparatus according to claim 13 , wherein the determining a sample guide video frame in the sample video frames according to the predicted sample label information and sample label information corresponding to the target sample video frame comprises:

calculating information accuracy between the predicted sample label information and the sample label information; and

determining a positive sample guide video frame and a negative sample guide video frame in the sample video frames according to the information accuracy,

wherein first information accuracy of performing information labelling on the target sample video frame according to the positive sample guide video frame is higher than second information accuracy of performing information labelling on the target sample video frame according to the negative sample guide video frame.

15. A non-transitory computer-readable storage medium, storing at least one instruction, at least one program, a code set, or an instruction set, the at least one instruction, the at least one program, the code set, or the instruction set being loaded and executed by a processor to implement:

obtaining a video;

performing feature extraction on a target video frame in the video, to obtain a target image feature of the target video frame;

determining, according to image feature matching degrees between the target video frame and labelled video frames, a guide video frame of the target video frame from the labelled video frames, the labelled video frames belonging to the video, the guide video frame being used for guiding the target video frame for information labelling, the image feature matching degrees being matching degrees between the target image feature and image features corresponding to the labelled video frames, and an image feature matching degree between the guide video frame and the target video frame being higher than image feature matching degrees between other labelled video frames and the target video frame; and

generating target label information corresponding to the target video frame according to label information corresponding to the guide video frame,

wherein determining the guide video frame of the target video frame from the labelled video frames comprises:

obtaining a labelled object image feature of a labelled object in an initially labelled video frame, the initially labelled video frame being a video frame with preset label information in the video, and the labelled object being an object comprising label information in the initially labelled video frame;

obtaining a candidate image feature from a memory pool of a memory selection network (MSN), the MSN comprising the memory pool and a selection network, and the memory pool storing the image features of the labelled video frames;

inputting the candidate image feature, the labelled object image feature, and the target image feature into the selection network to obtain an image feature score outputted by the selection network, the image feature score indicating an image feature matching degree between the candidate image feature and the target image feature; and

determining a labelled video frame corresponding to a highest image feature score as the guide video frame; and

the at least one instruction, the at least one program, the code set, or the instruction set further cause the processor to implement: storing the target image feature of the target video frame into the memory pool.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2021
From: WU, RUIZHENG; JIA, JIAYA
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 057468/0769 →
Priority Claims (1)
CN 201910807774.8 · Aug 29, 2019 · national
Continuity (2)
Continuation PCTCN2020106575 · Aug 3, 2020
Related Publication 20210406553A1 · Dec 30, 2021