IP Library › Granted Patent US 11,676,385
Granted Patent B1
US 11,676,385 · App. 17/816,990 · Granted Jun 13, 2023

Processing method and apparatus, terminal device and medium

Inventors: Ye Yuan (Los Angeles, CA); Yufei Wang (Los Angeles, CA); Longyin Wen (Los Angeles, CA)
Assignee: LEMON INC.
G06V20/46G06V10/462G06V20/635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,676,385
App. No.
17/816,990
Granted
Jun 13, 2023
Kind
B1
Abstract

A target video and video description information corresponding to the target video are acquired; salient object information of the target video is determined; a key frame category of the video description information is determined; and the target video, the video description information, the salient object information and the key frame category are input into a processing model to obtain a timestamp of an image corresponding to the video description information in the target video.

Claims (58)

1. A processing method, comprising:

acquiring a target video and video description information of the target video;

determining salient object information of the target video;

extracting at least one key frame from the target video according to the video description information, and determining a category of each key frame of the at least one key frame;

inputting the target video into a first information extraction module in the processing model to obtain image information and first text information of the target video;

inputting the video description information into a second information extraction module in the processing model to obtain second text information of the target video, wherein the salient object information, the image information and the first text information are video-related features and the second text information and the category of each key frame are description-related features;

inputting the video-related features and the description-related features into the retrieval module in the processing model to perform matching processes on the video-related features and the description-related features to obtain, from the target video, frames matching the description-related features, candidate timestamps of the frames, and matching degrees of the frames to the description-related features; and

determining a timestamp of an image corresponding to the video description information in the target video according to the matching degrees corresponding to the plurality of candidate timestamps.

2. The method according to claim 1 , wherein inputting the target video into the first information extraction module in the processing model to obtain the image information and the first text information of the target video comprises:

after inputting the target video into the first information extraction module in the processing model, obtaining a first target object of the target video through sparse frame extraction;

performing image information extraction on the first target object to obtain the image information;

extracting subtitle information of the target video; and

performing text information extraction on the subtitle information to obtain the first text information.

3. The method according to claim 1 , wherein determining the salient object information of the target video comprises:

performing sparse frame extraction processing on the target video to obtain a second target object; and

determining salient object information corresponding to the second target object.

4. The method according to claim 3 , wherein the second target object comprises a frame and/or a video clip.

5. The method according to claim 1 , wherein determining the category of each key frame comprises:

inputting the video description information into a key frame category prediction model to obtain the category of each key frame.

6. A processing apparatus, comprising:

one or more processors; and

a storage apparatus configured to store one or more programs;

wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement:

acquiring a target video and video description information of the target video;

determining salient object information of the target video;

extracting at least one key frame from the target video according to the video description information, and determining a category of each key frame of the at least one key frame;

inputting the target video into a first information extraction module in the processing model to obtain image information and first text information of the target video;

inputting the video description information into a second information extraction module in the processing model to obtain second text information of the target video, wherein the salient object information, the image information and the first text information are video-related features and the second text information and the category of each key frame are description-related features;

inputting the video-related features and the description-related features into the retrieval module in the processing model to perform matching processes on the video-related features and the description-related features to obtain, from the target video, frames matching the description-related features, candidate timestamps of the frames, and matching degrees of the frames to the description-related features; and

determining a timestamp of an image corresponding to the video description information in the target video according to the matching degrees corresponding to the plurality of candidate timestamps.

7. The apparatus according to claim 6 , wherein the one or more processors inputs the target video into the first information extraction module in the processing model to obtain the image information and the first text information by:

after inputting the target video into the first information extraction module in the processing model, obtaining a first target object of the target video through sparse frame extraction;

performing image information extraction on the first target object to obtain the image information;

extracting subtitle information of the target video; and

performing text information extraction on the subtitle information to obtain the first text information.

8. The apparatus according to claim 6 , wherein the one or more processors determines the salient object information of the target video by:

performing sparse frame extraction processing on the target video to obtain a second target object; and

determining salient object information corresponding to the second target object.

9. The apparatus according to claim 8 , wherein the second target object comprises a frame and/or a video clip.

10. The apparatus according to claim 6 , wherein the one or more processors determines the category of each key frame comprises:

inputting the video description information into a key frame category prediction model to obtain the category of each key frame.

11. A non-transitory computer-readable storage medium storing a computer program which, when executed by a processor, implements:

acquiring a target video and video description information corresponding to the target video;

determining salient object information of the target video;

extracting at least one key frame from the target video according to the video description information, and determining a category of each key frame of the at least one key frame;

inputting the target video into a first information extraction module in the processing model to obtain image information and first text information of the target video;

inputting the video description information into a second information extraction module in the processing model to obtain second text information of the target video, wherein the salient object information, the image information and the first text information are video-related features and the second text information and the category of each key frame are description-related features;

inputting the video-related features and the description-related features into the retrieval module in the processing model to perform matching processes on the video-related features and the description-related features to obtain, from the target video, frames matching the description-related features, candidate timestamps of the frames, and matching degrees of the frames to the description-related features; and

determining a timestamp of an image corresponding to the video description information in the target video according to the matching degrees corresponding to the plurality of candidate timestamps.

12. The non-transitory computer-readable storage medium according to claim 11 , wherein the computer program inputs the target video into the first information extraction module in the processing model to obtain the image information and the first text information by:

after inputting the target video into the first information extraction module in the processing model, obtaining a first target object of the target video through sparse frame extraction;

performing image information extraction on the first target object to obtain the image information;

extracting subtitle information of the target video; and

performing text information extraction on the subtitle information to obtain the first text information.

13. The non-transitory computer-readable storage medium according to claim 11 , wherein the computer program determines the salient object information of the target video by:

performing sparse frame extraction processing on the target video to obtain a second target object; and

determining salient object information corresponding to the second target object.

14. The non-transitory computer-readable storage medium according to claim 13 , wherein the second target object comprises a frame and/or a video clip.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 17, 2023
From: YUAN, YE; WANG, YUFEI; WEN, LONGYIN
To: BYTEDANCE INC.
Reel/Frame 062731/0273 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 17, 2023
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 062731/0330 →
Priority Claims (1)
CN 202210365435.0 · Apr 7, 2022 · national
Cited By (2)
US 12,277,766 US 12,556,781