IP Library Granted Patent US 12700429
Granted Patent B2
US 12700429 · App. 18/832,049 · Granted Aug 4, 2026

Methods, devices, readable media and electronic devices for video processing

Inventors: Jia Sun (Beijing, CN); Zehuan Yuan (Beijing, CN)
Assignee: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
G11B27/036G06V20/41G06V20/49G06V20/62G06V30/19093
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12700429
App. No.
18/832,049
Granted
Aug 4, 2026
Kind
B2
Abstract

The disclosure relates to a video processing method, device, readable medium and electronic equipment. The method includes: obtaining key information corresponding to a video to be processed; extracting one or more target video clips from the video to be processed based on the key information; and obtaining a target video by adding the target video clip to a position preceding a specified video frame in the video to be processed, the specified video frame being any video frame of a preset number of video frames in the video to be processed.

Claims (72)

1 . A method for video processing, comprising:

obtaining key information corresponding to a video to be processed;

extracting one or more target video clips from the video to be processed based on the key information, wherein the extracting one or more target video clips from the video to be processed based on the key information comprises:

obtaining one or more candidate video clips of the video to be processed based on image information and audio information of the video to be processed, wherein the one or more candidate video clips comprise a piece of complete semantic information in the video to be processed;

obtaining a similarity between each candidate video clip and the key information; and

determining one or more candidate video clips with a highest similarity as the one or more target video clips; and

obtaining a target video by adding the target video clip to a position preceding a specified video frame in the video to be processed, the specified video frame being any video frame of the first N video frames in the video to be processed, wherein N is a fixed number.

2 . The method of claim 1 , wherein the adding the target video clip to a position preceding the specified video frame in the video to be processed to obtain the target video comprises:

updating the target video clip based on predetermined description information, the predetermined description information being used to characterize the key information of the target video clip corresponding to the video to be processed; and

obtaining the target video by adding the updated target video clip to the position preceding the specified video frame in the video to be processed.

3 . The method of claim 1 , wherein the obtaining one or more candidate video clips of the video to be processed based on the image information and audio information of the video to be processed comprises:

obtaining one or more pending video clips of the video to be processed based on the image information, each of the pending video clips comprising one or more frames of images;

obtaining one or more pending audio clips of the video to be processed based on the audio information; and

determining the one or more candidate video clips based on the pending video clip and the pending audio clip.

4 . The method of claim 3 , wherein the obtaining one or more pending video clips of the video to be processed based on the image information comprises:

obtaining an image text of each frame image of the video to be processed;

calculating a text similarity of image texts between adjacent frame images; and

determining the pending video clip based on the text similarity.

5 . The method of claim 3 , wherein the obtaining one or more pending audio clips of the video to be processed based on the audio information comprises:

obtaining an audio text corresponding to the audio information;

performing a sentence segmentation inference on the audio text to obtain sentence segmentation information in the audio text; and

determining the one or more pending audio clips based on the sentence segmentation information.

6 . The method of claim 3 , wherein the adding the target video clip to a position preceding the specified video frame in the video to be processed to obtain the target video comprises:

updating the target video clip based on predetermined description information, the predetermined description information being used to characterize the key information of the target video clip corresponding to the video to be processed; and

obtaining the target video by adding the updated target video clip to the position preceding the specified video frame in the video to be processed.

7 . The method of claim 3 , wherein the determining the candidate video clip based on the pending video clip and the pending audio clip comprises:

determining a correspondence between the pending audio clip and the pending video clip; and

for each pending video clip, based on a pending audio clip corresponding to said pending video clip, performing an integrity correction on said pending video clip to obtain a candidate video clip.

8 . The method of claim 7 , wherein the adding the target video clip to a position preceding the specified video frame in the video to be processed to obtain the target video comprises:

updating the target video clip based on predetermined description information, the predetermined description information being used to characterize the key information of the target video clip corresponding to the video to be processed; and

obtaining the target video by adding the updated target video clip to the position preceding the specified video frame in the video to be processed.

9 . A non-transitory computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processing device, cause the processing device to:

obtain key information corresponding to a video to be processed;

extract one or more target video clips from the video to be processed based on the key information, wherein the extracting one or more target video clips from the video to be processed based on the key information comprises:

obtaining one or more candidate video clips of the video to be processed based on image information and audio information of the video to be processed, wherein the one or more candidate video clips comprise a piece of complete semantic information in the video to be processed;

obtaining a similarity between each candidate video clip and the key information; and

determining one or more candidate video clips with a highest similarity as the one or more target video clips; and

obtain a target video by adding the target video clip to a position preceding a specified video frame in the video to be processed, the specified video frame being any video frame the first N video frames in the video to be processed, wherein N is a fixed number.

10 . The non-transitory computer-readable medium of claim 9 , wherein the processing device is further caused to:

obtain one or more pending video clips of the video to be processed based on image information, each of the pending video clips comprising one or more frames of images;

obtain one or more pending audio clips of the video to be processed based on the audio information; and

determine one or more candidate video clips based on the pending video clip and the pending audio clip.

11 . An electronic device, comprising:

a storage device having a computer program stored thereon;

a processing device for executing the computer program on a storage device to perform:

obtaining key information corresponding to a video to be processed;

extracting one or more target video clips from the video to be processed based on the key information, wherein the extracting one or more target video clips from the video to be processed based on the key information comprises:

obtaining one or more candidate video clips of the video to be processed based on image information and audio information of the video to be processed, wherein the one or more candidate video clips comprise a piece of complete semantic information in the video to be processed;

obtaining a similarity between each candidate video clip and the key information; and

determining one or more candidate video clips with a highest similarity as the one or more target video clips; and

obtaining a target video by adding the target video clip to a position preceding a specified video frame in the video to be processed, the specified video frame being any video frame of the first N video frames in the video to be processed, wherein N is a fixed number.

12 . The electronic device of claim 11 , wherein the obtaining one or more candidate video clips of the video to be processed based on the image information and audio information of the video to be processed comprises:

obtaining one or more pending video clips of the video to be processed based on the image information, each of the pending video clips comprising one or more frames of images;

obtaining one or more pending audio clips of the video to be processed based on the audio information; and

determining the one or more candidate video clips based on the pending video clip and the pending audio clip.

13 . The electronic device of claim 12 , wherein the obtaining one or more pending video clips of the video to be processed based on the image information comprises:

obtaining an image text of each frame image of the video to be processed;

calculating a text similarity of image texts between adjacent frame images; and

determining the pending video clip based on the text similarity.

14 . The electronic device of claim 12 , wherein the obtaining one or more pending audio clips of the video to be processed based on the audio information comprises:

obtaining an audio text corresponding to the audio information;

performing a sentence segmentation inference on the audio text to obtain sentence segmentation information in the audio text; and

determining the one or more pending audio clips based on the sentence segmentation information.

15 . The electronic device of claim 12 , wherein the determining the candidate video clip based on the pending video clip and the pending audio clip comprises:

determining a correspondence between the pending audio clip and the pending video clip; and

for each pending video clip, based on a pending audio clip corresponding to said pending video clip, performing an integrity correction on said pending video clip to obtain a candidate video clip.

16 . The electronic device of claim 15 , wherein the adding the target video clip to a position preceding the specified video frame in the video to be processed to obtain the target video comprises:

updating the target video clip based on predetermined description information, the predetermined description information being used to characterize the key information of the target video clip corresponding to the video to be processed; and

obtaining the target video by adding the updated target video clip to the position preceding the specified video frame in the video to be processed.

17 . The electronic device of claim 11 , wherein the adding the target video clip to a position preceding the specified video frame in the video to be processed to obtain the target video comprises:

updating the target video clip based on predetermined description information, the predetermined description information being used to characterize the key information of the target video clip corresponding to the video to be processed; and

obtaining the target video by adding the updated target video clip to the position preceding the specified video frame in the video to be processed.