IP Library › Granted Patent US 11,646,050
Granted Patent B2
US 11,646,050 · App. 17/212,037 · Granted May 9, 2023

Method and apparatus for extracting video clip

Inventors: Qinyi Zhang (Beijing, CN); Caihong Ma (Beijing, CN)
Assignee: Beijing Baidu Netcom Science and Technology Co., Ltd.
G10L25/54G06F16/7834G06F18/2413G06N3/04G06V10/454G06V10/82G06V20/41G06V20/46G06V20/49
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,646,050
App. No.
17/212,037
Granted
May 9, 2023
Kind
B2
Abstract

The present disclosure discloses a method and apparatus for extracting a video clip, relates to the field of artificial intelligence technology such as video processing, audio processing, and cloud computing. The method includes: acquiring a video, and extracting an audio stream in the video; determining a confidence that audio data in each preset period in the audio stream comprises a preset feature; and extracting a target video clip corresponding to a location of a target audio clip in the video; wherein the target audio clip is an audio clip within a continuous preset period, and has a confidence that the audio data includes the preset feature, which is larger than a preset confidence threshold. This method may improve the accuracy of extracting a video clip.

Claims (48)

1. A method for extracting a video clip, the method comprising:

acquiring a video, and extracting an audio stream in the video;

determining a confidence that audio data in each preset period in the audio stream comprises a preset feature; and

extracting a target video clip corresponding to a location of a target audio clip in the video; wherein the target audio clip is an audio clip within a continuous preset period, and has a confidence that the audio data comprises the preset feature, which is larger than a preset confidence threshold.

2. The method according to claim 1 , wherein the determining comprises:

sliding on the audio stream with a preset time step using a window of a preset length, to extract an audio stream clip with each of the preset length;

determining a confidence that each of the audio stream clip comprises the preset feature; and

determining, for the audio data in each preset period, the confidence that the audio data of the preset period comprises the preset feature based on the confidence that each of the audio stream clip to which the audio data of the preset period belongs comprises the preset feature.

3. The method according to claim 1 , wherein the preset confidence threshold is a plurality of preset confidence thresholds, and the extracting the target video clip comprises:

determining the audio clip within a continuous preset period, the confidence of which is above the preset confidence threshold, for each preset confidence threshold of the preset confidence thresholds;

determining the target audio clip in a plurality of audio clips determined based on the preset confidence thresholds; and

extracting the target video clip corresponding to the location of the target audio clip in the video.

4. The method according to claim 1 , wherein the determining the confidence comprises:

using a neural network classification model to determine the confidence that the audio data of each preset period in the audio stream comprises the preset feature.

5. The method according to claim 1 , wherein the preset feature comprises a feature representing that a spectrum change of the audio data in the audio stream exceeds a preset spectrum change threshold.

6. An electronic device, comprising:

at least one processor; and

a memory, communicatively connected to the at least one processor; wherein,

the memory, storing instructions executable by the at least one processor, the instructions, when executed by the at least one processor, cause the at least one processor to perform an operation for extracting a video clip, comprising:

acquiring a video, and extracting an audio stream in the video;

determining a confidence that audio data in each preset period in the audio stream comprises a preset feature; and

extracting a target video clip corresponding to a location of a target audio clip in the video; wherein the target audio clip is an audio clip within a continuous preset period, and has a confidence that the audio data comprises the preset feature, which is larger than a preset confidence threshold.

7. The device according to claim 6 , wherein the determining comprises:

sliding on the audio stream with a preset time step using a window of a preset length, to extract an audio stream clip with each of the preset length;

determining a confidence that each of the audio stream clip comprises the preset feature; and

determining, for the audio data in each preset period, the confidence that the audio data of the preset period comprises the preset feature based on the confidence that each of the audio stream clip to which the audio data of the preset period belongs comprises the preset feature.

8. The device according to claim 6 , wherein the preset confidence threshold is a plurality of preset confidence thresholds, and the extracting the target video clip comprises:

determining the audio clip within a continuous preset period, the confidence of which is above the preset confidence threshold, for each preset confidence threshold of the preset confidence thresholds;

determining the target audio clip in a plurality of audio clips determined based on the preset confidence thresholds; and

extracting the target video clip corresponding to the location of the target audio clip in the video.

9. The device according to claim 6 , wherein the determining the confidence comprises:

using a neural network classification model to determine the confidence that the audio data of each preset period in the audio stream comprises the preset feature.

10. The device according to claim 6 , wherein the preset feature comprises a feature representing that a spectrum change of the audio data in the audio stream exceeds a preset spectrum change threshold.

11. A non-transitory computer readable storage medium, storing computer instructions, the computer instructions, being used to cause the computer to perform an operation for extracting a video clip, comprising:

acquiring a video, and extracting an audio stream in the video;

determining a confidence that audio data in each preset period in the audio stream comprises a preset feature; and

extracting a target video clip corresponding to a location of a target audio clip in the video; wherein the target audio clip is an audio clip within a continuous preset period, and has a confidence that the audio data comprises the preset feature, which is larger than a preset confidence threshold.

12. The medium according to claim 11 , wherein the determining comprises:

sliding on the audio stream with a preset time step using a window of a preset length, to extract an audio stream clip with each of the preset length;

determining a confidence that each of the audio stream clip comprises the preset feature; and

determining, for the audio data in each preset period, the confidence that the audio data of the preset period comprises the preset feature based on the confidence that each of the audio stream clip to which the audio data of the preset period belongs comprises the preset feature.

13. The medium according to claim 11 , wherein the preset confidence threshold is a plurality of preset confidence thresholds, and the extracting the target video clip comprises:

determining the audio clip within a continuous preset period, the confidence of which is above the preset confidence threshold, for each preset confidence threshold of the preset confidence thresholds;

determining the target audio clip in a plurality of audio clips determined based on the preset confidence thresholds; and

extracting the target video clip corresponding to the location of the target audio clip in the video.

14. The medium according to claim 11 , wherein the determining the confidence comprises:

using a neural network classification model to determine the confidence that the audio data of each preset period in the audio stream comprises the preset feature.

15. The medium according to claim 11 , wherein the preset feature comprises a feature representing that a spectrum change of the audio data in the audio stream exceeds a preset spectrum change threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2021
From: ZHANG, QINYI; MA, CAIHONG
To: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
Reel/Frame 056681/0751 →
Priority Claims (1)
CN 202011064001.4 · Sep 30, 2020 · national
Continuity (1)
Related Publication 20210209371A1 · Jul 8, 2021