IP Library Granted Patent US 12,400,428
Granted Patent B2
US 12,400,428 · App. 18/149,035 · Granted Aug 26, 2025

Automatic classification method and system of teaching videos based on different presentation forms

Inventors: Qiusha Min (Wuhan, CN); Ziyi Li (Wuhan, CN)
Assignee: Central China Normal University
G06V10/764G06V10/44G06V10/82G06V20/46G06V40/161
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,428
App. No.
18/149,035
Granted
Aug 26, 2025
Kind
B2
Abstract

The present disclosure belongs to the technical field of artificial intelligence, and discloses an automatic classification method and system of teaching videos based on different presentation forms, which with three convolutional neural network models, may accurately locate the information required for teaching video classification by two self-trained YOLOV4 target detection neural network models and human body key point detection technology, solves the problem that the background and character features of teaching videos do not change significantly, and improves the accuracy of feature extraction. The structure of the self-built convolutional neural network models is suitable for classification of Interview type and Head type teaching videos. The depth of the network is just appropriate compared with several classical video classification algorithms, which reduces the energy consumption of computer hardware. Using other related image data sets preprocessed as the required training sets breaks through the bottleneck in data sets of teaching video.

Claims (25)

1. An automatic classification method of teaching videos based on different presentation forms, comprising:

extracting classroom features from key frames of a video using a self-trained YOLOV4 target detection network model 1, and determining whether the video is a classroom recording type by the outputted classroom features;

determining whether the video is a pure PPT type by the information outputted from a self-trained YOLOV4 target detection network model 2;

distinguishing PPT plus teacher image type videos from studio recording type videos according to human body key point detection; and

distinguishing features of interview type videos from features of head type videos using a self-built convolutional neural network model.

2. The automatic classification method of teaching videos based on different presentation forms according to claim 1 , specifically comprising:

(1) collecting six types of teaching videos, classifying the collected video sets according to six types of teaching videos, and extracting video key frames;

(2) after extracting the video key frames, preprocessing the video key frames to form video folders, each folder being composed of corresponding video key frames as a test set for teaching video detection;

(3) detecting the preprocessed key frames in the folders of videos by two self-trained YOLOV4 target detection network models, and determining whether the video types are pure PPT type, PPT plus teacher image type, classroom recording type and studio recording type through the outputted information;

(4) cutting the face parts of key frames in the interview type and head type videos in a size of 28×28 through face detection technology;

(5) since the interview type and the head type teaching videos have some differences in head pose features, collecting and classifying public data sets of face images and head poses; and

(6) inputting the key frames in the remaining folder into the self-built convolutional neural network model for classification detection.

3. The automatic classification method of teaching videos based on different presentation forms according to claim 2 , wherein the six types of teaching videos in step (1) comprise pure PPT type, PPT plus teacher image type, classroom recording type, studio recording type, interview type and head type.

4. The automatic classification method of teaching videos based on different presentation forms according to claim 2 , wherein the preprocessing the video key frames in step (2) comprises unifying the picture size to 416×416 and removing average values from the images.

5. The automatic classification method of teaching videos based on different presentation forms according to claim 2 , wherein in step (5), the collected public data sets are used as training sets, validation sets and test sets for training and distinguishing the two video types, three types of data are input into the self-built convolutional neural network respectively, and after the optimal weight is obtained, the video folders formed by the key frames extracted from the two types of videos are used as the final test sets for final detection.

6. An automatic classification system of teaching videos based on different presentation forms for the automatic classification method of teaching videos based on different presentation forms according to claim 1 , comprising:

a self-trained YOLOV4 target detection network model 1 unit: a large number of collected images with classroom features are used as data sets, classroom features on each image are marked, and a YOLOV4 target detection network model is trained and optimized using the marked tags, so that the YOLOV4 target detection network model 1 detects classroom features in images or videos;

a self-trained YOLOV4 target detection network model 2 unit: a YOLOV4 target detection network model is trained using a public COCO data set, so that the YOLOV4 target detection network model 2 outputs character information in images or videos; and

a self-built convolutional neural network model unit: Prelu is used as an activation function, an Adabound optimizer is used, a data enhancement layer is added to the first layer, each layer uses Batch Normalization for batch normalization, and the last layer uses a softmax function for classifying key frames of teaching videos.

7. The automatic classification system of teaching videos based on different presentation forms according to claim 6 , wherein the self-built convolutional neural network model unit comprises 5 convolution layers, 1 pooling layer, 1 Dropout layer and 2 fully connected layers connected in turn,

3 convolution layers having a convolution core size of 3×3 and a step size of 1, 2 convolution layers having a convolution core size of 3×3 and a step size of 2, and the convolution layers with the step size of 1 and the convolution layers with the step size of 2 being arranged alternately; and

the pooling layer having a size of 2×2 and a step size of 2, and being connected behind the last convolution layer: the first fully connected layer having a size of 256; the parameter of the Dropout layer being 0.3; the size of the second fully connected layer being the size of the classified video type; and the size of an input layer being 28×28×3.

8. A computer program product stored on a non-transitory computer-readable medium, comprising a computer-readable program, when executed on an electronic device, providing a user input interface to apply the automatic classification method of teaching videos based on different presentation forms according to claim 1 .

9. A non-transitory computer-readable storage medium for storing instructions, wherein when the instructions are run on a computer, the computer applies the automatic classification method of teaching videos based on different presentation forms according to claim 1 .

10. An information data processing terminal, configured to implement the automatic classification method of teaching videos based on different presentation forms according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 30, 2022
From: MIN, QIUSHA; LI, ZIYI
To: CENTRAL CHINA NORMAL UNIVERSITY
Reel/Frame 062248/0223 →
Priority Claims (1)
CN 202210249839.3 · Mar 14, 2022 · national
Continuity (1)
Related Publication 20230290118A1 · Sep 14, 2023
References Cited (8)
US 20230290118A1 · Min · 2023 [cited by examiner]
US 20250181985A1 · Chatterjee · 2025 [cited by examiner]
CN 106851419A · 2017 [cited by applicant]
CN 111666748A · 2020 [cited by applicant]
CN 112861802A · 2021 [cited by applicant]
Wang, Biao, and Xiao Guo. “Analysis of Classroom Teaching Status Based on Target Detection Model.” 2019 IEEE 10th International Conference on Software Engineering and Service Science (ICSESS). IEEE, 2019. (Year: 2019). [cited by examiner]
Köse, Erman, Elif Talbeyaz, and Selçuk Karaman. “Classification of instructional videos.” Technology, Knowledge and Learning 26.4 (2021): 1079-1109. (Year: 2021). [cited by examiner]
Choe, Ronny C., et al. “Student satisfaction and learning outcomes in asynchronous online lecture videos.” CBE—Life Sciences Education 18.4 (2019): ar55. (Year: 2019). [cited by examiner]