IP Library › Granted Patent US 12,513,365
Granted Patent B2
US 12,513,365 · App. 17/772,392 · Granted Dec 30, 2025

Method for training content moderation model, method for moderating video content, computer device, and storage medium

Inventors: Feng Shi (Guangzhou, CN); Zhenqiang Liu (Guangzhou, CN)
Assignee: BIGO TECHNOLOGY PTE. LTD.
H04N21/466G06T3/40G06V10/22G06V10/774G06V10/82G06V20/41G06V20/46G06V20/49H04N21/4542
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,513,365
App. No.
17/772,392
Granted
Dec 30, 2025
Kind
B2
Abstract

Provided is a method for training a content moderation mode. The method includes extracting part of image data of a sample video file as sample image data; positioning a time point of the sample image data in the sample video file in the case that the sample image data contains offensive content; extracting salient image region data from the image data around the time point; and training the content moderation model based on the image region data and the sample image data.

Claims (62)

1 . A method for training a content moderation model, comprising:

extracting part of image data of a sample video file as sample image data;

positioning a time point of the sample image data on a timeline of the sample video file in response to a determination that the sample image data contains offensive content;

extracting salient image region data from the image data around the time point; and

training the content moderation model based on the image region data and the sample image data;

wherein extracting the salient image region data from the image data around the time point comprises:

determining a time range containing the time point;

looking up a salient region detection model configured to identify a salient image region in the image data; and

identifying the salient image region data in the image data by inputting the image data within the time range into the salient region detection model.

2 . The method according to claim 1 , wherein extracting part of image data of the sample video file as the sample image data comprises:

partitioning the sample video file into at least two sample video segments; and

extracting part of the image data from each sample video segment as the sample image data.

3 . The method according to claim 2 , wherein extracting part of image data of the sample video file as the sample image data comprises at least one of:

chronologically ranking the sample image data; and

scaling the sample image data to a preset size.

4 . The method according to claim 1 , wherein positioning the time point of the sample image data on the timeline of the sample video file in response to the determination that the sample image data contains the offensive content comprises:

looking up an offense discrimination model configured to identify an image offense score of the content in the image data;

identifying the image offense score of content in the sample image data by inputting the sample image data into the offense discrimination model;

selecting the sample image data with the image offense score meeting a preset offense condition; and

determining the time point of the sample image data meeting the preset offense condition in the sample video file.

5 . The method according to claim 4 , wherein looking up the offense discrimination model comprises:

determining an offense category marked on the sample video file and representing the offensive content; and

looking up the offense discrimination model corresponding to the offense category, wherein the offense discrimination model is configured to identify the image offense score of the content that belongs to the offense category in the image data.

6 . The method according to claim 4 , wherein selecting the sample image data with the image offense score meeting the preset offense condition comprises:

determining, whether the sample image data comprises the image offense score that is greater than a preset image score threshold;

determining that the image offense score meets the preset offense condition in the case that the sample image data comprises the image offense score that is greater than the preset image score threshold; and

determining that the image offense score of a maximum value meets a preset violation condition in the case that the sample image data does not comprise the image offense score that is greater than the preset image score threshold.

7 . The method according to claim 1 , wherein training the content moderation model based on the image region data and the sample image data comprises:

determining an offense category marked on the sample video file and representing the offensive content;

acquiring a deep neural network and a pre-trained model;

initializing the deep neural network by the pre-trained model;

training, by backpropagation, the deep neural network as the content moderation model based on the image region data, the sample image data, and the offense category.

8 . A method for moderating video content, comprising:

extracting part of image data of a target video file as target image data;

positioning a time point of the target image data on a timeline of the target video file in response to a determination that the target image data contains offensive content;

extracting salient image region data from the image data around the time point; and

moderating content of the target video file by inputting the image region data and the target image data into a preset content moderation model;

wherein extracting the salient image region data from the image data around the time point comprises:

determining a time range containing the time point;

looking up a salient region detection model configured to identify a salient image region in the image data; and

identifying the salient image region data in the image data by inputting the image data within the time range into the salient region detection model.

9 . The method according to claim 8 , wherein inputting the image region data and the target image data into the preset content moderation model to moderate the content of the target video file comprises:

determining a file offense score of the content of the target video file by inputting the image region data and the target image data into the preset content moderation model in the case that the content of the target video file belongs to a preset offense category;

determining a file score threshold;

determining the content of the target video file to be legal in the case that the file offense score is less than or equal to the file score threshold.

10 . The method according to claim 9 , wherein moderating the content of the target video file by inputting the image region data and the target image data into the preset content moderation model further comprises:

distributing the target video file to a designated client in the case that the file offense score is greater than the file score threshold;

determining the content of the target video file to be legal in the case that first moderation information is received from the client; and

determining the content of the target video file to be offensive in the case that second moderation information is received from the client.

11 . The method according to claim 9 , wherein determining the file score threshold comprises:

determining a total quantity of target video files with a previous time period, wherein the file offense score of the target video file has been determined;

generating the file score threshold, such that a ratio of a moderation quantity to the total quantity matches a preset push ratio, wherein the moderation quantity is a quantity of the target video files of which file offense scores are greater than the file score threshold.

12 . A computer device for training content moderation model, comprising:

one or more processors;

a memory configured to store one or more programs;

wherein the one or more processors, when running the one or more programs, is caused to perform the method for training the content moderation model as defined in claim 1 .

13 . A non-volatile computer readable storage medium, storing a computer program, wherein the computer program, when run by a processor of a computer device, causes the computer device to perform the method for training the content moderation model as defined in claim 1 .

14 . A computer device for moderating video content, comprising:

one or more processors;

a memory configured to store one or more programs;

wherein the one or more processors, when running the one or more programs, is caused to perform the method for moderating the video content as defined in claim 8 .

15 . A non-volatile computer readable storage medium, storing a computer program, wherein the computer program, when run by a processor of a computer device, causes the computer device to perform the method for moderating the video content as defined in claim 8 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2022
From: SHI, FENG; LIU, ZHENQIANG
To: BIGO TECHNOLOGY PTE. LTD.
Reel/Frame 059741/0206 →
Priority Claims (1)
CN 201911051711.0 · Oct 31, 2019 · national
Continuity (1)
Related Publication 20220377421A1 · Nov 24, 2022
References Cited (43)
US 10962939B1 · Das · 2021 [cited by examiner]
US 11093698B2 · Liu et al. · 2021 [cited by applicant]
US 11354900B1 · Li · 2022 [cited by examiner]
US 20170289624A1 · Avila et al. · 2017 [cited by applicant]
US 20190012568A1 · Kumar et al. · 2019 [cited by applicant]
US 20200366959A1 · Pau · 2020 [cited by examiner]
CN 101315663A · 2008 [cited by applicant]
CN 103440494A · 2013 [cited by applicant]
CN 103902954A · 2014 [cited by applicant]
CN 105844238A · 2016 [cited by applicant]
CN 105844239A · 2016 [cited by applicant]
CN 105930841A · 2016 [cited by applicant]
CN 106708949A · 2017 [cited by applicant]
CN 106778590A · 2017 [cited by applicant]
CN 107341505A · 2017 [cited by applicant]
CN 108121961A · 2018 [cited by applicant]
CN 108124191A · 2018 [cited by applicant]
CN 108419091A · 2018 [cited by examiner]
CN 109151499A · 2019 [cited by applicant]
CN 109284784A · 2019 [cited by applicant]
CN 109409241A · 2019 [cited by applicant]
CN 109753975A · 2019 [cited by applicant]
CN 109862394A · 2019 [cited by examiner]
CN 110276398A · 2019 [cited by applicant]
CN 110796098A · 2020 [cited by applicant]
RU 2446460C1 · 2012 [cited by applicant]
RU 2530671C1 · 2014 [cited by applicant]
WO 2017162919A1 · 2017 [cited by applicant]
Behrad, Alierza et al., “Content-based obscene video recognition by combining 3D spatiotemporal and motion-based features”, Eurasip Journal on Image and Video Processing, vol. 2012, No. 1, abstract, figure 1, sec. 3.2, … [cited by examiner]
Singh, Shubham et al., “KidsGUARD: Fine Grained Approach for Child Unsafe Video Representation and Detection” , Applied Computing, ACM, 2, pp. 2104-2111, abstract, sec. 3.1, 3.2; Apr. 8, 2019 (Year: 2019). [cited by examiner]
Behrad et al.: Content-based obscene video recognition by combining 3D spatiotemporal and motion-based features. Eurasip Journal on Image and Video Processing, 2012 (Year: 2012). [cited by examiner]
Extended European Search Report Communication Pursuant to Rule 62 EPC, dated Oct. 24, 2022 in Patent Application No. EP 20881536.5, which is a foreign counterpart to this U.S. Application. [cited by applicant]
Behrad, Alierza et al., “Content-based obscene video recognition by combining 3D spatiotemporal and motion-based features”, Eurasip Journal on Image and Video Processing, vol. 2012, No. 1, abstract, figure 1, sec. 3.2, … [cited by applicant]
Choi, Byeongcheol et al., “Design and performance evaluation of temporal motion and color energy features for objectionable video classification”, Advanced Information Management and Service (Ims), 2010 6th Internationa… [cited by applicant]
Moreira, Daniel et al., “Multimodal data fusion for sensitive scene localization”, Information Fusion, Elsevier, US, vol. 45, pp. 307-323, the whole document; Mar. 7, 2018. [cited by applicant]
Singh, Shubham et al., “KidsGUARD: Fine Grained Approach for Child Unsafe Video Representation and Detection”, Applied Computing, ACM, 2 Penn Plaza, Suite 701New YorkNY10121-0701USA, pp. 2104-2111, abstract, sec. 3.1, 3… [cited by applicant]
Yan, Jianqiang et al., “Pornographic video detection with MapReduce”, International Journal of Machine Learning and Cybernetics, Springer Berlin Heidelberg, Berlin/ Heidelberg, vol. 9, No. 12, pp. 2105-2115, abstract, f… [cited by applicant]
Yatawatte, Hasini et al., “Content Based Video Retrieval for Obscene Adult Content Detection”, , SAT 2015 18th International Conference, Austin, TX, USA, Sep. 24-27, 2015; [Lecture Notes in Computer Science; Lect. Notes… [cited by applicant]
Communication pursuant to Article 94(3) EPC of European application No. 20881536.5 issued on Nov. 21, 2022. [cited by applicant]
Russian Search Report, dated Nov. 15, 2022 in Patent Application No. 2022114373, which is a foreign counterpart to this U.S. Application. [cited by applicant]
International Search Report of the International Searching Authority for State Intellectual Property Office of the People's Republic of China in PCT application No. PCT/CN2020/107353 issued on Oct. 30, 2020, which is an… [cited by applicant]
The State Intellectual Property Office of People's Republic of China, First Office Action in Patent Application No. CN201911051711.0 issued on Mar. 29, 2021, which is a foreign counterpart application corresponding to t… [cited by applicant]
Notification of Completion of Formalities for Patent Register and Notification to Grant Patent Right for Invention Application No. 201911051711.0 Issued on Jun. 30, 2021, to which this application claims priority. [cited by applicant]