IP Library › Granted Patent US 12,347,195
Granted Patent B2
US 12,347,195 · App. 17/688,987 · Granted Jul 1, 2025

Video text processing method, apparatus, and computer-readable storage medium

Inventors: Hao Song (Shenzhen, CN); Shan Huang (Shenzhen, CN)
Assignee: Tencent Technology (Shenzhen) Company Limited
G06V20/46G06V10/24G06V10/25G06V10/761
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,195
App. No.
17/688,987
Granted
Jul 1, 2025
Kind
B2
Abstract

A video processing method is provided. The method includes extracting at least two adjacent video frame images from a frame image sequence corresponding to a video, positioning a text region of each video frame image in the at least two adjacent video frame images, determining a degree of similarity between text regions of each video frame image in the at least two adjacent video frame images, determining, based on the degree of similarity, a key video frame segment comprising a same text in the video, and determining a text key frame in the video based on the key video frame segment.

Claims (86)

1. A video processing method, performed by at least one processor, the method comprising:

extracting at least two adjacent video frame images from a frame image sequence corresponding to a video;

positioning a text region of each video frame image in the at least two adjacent video frame images, wherein the positioning the text region of each video frame image in the at least two adjacent video frame images comprises positioning, by using a preset key frame model, the text region of each video frame image in the at least two adjacent video frame images, wherein the preset key frame model is obtained by:

obtaining a training sample, the training sample comprising an adjacent sample video frame image, a text annotation region, and a degree of annotation similarity;

obtaining a text prediction region of each sample video frame image in the at least two adjacent video frame images and a degree of prediction similarity between text prediction regions of each sample video frame image in the at least two adjacent video frame images based on an original key frame model;

obtaining a loss function value based on:

a first difference between the text prediction region of each sample video frame image in the at least two adjacent video frame images and the text annotation region, and

a second difference between the degree of prediction similarity and the degree of annotation similarity; and

obtaining the preset key frame model by continuously performing iterative training on the original key frame model based on the loss function value, until a preset training cut-off condition is met;

determining a degree of similarity between text regions of each video frame image in the at least two adjacent video frame images;

determining, based on the degree of similarity, a key video frame segment comprising a same text in the video;

determining a text key frame in the video based on the key video frame segment; and

after the determining the text key frame in the video based on the key video frame segment, transmitting the text key frame to a display device, to display video information corresponding to the text key frame by using the display device.

2. The method of claim 1 , wherein the extracting the at least two adjacent video frame images comprises:

obtaining the frame image sequence by decoding the video; and

obtaining the at least two adjacent video frame images by obtaining a current video frame image and a subsequent video frame image in the frame image sequence.

3. The method of claim 1 , wherein, after the determining the text key frame in the video based on the key video frame segment, the method further comprises:

obtaining target text information by obtaining text information of the text key frame; and

obtaining an audit result by auditing the video based on the target text information.

4. The method of claim 1 , wherein, after the continuously performing iterative training on the original key frame model by using the loss function value, the method further comprises:

optimizing, based on a new training sample being obtained, the preset key frame model using the new training sample;

wherein the positioning the text region of each video frame image in the at least two adjacent video frame images further comprises positioning, by using the optimized preset key frame model, the text region of each video frame image in the at least two adjacent video frame images; and

wherein the determining the degree of similarity between the text regions of each video frame image in the at least two adjacent video frame images by using the preset key frame model further comprises determining the degree of similarity between the text regions of each video frame image in the at least two adjacent video frame images.

5. The method of claim 1 , further comprising:

extracting another at least two adjacent video frame images from the frame image sequence corresponding to the video;

calculating a text inclusion value of each video frame image in the another at least two adjacent video frame images, the text inclusion value indicating whether the corresponding video frame image includes text;

determining that the text inclusion value of each video frame image in the another at least two adjacent video frame images each does not meet a preset inclusion value; and

in response to determining that the text inclusion value of each video frame image in the another at least two adjacent video frame images each does not meet the preset inclusion value, discarding the another at least two adjacent video frame images from key video frame segment processing.

6. The method of claim 1 , wherein the positioning the text region of each video frame image in the at least two adjacent video frame images further comprises:

obtaining an initial feature of each video frame image in the at least two adjacent video frame images;

obtaining a text mask feature of the initial feature;

calculating a text inclusion value of each video frame image in the at least two adjacent video frame images based on the text mask feature; and

determining the text region of each video frame image in the at least two adjacent video frame images based on all text inclusion values corresponding to the at least two adjacent video frame images are greater than a preset inclusion value.

7. The method of claim 6 , wherein the determining the degree of similarity comprises:

fusing the initial feature of each video frame image in the at least two adjacent video frame images and the text mask feature corresponding to the text region of each video frame image in the at least two adjacent video frame images into a key frame feature of each video frame image in the at least two adjacent video frame images;

obtaining a feature difference between key frame features of each video frame image in the at least two adjacent video frame images; and

determining the degree of similarity between the text regions of each video frame image in the at least two adjacent video frame images based on the feature difference.

8. The method of claim 6 , wherein the obtaining the text mask feature of the initial feature comprises:

determining a text weight value of the initial feature; and

obtaining the text mask feature of the initial feature based on the text weight value.

9. An apparatus, comprising:

at least one memory configured to store computer program code; and

at least one processor configured to access said computer program code and operate as instructed by said computer program code, thereby causing the apparatus to:

extract at least two adjacent video frame images from a frame image sequence corresponding to a video;

position a text region of each video frame image in the at least two adjacent video frame images, wherein the positioning the text region of each video frame image in the at least two adjacent video frame images comprises positioning, by using a preset key frame model, the text region of each video frame image in the at least two adjacent video frame images, wherein the preset key frame model is obtained by:

obtaining a training sample, the training sample comprising an adjacent sample video frame image, a text annotation region, and a degree of annotation similarity;

obtaining a text prediction region of each sample video frame image in the at least two adjacent video frame images and a degree of prediction similarity between text prediction regions of each sample video frame image in the at least two adjacent video frame images based on an original key frame model;

obtaining a loss function value based on:

a first difference between the text prediction region of each sample video frame image in the at least two adjacent video frame images and the text annotation region, and

a second difference between the degree of prediction similarity and the degree of annotation similarity; and

obtaining the preset key frame model by continuously performing iterative training on the original key frame model based on the loss function value, until a preset training cut-off condition is met;

determine a degree of similarity between text regions of each video frame image in the at least two adjacent video frame images;

determine, based on the degree of similarity, a key video frame segment comprising a same text in the video;

determine a text key frame in the video based on the key video frame segment; and

after the determining the text key frame in the video based on the key video frame segment, transmit the text key frame to a display device, to display video information corresponding to the text key frame by using the display device.

10. The apparatus of claim 9 , wherein the apparatus is further caused to:

obtain the frame image sequence by decoding the video; and

obtain the at least two adjacent video frame images by obtaining a current video frame image and a subsequent video frame image in the frame image sequence.

11. The apparatus of claim 9 , wherein the apparatus is further caused to:

obtain an initial feature of each video frame image in the at least two adjacent video frame images;

obtain a text mask feature of the initial feature;

calculate a text inclusion value of each video frame image in the at least two adjacent video frame images based on the text mask feature; and

determine the text region of each video frame image in the at least two adjacent video frame images based on all text inclusion values corresponding to the at least two adjacent video frame images are greater than a preset inclusion value.

12. The apparatus of claim 11 , wherein the apparatus is further caused to:

fuse the initial feature of each video frame image in the at least two adjacent video frame images and the text mask feature corresponding to the text region of each video frame image in the at least two adjacent video frame images into a key frame feature of each video frame image in the at least two adjacent video frame images;

obtain a feature difference between key frame features of each video frame image in the at least two adjacent video frame images; and

determine the degree of similarity between the text regions of each video frame image in the at least two adjacent video frame images based on the feature difference.

13. The apparatus of claim 11 , wherein the obtaining the text mask feature of the initial feature comprises:

determining a text weight value of the initial feature; and

obtaining the text mask feature of the initial feature based on the text weight value.

14. The apparatus of claim 9 , wherein the apparatus is further caused to:

obtain target text information by obtaining text information of the text key frame; and

obtain an audit result by auditing the video based on the target text information.

15. A non-transitory computer-readable storage medium storing computer instructions that, when executed by at least one processor, cause a video processing computing device to:

extract at least two adjacent video frame images from a frame image sequence corresponding to a video;

position a text region of each video frame image in the at least two adjacent video frame images, wherein the positioning the text region of each video frame image in the at least two adjacent video frame images comprises positioning, by using a preset key frame model, the text region of each video frame image in the at least two adjacent video frame images, wherein the preset key frame model is obtained by:

obtaining a training sample, the training sample comprising an adjacent sample video frame image, a text annotation region, and a degree of annotation similarity;

obtaining a text prediction region of each sample video frame image in the at least two adjacent video frame images and a degree of prediction similarity between text prediction regions of each sample video frame image in the at least two adjacent video frame images based on an original key frame model;

obtaining a loss function value based on:

a first difference between the text prediction region of each sample video frame image in the at least two adjacent video frame images and the text annotation region, and

a second difference between the degree of prediction similarity and the degree of annotation similarity; and

obtaining the preset key frame model by continuously performing iterative training on the original key frame model based on the loss function value, until a preset training cut-off condition is met;

determine a degree of similarity between text regions of each video frame image in the at least two adjacent video frame images;

determine, based on the degree of similarity, a key video frame segment comprising a same text in the video;

determine a text key frame in the video based on the key video frame segment; and

after the determining the text key frame in the video based on the key video frame segment, transmit the text key frame to a display device, to display video information corresponding to the text key frame by using the display device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2022
From: SONG, HAO; HUANG, SHAN
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 059192/0347 →
Priority Claims (1)
CN 202010096614.X · Feb 17, 2020 · national
Continuity (2)
Continuation PCTCN2020126832 · Nov 5, 2020
Related Publication 20220198800A1 · Jun 23, 2022
References Cited (18)
US 6243419B1 · Satou · 2001 [cited by examiner]
US 7787705B2 · Sun · 2010 [cited by examiner]
US 11138440B1 · Wang · 2021 [cited by examiner]
US 20110110592A1 · Wada · 2011 [cited by examiner]
US 20180082127A1 · Carlson et al. · 2018 [cited by applicant]
US 20180210774A1 · Young · 2018 [cited by examiner]
US 20190034312A1 · Karunamoorthy et al. · 2019 [cited by applicant]
CN 106454151A · 2017 [cited by applicant]
CN 108495162A · 2018 [cited by applicant]
CN 109803180A · 2019 [cited by applicant]
CN 110147745A · 2019 [cited by applicant]
CN 110399798A · 2019 [cited by applicant]
CN 111294646A · 2020 [cited by applicant]
First Office Action of Chinese Application No. 202010096614.X dated Jan. 27, 2021. [cited by applicant]
Second Office Action of Chinese Application No. 202010096614.X dated Aug. 2, 2021. [cited by applicant]
Third Office Action of Chinese Application No. 202010096614.X dated Oct. 26, 2021. [cited by applicant]
International Search Report of PCT/CN2020/126832 dated Feb. 4, 2021 [PCT/ISA/210]. [cited by applicant]
Written Opinion of PCT/CN2020/126832 dated Feb. 4, 2021 [PCT/ISA/237]. [cited by applicant]