IP Library Granted Patent US 12713103
Granted Patent B2
US 12713103 · App. 18/320,302 · Granted Aug 18, 2026

Subtitle processing method and apparatus of multimedia file, electronic device, and computer-readable storage medium

Inventors: Dan He (Shenzhen, CN); Shuyu Gong (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
H04N21/4884H04N21/44008
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12713103
App. No.
18/320,302
Granted
Aug 18, 2026
Kind
B2
Abstract

A subtitle processing method includes: playing the multimedia file in response to a play trigger operation, the multimedia file being associated with a plurality of subtitles, a type of the multimedia file being a video file or an audio file, and displaying the plurality of subtitles sequentially in a human-computer interaction interface during playing the multimedia file, a pattern of the plurality of subtitles being related to a content of the multimedia file.

Claims (84)

1 . A subtitle processing method of a multimedia file, the method being performed by an electronic device and comprising:

playing the multimedia file in response to a play trigger operation, the multimedia file being associated with a plurality of subtitles, a type of the multimedia file being a video file or an audio file, wherein the multimedia file comprises a plurality of segments, one of the segments comprises a plurality of sub-segments, each sub-segment has a content feature of a dynamic dimension of the sub-segment, and different sub-segments have different content features of the dynamic dimension; and

displaying the plurality of subtitles sequentially in a human-computer interaction interface during playing the multimedia file, a pattern of the plurality of subtitles being related to a content of the multimedia file, comprising:

for one of the segments of the multimedia file:

displaying at least one subtitle associated with the segment sequentially in the human-computer interaction interface based on a pattern adapted to a content feature of at least one dimension of the segment, comprising:

for one sub-segment of the segment: displaying at least one subtitle associated with the sub-segment based on a pattern adapted to the content feature of the dynamic dimension that the sub-segment has.

2 . The method according to claim 1 , wherein the displaying the plurality of subtitles sequentially in a human-computer interaction interface, comprises:

displaying the plurality of subtitles to which the pattern is applied sequentially in the human-computer interaction interface, the pattern being adapted to a content feature of at least one dimension of the multimedia file, and the content feature of the at least one dimension comprising: a style, an object, a scenario, a plot, and a hue.

3 . The method according to claim 2 , wherein the method further comprises:

acquiring the content feature of the at least one dimension of the multimedia file; and

performing a pattern conversion on a plurality of original subtitles associated with the multimedia file based on the content feature of the at least one dimension to obtain a plurality of new subtitles, the plurality of new subtitles being used as the plurality of subtitles to be displayed in the human-computer interaction interface.

4 . The method according to claim 3 , wherein the performing a pattern conversion on a plurality of original subtitles associated with the multimedia file based on the content feature of the at least one dimension to obtain a plurality of new subtitles, comprises:

calling a subtitle model based on a value corresponding to the content feature of the at least one dimension and the plurality of original subtitles associated with the multimedia file to obtain a plurality of new subtitles,

the subtitle model being a generative model, wherein the generative model is trained in a generative adversarial network formed by the generative model and a discriminative model.

5 . The method according to claim 1 , wherein

types of the segments comprise at least one of the following: an object segment, a scenario segment, and a plot segment.

6 . The method according to claim 5 , wherein the displaying at least one subtitle associated with the segment sequentially in the human-computer interaction interface based on a pattern adapted to a content feature of at least one dimension of the segment, comprises:

acquiring a content feature of a static dimension of the segment, a content feature of a static dimension of the object segment comprising at least one of the following object properties of a sounding object in the object segment: a role type, a gender, and an age; a feature of a static dimension of the scenario segment comprising a scenario type of the scenario segment; a feature of a static dimension of the plot segment comprising a plot progress of the plot segment; and

displaying at least one subtitle associated with the segment synchronously in the human-computer interaction interface based on a pattern adapted to the content feature of the static dimension of the segment, the pattern keeping unchanged during playing the segment.

7 . The method according to claim 5 , wherein

the plurality of sub-segments have a content feature of a static dimension of the segment and the content feature of the dynamic dimension of the segment; and

the displaying at least one subtitle associated with the segment sequentially in the human-computer interaction interface based on a pattern adapted to a content feature of at least one dimension of the segment, comprises: for one sub-segment of the segment:

displaying at least one subtitle associated with the sub-segment based on the pattern adapted to the content feature of the static dimension and the content feature of the dynamic dimension that the sub-segment has.

8 . The method according to claim 7 , wherein

the content feature of the static dimension of the object segment comprises at least one of the following object properties: a role type, a gender, and an age of the sounding object in the object segment; the content feature of the dynamic dimension of the object segment comprises the following object properties: a mood of the sounding object in the object segment;

a content feature of the static dimension of the plot segment comprises a plot type of the plot segment, and a content feature of a dynamic dimension of the plot segment comprises at least one of the following: scenario types of different scenarios appearing in the scenario segment, and object properties of different sounding objects appearing in the scenario segment; and

a content feature of the static dimension of the scenario segment comprises: a type of a scenario to which the scenario segment relates; a content feature of a dynamic dimension of the scenario segment comprises at least one of the following: object properties of different sounding objects appearing in the scenario segment, and types of different plots appearing in the scenario segment.

9 . The method according to claim 1 , wherein

when the at least one dimension is a plurality of dimensions, the displaying at least one subtitle associated with the segment sequentially in the human-computer interaction interface based on a pattern adapted to a content feature of at least one dimension of the segment, comprises:

fusing content features of the plurality of dimensions of the segment to obtain a fusion content feature; and

performing a pattern conversion on at least one original subtitle associated with the segment based on the fusion content feature to obtain at least one new subtitle, the at least one new subtitle being used as the at least one subtitle to be displayed in the human-computer interaction interface.

10 . The method according to claim 3 , wherein the acquiring the content feature of the at least one dimension of the multimedia file, comprises:

calling a content feature identification model to perform a content feature identification on the content of the multimedia file to obtain a content feature of at least one dimension of the multimedia file;

the content feature identification model being obtained by training based on a sample multimedia file and a label labeled for a content of the sample multimedia file.

11 . The method according to claim 3 , wherein

when the multimedia file is the video file, the acquiring the content feature of the at least one dimension of the multimedia file, comprises:

performing the following processing for a target object appearing in the video file:

preprocessing a target video frame where the target object is located;

performing a feature extraction on the target video frame being preprocessed to obtain an image feature corresponding to the target video frame; and

performing a dimension reduction on the image feature, and classifying the image features after the dimension reduction through a trained classifier to obtain an object property of the target object.

12 . The method according to claim 3 , wherein

when the multimedia file is the video file, the acquiring the content feature of the at least one dimension of the multimedia file, comprises:

performing the following processing for a target object appearing in the video file:

extracting a local binary pattern feature corresponding to the target video frame where the target object is located, and performing a dimension reduction on the local binary pattern feature;

extracting a histogram of oriented gradient feature corresponding to the target video frame, and performing a dimension reduction on the histogram of oriented gradient feature;

performing a canonical correlation analysis on the local binary pattern feature and the histogram of oriented gradient image after the dimension reduction to obtain an analysis result; and

performing a regression on the analysis result to obtain the object property of the target object.

13 . The method according to claim 3 , wherein

when the multimedia file is the video file, the acquiring the content feature of the at least one dimension of the multimedia file, comprises:

performing the following processing for a target object appearing in the video file:

normalizing the target video frame where the target object is located, and partitioning the target video frame being normalized to obtain a plurality of sub-regions;

extracting local binary pattern features corresponding to the sub-regions, and performing statistics on a plurality of the local binary pattern features to obtain a local histogram statistic feature corresponding to the target video frame; and

performing a local sparse reconstruction representation on the local histogram statistic feature through a local feature library of a training set, and performing a local reconstruction residual weighting identification on a local sparse reconstruction representation result to obtain the object property of the target object.

14 . The method according to claim 11 , wherein

when a plurality of objects appear in the video file, determining the target object from the plurality of objects in one of the following ways:

determining an object having a longest appearance time in the video file as the target object;

determining an object satisfying a user preference in the video file as the target object; and

determining an object related to a user interaction in the video file as the target object.

15 . A subtitle processing apparatus of a multimedia file, comprising:

at least one memory, configured to store an executable instruction; and

at least one processor, configured to execute the executable instruction stored in the at least one memory and perform:

playing the multimedia file in response to a play trigger operation, the multimedia file being associated with a plurality of subtitles, a type of the multimedia file being a video file or an audio file, wherein the multimedia file comprises a plurality of segments, one of the segments comprises a plurality of sub-segments, each sub-segment has a content feature of a dynamic dimension of the sub-segment, and different sub-segments have different content features of the dynamic dimension; and

displaying the plurality of subtitles sequentially in a human-computer interaction interface during playing the multimedia file, a pattern of the plurality of subtitles being related to a content of the multimedia file, comprising:

for one of the segments of the multimedia file:

displaying at least one subtitle associated with the segment sequentially in the human-computer interaction interface based on a pattern adapted to a content feature of at least one dimension of the segment, comprising:

for one sub-segment of the segment: displaying at least one subtitle associated with the sub-segment based on a pattern adapted to the content feature of the dynamic dimension that the sub-segment has.

16 . The apparatus according to claim 15 , wherein the displaying the plurality of subtitles sequentially in a human-computer interaction interface, comprises:

displaying the plurality of subtitles to which the pattern is applied sequentially in the human-computer interaction interface, the pattern being adapted to a content feature of at least one dimension of the multimedia file, and the content feature of the at least one dimension comprising: a style, an object, a scenario, a plot, and a hue.

17 . The apparatus according to claim 16 , wherein the at least one processor is further configured to perform:

acquiring the content feature of the at least one dimension of the multimedia file; and

performing a pattern conversion on a plurality of original subtitles associated with the multimedia file based on the content feature of the at least one dimension to obtain a plurality of new subtitles, comprising:

calling a subtitle model based on a value corresponding to the content feature of the at least one dimension and the plurality of original subtitles associated with the multimedia file to obtain a plurality of new subtitles,

the subtitle model being a generative model, wherein the generative model is trained in a generative adversarial network formed by the generative model and a discriminative model.

18 . The apparatus according to claim 15 , wherein

types of the segment comprise at least one of the following: an object segment, a scenario segment, and a plot segment.

19 . The apparatus according to claim 18 , wherein the displaying at least one subtitle associated with the segment sequentially in the human-computer interaction interface based on a pattern adapted to a content feature of at least one dimension of the segment, comprises:

acquiring a content feature of a static dimension of the segment, a content feature of a static dimension of the object segment comprising at least one of the following object properties of a sounding object in the object segment: a role type, a gender, and an age; a feature of a static dimension of the scenario segment comprising a scenario type of the scenario segment; a feature of a static dimension of the plot segment comprising a plot progress of the plot segment; and

displaying at least one subtitle associated with the segment synchronously in the human-computer interaction interface based on a pattern adapted to the content feature of the static dimension of the segment, the pattern keeping unchanged during playing the segment.

20 . A non-transitory computer-readable storage medium storing an executable instruction, the executable instruction when executed by at least one processor, cause the at least one processor to perform:

playing the multimedia file in response to a play trigger operation, the multimedia file being associated with a plurality of subtitles, a type of the multimedia file being a video file or an audio file, wherein the multimedia file comprises a plurality of segments, one of the segments comprises a plurality of sub-segments, each sub-segment has a content feature of a dynamic dimension of the sub-segment, and different sub-segments have different content features of the dynamic dimension; and

displaying the plurality of subtitles sequentially in a human-computer interaction interface during playing the multimedia file, a pattern of the plurality of subtitles being related to a content of the multimedia file, comprising:

for one of the segments of the multimedia file:

displaying at least one subtitle associated with the segment sequentially in the human-computer interaction interface based on a pattern adapted to a content feature of at least one dimension of the segment, comprising:

for one sub-segment of the segment: displaying at least one subtitle associated with the sub-segment based on a pattern adapted to the content feature of the dynamic dimension that the sub-segment has.