IP Library Granted Patent US 11,272,257
Granted Patent B2
US 11,272,257 · App. 16/936,514 · Granted Mar 8, 2022

Method and apparatus for pushing subtitle data, subtitle display method and apparatus, device and medium

Inventors: Zi Heng Luo (Shenzhen, CN); Xiu Ming Zhu (Shenzhen, CN); Xiao Hua Hu (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LTD
H04N21/4884G10L15/005G10L15/26H04N21/2187H04N21/4307H04N21/4856
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,272,257
App. No.
16/936,514
Granted
Mar 8, 2022
Kind
B2
Abstract

A method and apparatus for pushing subtitle data in a live scenario. The method includes: obtaining video stream data and audio stream data, the audio stream data being data corresponding to an audio part in the video stream data; generating the subtitle data according to the audio stream data, the subtitle data comprising a subtitle text corresponding to a speech in the audio stream data and time information of the subtitle text; and pushing, in response to pushing the video stream data to a user terminal, the subtitle data to the user terminal, the subtitle data instructing the user terminal to synchronously display the subtitle text with live pictures in the video stream data and the audio part in the audio stream data according to the time information of the subtitle text.

Claims (85)

1. A method for pushing subtitle data, performed by a computer device, the method comprising:

obtaining video stream data and audio stream data, the audio stream data being data corresponding to an audio part in the video stream data;

generating the subtitle data according to the audio stream data, the subtitle data comprising a subtitle text corresponding to a speech in the audio stream data and time information of the subtitle text;

pushing, in response to pushing the video stream data to a user terminal, the subtitle data to the user terminal, the subtitle data instructing the user terminal to synchronously display the subtitle text with live pictures in the video stream data and the audio part in the audio stream data according to the time information of the subtitle text; and

receiving a subtitle obtaining request from the user terminal, the subtitle obtaining request comprises a time identifier, the time identifier indicating time information of the requested subtitle data,

wherein the generating the subtitle data according to the audio stream data comprises generating the subtitle data according to the audio stream data through a target service, the target service including a subtitle generation service,

wherein the subtitle obtaining request further includes a service identifier used for indicating a subtitle generation service,

wherein the pushing the subtitle data to the user terminal further comprises pushing the subtitle data to the user terminal based on determining that the subtitle generation service indicated by the service identifier is the target service, and

wherein the method further comprises:

detecting whether sequence numbers of data blocks in the subtitle data are consecutive;

based on determining that the sequence numbers of the data blocks in the subtitle data are not consecutive, requesting data blocks corresponding to missing sequence numbers from the target service, the missing sequence numbers being sequence numbers that are missing between a sequence number of the first data block and a sequence number of the last data block in the subtitle data, receiving the data blocks corresponding to the missing sequence numbers from the target service; and

rearranging the subtitle data based on the received data blocks corresponding to the missing sequence numbers wherein the pushing the subtitle data to the user terminal further comprises:

querying whether the subtitle data corresponding to the time information indicated by the time identifier is cached;

based on determining that the subtitle data corresponding to the time information is cached, pushing the cached subtitle data to the user terminal; and

delaying the pushing of the video stream data to the user terminal to synchronize the subtitle text with the video stream data.

2. The method according to claim 1 , wherein the subtitle obtaining request further comprises language indication information indicating a subtitle language of the subtitle data, and

wherein the pushing, in response to the pushing the video stream data to the user terminal, the subtitle data to the user terminal comprises:

determining whether the subtitle language indicated by the language indication information is a language corresponding to the subtitle text; and

based on determining that the language indication information is the language corresponding to the subtitle text, pushing the subtitle data to the user terminal.

3. The method according to claim 1 , further comprising:

based on determining that the subtitle data is not found, extracting the subtitle data from a subtitle database; and

caching the extracted subtitle data.

4. The method according to claim 2 , further comprising:

determining a next request time according to the time information of the subtitle data pushed to the user terminal; and

transmitting request indication information to the user terminal, the request indication information instructing the user terminal to transmit a new subtitle obtaining request when the next request time arrives.

5. The method according to claim 1 , wherein the generating the subtitle data according to the audio stream data comprises:

performing a speech recognition on the audio stream data to obtain a speech recognized text; and

generating the subtitle data according to the speech recognized text.

6. The method according to claim 5 , wherein the performing speech recognition on the audio stream data to obtain the speech recognized text comprises:

performing a speech start and end detection on the audio stream data to obtain a speech start frame and a speech end frame in the audio stream data, the speech start frame being an audio frame at the start of a speech segment and the speech end frame being an audio frame at the end of the speech segment; and

performing the speech recognition on target speech data in the audio stream data to obtain the speech recognized text corresponding to the target speech data, the target speech data comprising a plurality of audio frames between any set of the speech start frame and the speech end frame in the audio stream data.

7. The method according to claim 6 , wherein the performing the speech recognition on the target speech data in the audio stream data further comprises:

performing a speech frame extraction at predetermined time intervals according to the time information of the plurality of audio frames in the target speech data to obtain at least one piece of speech subdata, the speech subdata comprising at least one audio frame, among the plurality of audio frames, between the speech start frame and a target audio frame in the target speech data when the speech frame extraction operation of the speech subdata corresponds to the time information in the target speech data;

performing the speech recognition on the at least one piece of speech subdata to obtain recognized subtext corresponding to the at least one piece of speech subdata; and

obtaining the recognized subtext corresponding to the at least one piece of speech subdata as the speech recognized text corresponding to the target speech data.

8. The method according to claim 5 , wherein the generating the subtitle data according to the speech recognized text comprises:

translating the speech recognized text into translated text corresponding to a target language;

generating the subtitle text according to the translated text, the subtitle text comprising at least one of the translated text or the speech recognized text; and

generating the subtitle data according to the subtitle text.

9. The method according to claim 1 , wherein the obtaining the video stream data and the audio stream data comprises:

transcoding a video stream through a transcoding process in a transcoding device to obtain the video stream data and the audio stream data with synchronized time information.

10. The method according to claim 1 , wherein the video stream data is live video stream data.

11. An apparatus for pushing subtitle data, comprising:

at least one memory storing compute program code; and

at least one processor configured to access the at least one memory and operate as instructed by the computer program code, the computer program code comprising:

stream obtaining code configured to cause the at least one processor to obtain video stream data and audio stream data, the audio stream data being data corresponding to an audio part in the video stream data;

subtitle data generation code configured to cause the at least one processor to generate the subtitle data according to the audio stream data, the subtitle data comprising a subtitle text corresponding to a speech in the audio stream data and time information of the subtitle text; and

subtitle pushing code configured to cause the at least one processor to:

push, in response to pushing the video stream data to a user terminal, the subtitle data to the user terminal, the subtitle data instructing the user terminal to synchronously display the subtitle text with live pictures in the video stream data and the audio part in the audio stream data according to the time information of the subtitle text;

receive a subtitle obtaining request from the user terminal, the subtitle obtaining request comprising a time identifier, the time identifier indicating time information of the requested subtitle data;

query whether the subtitle data corresponding to the time information indicated by the time identifier is cached;

based on determining that the subtitle data corresponding to the time information is cached, push the cached subtitle data to the user terminal; and

delay the pushing of the video stream data to the user terminal to synchronize the subtitle text with the video stream data,

wherein the subtitle data generation code is further configured to cause the at least one processor to generate the subtitle data according to the audio stream data through a target service, the target service including a subtitle generation service,

wherein the subtitle obtaining request further includes a service identifier used for indicating a subtitle generation service,

wherein the subtitle pushing code is further configured to cause the at least one processor to push the subtitle data to the user terminal based on determining that the subtitle generation service indicated by the service identifier is the target service, and wherein the apparatus further comprises sequence number detection code configured to cause the at least one processor to:

detect whether sequence numbers of data blocks in the subtitle data are consecutive, based on determining that the sequence numbers of the data blocks in the subtitle data are not consecutive, request data blocks corresponding to missing sequence numbers from the target service, the missing sequence numbers being sequence numbers that are missing between a sequence number of the first data block and a sequence number of the last data block in the subtitle data;

receive the data blocks corresponding to the missing sequence numbers from the target service; and

rearrange the subtitle data based on the received data blocks corresponding to the missing sequence numbers.

12. The apparatus according to claim 11 , wherein the subtitle obtaining request further comprises language indication information indicating a subtitle language of the subtitle data, and

wherein the subtitle pushing code is further configured to cause the at least one processor to:

determine whether the subtitle language indicated by the language indication information is a language corresponding to the subtitle text; and

based on determining that the language indication information is the language corresponding to the subtitle text, push the subtitle data to the user terminal.

13. The apparatus according to claim 12 , further comprising:

time determining code configured to cause the at least one processor to determine a next request time according to the time information of the subtitle data pushed to the user terminal; and

indication information transmitting code configured to cause the at least one processor to transmit request indication information to the user terminal, the request indication information instructing the user terminal to transmit a new subtitle obtaining request when the next request time arrives.

14. The apparatus according to claim 11 , wherein the subtitle data generation code is further configured to cause the at least one processor to:

perform a speech recognition on the audio stream data to obtain a speech recognized text; and

generate the subtitle data according to the speech recognized text.

15. A non-transitory computer-readable storage medium, storing at least one instruction, the at least one instruction, when loaded and executed by a processor, the processor is configured to:

obtain video stream data and audio stream data, the audio stream data being data corresponding to an audio part in the video stream data;

generate the subtitle data according to the audio stream data, the subtitle data comprising a subtitle text corresponding to a speech in the audio stream data and time information of the subtitle text; and

push, in response to pushing the video stream data to a user terminal, the subtitle data to the user terminal, the subtitle data instructing the user terminal to synchronously display the subtitle text with live pictures in the video stream data and the audio part in the audio stream data according to the time information of the subtitle text; and

receive a subtitle obtaining request from the user terminal, the subtitle obtaining request comprising a time identifier, the time identifier indicating time information of the requested subtitle data,

wherein the subtitle obtaining request further includes a service identifier used for indicating a subtitle generation service, and

wherein the processor is further configured to:

generate the subtitle data according to the audio stream data through a target service, the target service including a subtitle generation service,

push the subtitle data to the user terminal based on determining that the subtitle generation service indicated by the service identifier is the target service, and

detect whether sequence numbers of data blocks in the subtitle data are consecutive;

based on determining that the sequence numbers of the data blocks in the subtitle data are not consecutive, request data blocks corresponding to missing sequence numbers from the target service, the missing sequence numbers being sequence numbers that are missing between a sequence number of the first data block and a sequence number of the last data block in the subtitle data:

receive the data blocks corresponding to the missing sequence numbers from the target service; and

rearrange the subtitle data based on the received data blocks corresponding to the missing sequence numbers receive a subtitle obtaining request from the user terminal, the subtitle obtaining request comprising a time identifier, the time identifier indicating time information of the requested subtitle data, and

query whether the subtitle data corresponding to the time information indicated by the time identifier is cached;

based on determining that the subtitle data corresponding to the time information is cached, push the cached subtitle data to the user terminal; and

delay the pushing of the video stream data to the user terminal to synchronize the subtitle text with the video stream data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2020
From: LUO, ZI HENG; ZHU, XIU MING; HU, XIAO HUA
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 053289/0603 →
Priority Claims (1)
CN 201810379453.8 · Apr 25, 2018 · national
Continuity (2)
Continuation PCTCN2019080299 · Mar 29, 2019
Related Publication 20200359104A1 · Nov 12, 2020
Cited By (3)
US 12,192,561 US 12,388,888 US 12,424,198