Multi-format single stream scalable coding for multi-language audio
Multi-format single stream scalable coding for multi-language audio includes separating background and speech audio of a video stream uploaded to an online video platform and separately encoding the background and speech audio to different coding layers using a scalable video coding schema. During encoding, the background audio is encoded to a base layer bitstream, different language versions of the speech audio are encoded to different enhancement layer bitstreams, and language selection precedence data is embedded to signal to a decoder which of those enhancement layer bitstreams to decode for playback of the video. During decoding, the appropriate enhancement layer bitstream is decoded to obtain speech audio in a desired language, and the base layer bitstream is decoded to obtain the background audio. The background audio and the speech audio are re-mixed into a combined audio stream, which is transmuxed with a video component to produce a media stream for playback.
1 . A method, comprising:
separating audio of an input video stream uploaded to an online video platform into background audio and speech audio;
converting the speech audio into multiple language speech audio versions;
encoding the background audio to a base layer bitstream;
encoding each of the multiple language speech audio versions to a different enhancement layer bitstream, wherein each of the different enhancement layer bitstreams represents the speech audio in one of multiple different languages;
combining, into an encoded audio stream, the base layer bitstream, each of the enhancement layer bitstreams, and language selection precedence data for the enhancement layer bitstreams, wherein the language selection precedence data is embedded within a header of the encoded audio stream; and
outputting the encoded audio stream for storage or further processing.
2 . The method of claim 1 , wherein separating the audio of the input video stream uploaded to the online video platform into the background audio and the speech audio comprises:
performing blind audio source separation against the audio of the input video stream.
3 . The method of claim 2 , wherein performing the blind audio source separation against the audio of the input video stream comprises:
using a machine learning model trained for speech and speaker identification to perform the blind audio source separation.
4 . The method of claim 1 , wherein converting the speech audio into the multiple language speech audio versions comprises:
converting the speech audio into text;
translating the text into each of the multiple different languages; and
converting, for each of the multiple different languages, the translated text into one of the multiple language speech audio versions.
5 . The method of claim 1 , wherein converting the speech audio into the multiple language speech audio versions comprises:
using a large language model trained for speech audio conversion to directly translate the speech audio into each of the multiple language speech audio versions.
6 . The method of claim 1 , comprising:
determining the language selection precedence data according to a prioritization of the multiple language speech audio versions.
7 . The method of claim 6 , wherein combining, into the encoded audio stream, the base layer bitstream, each of the enhancement layer bitstreams, and the language selection precedence data for the enhancement layer bitstreams comprises:
embedding the language selection precedence data within metadata of the encoded audio stream or within supplemental enhancement information within the encoded audio stream.
8 . The method of claim 6 , wherein the language selection precedence data indicates, to a decoder, an enhancement layer bitstream of the encoded audio stream to decode for playback of the input video stream.
9 . The method of claim 1 , wherein the speech audio corresponds to one or both of diegetic speech or non-diegetic speech.
10 . The method of claim 1 , wherein the background audio is encoded to the base layer bitstream using a first audio channel format and the speech audio is encoded to the enhancement layer bitstreams using a second audio channel format.
11 . A method, comprising:
obtaining an encoded audio stream including a base layer bitstream, multiple enhancement layer bitstreams each representing speech audio of an input video stream used to produce the encoded audio stream in one of multiple different languages, and language selection precedence data embedded within a header of the encoded audio stream;
decoding, from the audio stream, the base layer bitstream into background audio;
decoding, from the audio stream, an enhancement layer bitstream indicated by the language selection precedence data into speech audio;
re-mixing the background audio and the speech audio into an audio stream;
combining the audio stream and a video stream into a single media stream; and
outputting the single media stream for playback or further processing.
12 . The method of claim 11 , wherein decoding the enhancement layer bitstream indicated by the language selection precedence data into the speech audio comprises:
reading the language selection precedence data from metadata of the encoded audio stream or supplemental enhancement information within the encoded audio bitstream; and
determining a priority language for the audio stream based on the language selection precedence data, wherein the enhancement layer bitstream corresponds to the priority language.
13 . The method of claim 12 , comprising:
decoding, from the audio stream, a different enhancement layer bitstream for playback within the single media stream, wherein the enhancement layer bitstream corresponds to a first language version of the speech audio and the different enhancement layer bitstream corresponds to a second language version of the speech audio.
14 . The method of claim 13 , wherein re-mixing the background audio and the speech audio into the audio stream comprises:
re-mixing a first chunk of the background audio and a first chunk of the speech audio in the first language into a first audio stream chunk, and
wherein the method comprises:
re-mixing a second chunk of the background audio and a second chunk of the speech audio in the second language into a second audio stream chunk.
15 . The method of claim 13 , wherein the different enhancement layer bitstream is decoded based on a selection, at a playback device to which the single media stream is output, of the second language.
16 . The method of claim 11 , wherein multiple enhancement layer bitstreams are decoded into different audio speech versions according to the language selection precedence data and the different audio speech versions are re-mixed with the background audio.
17 . A system, comprising:
one or more servers used with an online video platform and configured to:
obtain an input video stream from a first device;
encode background audio of the input video stream to a base layer bitstream;
encode each of multiple language versions of speech audio of the input video stream to a different enhancement layer bitstream, wherein each of the different enhancement layer bitstreams represents the speech audio in one of multiple different languages;
combine, into an encoded audio stream, the base layer bitstream, each of the enhancement layer bitstreams, and language selection precedence data for the enhancement layer bitstreams, wherein the language selection precedence data is embedded within a header of the encoded audio stream; and
output the encoded audio stream for decoding at a second device responsive to a playback request for a video associated with the encoded audio stream,
wherein the encoded audio stream configures the second device to decode one of the enhancement layer bitstreams for playback of a language version of the speech audio along with the background audio according to the language selection precedence data.
18 . The system of claim 17 , wherein the one or more servers are configured to:
separate audio of the input video stream into the background audio and the speech audio; and
convert the speech audio into the multiple language versions of the speech audio.
19 . The system of claim 17 , wherein the language selection precedence data is embedded within metadata of the encoded audio stream or within supplemental enhancement information within the encoded audio stream.
20 . The system of claim 17 , wherein the language selection precedence data is determined according to a prioritization of the multiple language versions of the speech audio.