IP Library › Granted Patent US 12,602,553
Granted Patent B2
US 12,602,553 · App. 18/245,802 · Granted Apr 14, 2026

Speech translation method, device, and storage medium

Inventors: Lei Li (Beijing, CN); Mingxuan Wang (Beijing, CN); Qianqian Dong (Beijing, CN); Chengqi Zhao (Beijing, CN)
Assignee: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
G06F40/47G10L15/02G10L15/063G10L15/16G10L15/1815
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,553
App. No.
18/245,802
Granted
Apr 14, 2026
Kind
B2
Abstract

Provided are a speech translation method, a device, and a storage medium. The method includes: extracting, through an encoder of an end-to-end speech translation model, the semantic feature of a to-be-processed speech; decoding, through a decoder of the end-to-end speech translation model, a source language text corresponding to the semantic feature from the semantic feature; decoding, through the decoder of the end-to-end speech translation model, the semantic feature according to the source language text to obtain a text sequence corresponding to the semantic feature; and splitting the text sequence to obtain a target language text corresponding to the to-be-processed speech.

Claims (60)

1 . A speech translation method, comprising:

extracting, through an encoder of an end-to-end speech translation model, a semantic feature of a to-be-processed speech;

decoding, through a decoder of the end-to-end speech translation model, a source language text corresponding to the semantic feature from the semantic feature;

decoding, through the decoder of the end-to-end speech translation model, the semantic feature according to the source language text to obtain a text sequence corresponding to the semantic feature, wherein the text sequence comprises the source language text and a target language text corresponding to the source language text; and

splitting the text sequence to obtain the target language text corresponding to the to-be-processed speech,

wherein a training process of the end-to-end speech translation model comprises:

pre-training, according to a text translation sample, the decoder of the end-to-end speech translation model to obtain an initial decoder; and

initializing the decoder of the end-to-end speech translation model based on the initial decoder; and training the initialized end-to-end speech translation model,

wherein the text translation sample comprises a source language sample text and a target language sample text corresponding to the source language sample text, and

wherein pre-training, according to the text translation sample, the decoder of the end-to-end speech translation model to obtain the initial decoder comprises:

splicing the source language sample text and the target language sample text to obtain a spliced sample sequence; and

pre-training, based on a masked cross-entropy loss function, the decoder by using the source language sample text and an all-zero vector as input of the decoder of the end-to-end speech translation model and the spliced sample sequence as desired output to obtain the initial decoder, wherein the masked cross-entropy loss function is configured to mask a prediction loss of a source language prediction text corresponding to the all-zero vector.

2 . The method according to claim 1 , wherein the encoder comprises a first encoder and a second encoder, and extracting, through the encoder of the end-to-end speech translation model, the semantic feature of the to-be-processed speech comprises:

extracting, through the first encoder, an acoustic feature of the to-be-processed speech;

detecting a blank frame and a duplicate frame in the acoustic feature, eliminating the blank frame and combining the duplicate frame to obtain a contracted acoustic feature; and

extracting, through the second encoder, the semantic feature in the contracted acoustic feature.

3 . The method according to claim 2 , wherein detecting the blank frame and the duplicate frame in the acoustic feature comprises:

detecting, based on a spike characteristic of a probability distribution of a connectionist temporal classification loss function, the blank frame and the duplicate frame in the acoustic feature.

4 . An electronic device, comprising a processor and a memory storing computer programs, wherein the computer programs, when executed by the processor, implement the following:

extracting, through an encoder of an end-to-end speech translation model, a semantic feature of a to-be-processed speech;

decoding, through a decoder of the end-to-end speech translation model, a source language text corresponding to the semantic feature from the semantic feature;

decoding, through the decoder of the end-to-end speech translation model, the semantic feature according to the source language text to obtain a text sequence corresponding to the semantic feature, wherein the text sequence comprises the source language text and a target language text corresponding to the source language text; and

splitting the text sequence to obtain the target language text corresponding to the to-be-processed speech,

wherein a training process of the end-to-end speech translation model comprises:

pre-training, according to a text translation sample, the decoder of the end-to-end speech translation model to obtain an initial decoder; and

initializing the decoder of the end-to-end speech translation model based on the initial decoder; and training the initialized end-to-end speech translation model,

wherein the text translation sample comprises a source language sample text and a target language sample text corresponding to the source language sample text, and

wherein pre-training, according to the text translation sample, the decoder of the end-to-end speech translation model to obtain the initial decoder comprises:

splicing the source language sample text and the target language sample text to obtain a spliced sample sequence; and

pre-training, based on a masked cross-entropy loss function, the decoder by using the source language sample text and an all-zero vector as input of the decoder of the end-to-end speech translation model and the spliced sample sequence as desired output to obtain the initial decoder, wherein the masked cross-entropy loss function is configured to mask a prediction loss of a source language prediction text corresponding to the all-zero vector.

5 . A non-transitory computer-readable storage medium storing computer programs, wherein the computer programs, when executed by a processor, implement the following:

extracting, through an encoder of an end-to-end speech translation model, a semantic feature of a to-be-processed speech;

decoding, through a decoder of the end-to-end speech translation model, a source language text corresponding to the semantic feature from the semantic feature;

decoding, through the decoder of the end-to-end speech translation model, the semantic feature according to the source language text to obtain a text sequence corresponding to the semantic feature, wherein the text sequence comprises the source language text and a target language text corresponding to the source language text; and

splitting the text sequence to obtain the target language text corresponding to the to-be-processed speech,

wherein a training process of the end-to-end speech translation model comprises:

pre-training, according to a text translation sample, the decoder of the end-to-end speech translation model to obtain an initial decoder; and

initializing the decoder of the end-to-end speech translation model based on the initial decoder; and training the initialized end-to-end speech translation model,

wherein the text translation sample comprises a source language sample text and a target language sample text corresponding to the source language sample text, and

wherein pre-training, according to the text translation sample, the decoder of the end-to-end speech translation model to obtain the initial decoder comprises:

splicing the source language sample text and the target language sample text to obtain a spliced sample sequence; and

pre-training, based on a masked cross-entropy loss function, the decoder by using the source language sample text and an all-zero vector as input of the decoder of the end-to-end speech translation model and the spliced sample sequence as desired output to obtain the initial decoder, wherein the masked cross-entropy loss function is configured to mask a prediction loss of a source language prediction text corresponding to the all-zero vector.

6 . The method according to claim 2 , wherein a training process of the first encoder comprises:

training, based on a connectionist temporal classification loss function, the first encoder by using a sample speech in a training sample set of the end-to-end speech translation model as input of the first encoder and a sample phoneme sequence corresponding to the sample speech as desired output.

7 . The electronic device according to claim 4 , wherein the encoder comprises a first encoder and a second encoder, and the computer programs implement extracting, through the encoder of the end-to-end speech translation model, the semantic feature of the to-be-processed speech by:

extracting, through the first encoder, an acoustic feature of the to-be-processed speech;

detecting a blank frame and a duplicate frame in the acoustic feature, eliminating the blank frame and combining the duplicate frame to obtain a contracted acoustic feature; and

extracting, through the second encoder, the semantic feature in the contracted acoustic feature.

8 . The electronic device according to claim 7 , wherein the computer programs implement detecting the blank frame and the duplicate frame in the acoustic feature by:

detecting, based on a spike characteristic of a probability distribution of a connectionist temporal classification loss function, the blank frame and the duplicate frame in the acoustic feature.

9 . The storage medium according to claim 5 , wherein the encoder comprises a first encoder and a second encoder, and the computer programs implement extracting, through the encoder of the end-to-end speech translation model, the semantic feature of the to-be-processed speech by:

extracting, through the first encoder, an acoustic feature of the to-be-processed speech;

detecting a blank frame and a duplicate frame in the acoustic feature, eliminating the blank frame and combining the duplicate frame to obtain a contracted acoustic feature; and

extracting, through the second encoder, the semantic feature in the contracted acoustic feature.

10 . The electronic device according to claim 4 , wherein a training process of the first encoder comprises:

training, based on a connectionist temporal classification loss function, the first encoder by using a sample speech in a training sample set of the end-to-end speech translation model as input of the first encoder and a sample phoneme sequence corresponding to the sample speech as desired output.

11 . The storage medium according to claim 9 , wherein detecting the blank frame and the duplicate frame in the acoustic feature comprises:

detecting, based on a spike characteristic of a probability distribution of a connectionist temporal classification loss function, the blank frame and the duplicate frame in the acoustic feature.

12 . The storage medium according to claim 5 , wherein a training process of the first encoder comprises:

training, based on a connectionist temporal classification loss function, the first encoder by using a sample speech in a training sample set of the end-to-end speech translation model as input of the first encoder and a sample phoneme sequence corresponding to the sample speech as desired output.

Priority Claims (1)
CN 202010987456.7 · Sep 18, 2020 · national
Continuity (1)
Related Publication 20240028841A1 · Jan 25, 2024
References Cited (17)
US 4803729A · Baker · 1989 [cited by examiner]
US 10249294B2 · Kim et al. · 2019 [cited by applicant]
US 11694677B2 · Lee · 2023 [cited by examiner]
US 20180075844A1 · Kim et al. · 2018 [cited by applicant]
US 20210056975A1 · Peng · 2021 [cited by examiner]
US 20210065690A1 · Indurthi · 2021 [cited by examiner]
CN 108231062A · 2018 [cited by applicant]
CN 110556100A · 2019 [cited by applicant]
CN 111326157A · 2020 [cited by applicant]
CN 111368559A · 2020 [cited by applicant]
CN 112183120A · 2021 [cited by applicant]
Bart: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension, Lewis et al., arXiv (Year: 2019). [cited by examiner]
Mass: Masked Sequence to Sequence Pre-training for Language Generation, Song et al., arXiv (Year: 2019). [cited by examiner]
Search Report issued Oct. 28, 2021 for PCT Application No. PCT/CN2021/116232, English translation, (4 pages). [cited by applicant]
Written Opinion for International Application No. PCT/CN2021/116232, mailed Oct. 28, 2021, 9 Pages. [cited by applicant]
First Search Report issued Jul. 11, 2023 in Chinese Application No. 202010987456.7, with English translation (4 pages). [cited by applicant]
First Office Action issued Jul. 13, 2023 in Chinese Application No. 202010987456.7, with English translation (12 pages). [cited by applicant]