IP Library Granted Patent US 12682888
Granted Patent B2
US 12682888 · App. 18/559,398 · Granted Jul 14, 2026

Method and device for generating speech recognition training set

Inventor: Li Fu (Beijing, CN)
Assignee: Jingdong Technology Holding Co., Ltd.
G10L15/063G06V10/764G06V10/774G06V20/41G06V20/63G06V30/153G10L25/57G10L25/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682888
App. No.
18/559,398
Granted
Jul 14, 2026
Kind
B2
Abstract

Disclosed in the present disclosure are a method and apparatus for generating a speech recognition training set. The method may include: acquiring a to-be-processed audio and a to-be-processed video, where the to-be-processed video comprises text information corresponding to the to-be-processed audio; recognizing the to-be-processed audio to obtain an audio text; recognizing text information in the to-be-processed video to obtain a video text; and using, based on consistency of the audio text with the video text, the to-be-processed audio as a speech sample and the video text as a label to obtain the speech recognition training set.

Claims (57)

1 . A method for generating a speech recognition training set, the method comprising:

acquiring a to-be-processed audio and a to-be-processed video, wherein the to-be-processed video comprises text information corresponding to the to-be-processed audio;

recognizing the to-be-processed audio to obtain an audio text;

recognizing text information in the to-be-processed video to obtain a video text; and

using, based on consistency of the audio text with the video text, the to-be-processed audio as a speech sample and the video text as a label to obtain the speech recognition training set, comprising:

for each video frame sequence of a plurality of video frame sequences, in the to-be-processed video, performing operations as follows: splicing the text information included in each video frame in the video frame sequence, in units of one video frame text in at least one video frame text recognized from video frames in the video frame sequence, to obtain a plurality of video frame sequence texts corresponding to the video frame sequence;

determining a target video frame sequence text, based on an editing distance between each video frame sequence text in the plurality of video frame sequence texts and a target audio clip text, wherein the target audio clip text is an audio clip text corresponding to an audio clip corresponding to the video frame sequence; and

using each audio clip of a plurality of audio clips in the to-be-processed audio as the speech sample, and the target video frame sequence text corresponding to the audio clip as the label, to obtain the speech recognition training set.

2 . The method according to claim 1 , wherein recognizing the to-be-processed audio to obtain the audio text, comprises:

deleting a silent portion in the to-be-processed audio based on a mute detection algorithm, to obtain the plurality of audio clips that are not mute; and

recognizing the plurality of audio clips to obtain a plurality of audio clip texts included in the audio text.

3 . The method according to claim 2 , wherein recognizing text information in the to-be-processed video to obtain the video text, comprises:

determining, from the to-be-processed video, the plurality of video frame sequences, each of the plurality of video frame sequences corresponding one-to-one to the plurality of audio clips; and

recognizing text information in each video frame in the plurality of video frame sequences to obtain the video frame text included in the video text.

4 . The method according to claim 1 , wherein splicing the text information included in each video frame in the video frame sequence, in units of one video frame text in at least one video frame text recognized from video frames in the video frame sequence, to obtain the plurality of video frame sequence texts corresponding to the video frame sequence, comprises:

for each video frame in the video frame sequence that comprises text information, performing operations as follows:

determining a plurality of to-be-spliced texts corresponding to the video frame, and splicing the plurality of to-be-spliced texts with at least one video frame text in the video frame to obtain a plurality of spliced texts; and

selecting a preset number of spliced texts from the plurality of spliced texts, based on an editing distance between the plurality of spliced texts and the target audio clip text, as the plurality of to-be-spliced texts corresponding to a next video frame of the video frame.

5 . The method according to claim 1 , wherein the method further comprises:

for each video frame sequence in the plurality of video frame sequences, in response to determining that the editing distance between the target video frame sequence text corresponding to the video frame sequence and the target audio clip text is greater than a preset distance threshold, deleting training samples corresponding to the video frame sequence in the speech recognition training set.

6 . An apparatus for generating a speech recognition training set, the apparatus comprising:

one or more processors; and

a memory, storing one or more programs thereon,

wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to perform operations, the operations comprising:

acquiring a to-be-processed audio and a to-be-processed video, wherein the to-be-processed video comprises text information corresponding to the to-be-processed audio;

recognizing the to-be-processed audio to obtain an audio text;

recognizing text information in the to-be-processed video to obtain a video text; and

using, based on consistency of the audio text with the video text, the to-be-processed audio as a speech sample and the video text as a label to obtain the speech recognition training set, comprising:

for each video frame sequence of a plurality of video frame sequences in the to-be-processed video, performing operations as follows: splicing the text information included in each video frame in the video frame sequence, in units of one video frame text in at least one video frame text recognized from video frames in the video frame sequence, to obtain a plurality of video frame sequence texts corresponding to the video frame sequence; determining a target video frame sequence text, based on an editing distance between each video frame sequence text in the plurality of video frame sequence texts and a target audio clip text, wherein the target audio clip text is an audio clip text corresponding to an audio clip corresponding to the video frame sequence; and

using each audio clip of a plurality of audio clips as the speech sample in the to-be-processed audio, and the target video frame sequence text corresponding to the audio clip as the label, to obtain the speech recognition training set.

7 . The apparatus according to claim 6 , wherein recognizing the to-be-processed audio to obtain the audio text, comprises:

deleting a silent portion in the to-be-processed audio based on a mute detection algorithm, to obtain the plurality of audio clips that are not mute; and

recognizing the plurality of audio clips to obtain a plurality of audio clip texts included in the audio text.

8 . The apparatus according to claim 7 , wherein recognizing text information in the to-be-processed video to obtain the video text, comprises:

determining, from the to-be-processed video, the plurality of video frame sequences, each of the plurality of video frame sequences corresponding one-to-one to the plurality of audio clips; and recognizing text information in each video frame in the plurality of video frame sequences to obtain the video frame text included in the video text.

9 . The apparatus according to claim 6 , wherein splicing the text information included in each video frame in the video frame sequence, in units of one video frame text in at least one video frame text recognized from video frames in the video frame sequence, to obtain the plurality of video frame sequence texts corresponding to the video frame sequence, comprises:

for each video frame in the video frame sequence that comprises text information, performing operations as follows: determining a plurality of to-be-spliced texts corresponding to the video frame, and splicing the plurality of to-be-spliced texts with at least one video frame text in the video frame to obtain a plurality of spliced texts; and selecting a preset number of spliced texts from the plurality of spliced texts, based on an editing distance between the plurality of spliced texts and the target audio clip text, as the plurality of to-be-spliced texts corresponding to a next video frame of the video frame.

10 . The apparatus according to claim 6 , wherein the operations further comprise:

for each video frame sequence in the plurality of video frame sequences, in response to determining that the editing distance between the target video frame sequence text corresponding to the video frame sequence and the target audio clip text is greater than a preset distance threshold, deleting training samples corresponding to the video frame sequence in the speech recognition training set.

11 . A non-transitory computer readable medium, storing a computer program thereon, wherein, the program, when executed by a processor, causes the processor to perform operations, the operations comprising:

acquiring a to-be-processed audio and a to-be-processed video, wherein the to-be-processed video comprises text information corresponding to the to-be-processed audio;

recognizing the to-be-processed audio to obtain an audio text;

recognizing text information in the to-be-processed video to obtain a video text; and

using, based on consistency of the audio text with the video text, the to-be-processed audio as a speech sample and the video text as a label to obtain a speech recognition training set, comprising:

for each video frame sequence of a plurality of video frame sequences in the to-be-processed video, performing operations as follows: splicing the text information included in each video frame in the video frame sequence, in units of one video frame text in at least one video frame text recognized from video frames in the video frame sequence, to obtain a plurality of video frame sequence texts corresponding to the video frame sequence; determining a target video frame sequence text, based on an editing distance between each video frame sequence text in the plurality of video frame sequence texts and a target audio clip text, wherein the target audio clip text is an audio clip text corresponding to an audio clip corresponding to the video frame sequence; and

using each audio clip of a plurality of audio clips as the speech sample in the to-be-processed audio, and the target video frame sequence text corresponding to the audio clip as the label, to obtain the speech recognition training set.

12 . The non-transitory computer readable medium according to claim 11 , wherein recognizing the to-be-processed audio to obtain the audio text, comprises:

deleting a silent portion in the to-be-processed audio based on a mute detection algorithm, to obtain the plurality of audio clips that are not mute; and

recognizing the plurality of audio clips to obtain a plurality of audio clip texts included in the audio text.

13 . The non-transitory computer readable medium according to claim 12 , wherein recognizing text information in the to-be-processed video to obtain the video text, comprises:

determining, from the to-be-processed video, the plurality of video frame sequences, each of the plurality of video frame sequences corresponding one-to-one to the plurality of audio clips; and recognizing text information in each video frame in the plurality of video frame sequences to obtain the video frame text included in the video text.

14 . The non-transitory computer readable medium according to claim 11 , wherein splicing the text information included in each video frame in the video frame sequence, in units of one video frame text in at least one video frame text recognized from video frames in the video frame sequence, to obtain the plurality of video frame sequence texts corresponding to the video frame sequence, comprises:

for each video frame in the video frame sequence that comprises text information, performing operations as follows:

determining a plurality of to-be-spliced texts corresponding to the video frame, and splicing the plurality of to-be-spliced texts with at least one video frame text in the video frame to obtain a plurality of spliced texts; and

selecting a preset number of spliced texts from the plurality of spliced texts, based on an editing distance between the plurality of spliced texts and the target audio clip text, as the plurality of to-be-spliced texts corresponding to a next video frame of the video frame.

15 . The non-transitory computer readable medium according to claim 11 , wherein the operations further comprise:

for each video frame sequence in the plurality of video frame sequences, in response to determining that the editing distance between the target video frame sequence text corresponding to the video frame sequence and the target audio clip text is greater than a preset distance threshold, deleting training samples corresponding to the video frame sequence in the speech recognition training set.