IP Library Granted Patent US 12,640,139
Granted Patent B2
US 12,640,139 · App. 18/585,204 · Granted May 26, 2026

Method and apparatus for improving performance of artificial intelligence model using speech recognition results as text input with delay times

Inventors: Seung Hi Kim (Daejeon, KR); Jeong Uk Bang (Daejeon, KR); Seung Yun (Daejeon, KR)
Assignee: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
G10L15/063G10L15/26G10L2015/0631
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,640,139
App. No.
18/585,204
Granted
May 26, 2026
Kind
B2
Abstract

The present disclosure relates to a method and device for improving the performance of an AI model that uses voice recognition results as text input. A method of training an AI model according to an embodiment of the present disclosure may include: generating first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription; generating second time information by adding a pre-configured delay time to the first time information; generating a modified transcription based on an end time of a last word among the plurality of words and the second time information; and performing training of the AI model based on a second training sample including the voice and the modified transcription.

Claims (41)

1 . A method of training an artificial intelligence (AI) model, the method comprising:

generating first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription;

generating second time information by adding a pre-configured delay time to the first time information;

generating a modified transcription based on an end time of a last word among the plurality of words and the second time information; and

performing training of the AI model based on a second training sample including the voice and the modified transcription,

wherein the pre-configured delay time is progressively increased in correspondence with an advancement of a stage of the training, and

wherein an amount of word removal for generating the modified transcription is increased in response to an increase of the pre-configured delay time.

2 . The method of claim 1 , wherein the pre-configured delay time is related to a delay time for text output of a voice recognizer.

3 . The method of claim 1 , wherein the first time information includes information on an end time for each word for the plurality of words.

4 . The method of claim 3 , wherein the second time information is generated by adding the pre-configured delay time to the end time for each word for the plurality of words.

5 . The method of claim 1 , wherein the modified transcription is generated by removing one or more words from among the plurality of words whose end time for each word is greater than or equal to the end time based on the second time information.

6 . The method of claim 1 , wherein the pre-configured delay time is set to a value of 0 in an initial section of the training.

7 . The method of claim 1 , wherein the pre-configured delay time is set between a 0 value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.

8 . The method of claim 1 , wherein the pre-configured delay time is set between a minimum delay time value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.

9 . The method of claim 1 , wherein the pre-configured delay time is set to increase up to a delay time of a voice recognizer as a stage of the training advanced.

10 . An apparatus for training an artificial intelligence (AI) model, the apparatus comprising:

a processor and a memory,

wherein the processor is configured to:

generate first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription;

generate second time information by adding a pre-configured delay time to the first time information;

generate a modified transcription based on an end time of a last word among the plurality of words and the second time information; and

perform training of the AI model based on a second training sample including the voice and the modified transcription,

wherein the pre-configured delay time is progressively increased in correspondence with an advancement of a stage of the training, and

wherein an amount of word removal for generating the modified transcription is increased in response to an increase of the pre-configured delay time.

11 . The apparatus of claim 10 , wherein the pre-configured delay time is related to a delay time for text output of a voice recognizer.

12 . The apparatus of claim 10 , wherein the first time information includes information on an end time for each word for the plurality of words.

13 . The apparatus of claim 12 , wherein the second time information is generated by adding the pre-configured delay time to the end time for each word for the plurality of words.

14 . The apparatus of claim 10 , wherein the modified transcription is generated by removing one or more words from among the plurality of words whose end time for each word is greater than or equal to the end time based on the second time information.

15 . The apparatus of claim 10 , wherein the pre-configured delay time is set to a value of 0 in an initial section of the training.

16 . The apparatus of claim 10 , wherein the pre-configured delay time is set between a 0 value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.

17 . The apparatus of claim 10 , wherein the pre-configured delay time is set between a minimum delay time value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.

18 . The apparatus of claim 10 , wherein the pre-configured delay time is set to increase up to a delay time of a voice recognizer as a stage of the training advanced.

19 . One or more non-transitory computer readable storage medium storing one or more instructions,

wherein the one or more instructions are executed by one or more processors and control an apparatus for training an artificial intelligence (AI) model to:

generate first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription;

generate second time information by adding a pre-configured delay time to the first time information;

generate a modified transcription based on an end time of a last word among the plurality of words and the second time information; and

perform training of the AI model based on a second training sample including the voice and the modified transcription,

wherein the pre-configured delay time is progressively increased in correspondence with an advancement of a stage of the training, and

wherein an amount of word removal for generating the modified transcription is increased in response to an increase of the pre-configured delay time.

20 . The computer readable storage medium of claim 19 , wherein the pre-configured delay time is related to a delay time for text output of a voice recognizer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2024
From: KIM, SEUNG HI; BANG, JEONG UK; YUN, SEUNG
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 066539/0315 →
Priority Claims (1)
KR 10-2023-0077050 · Jun 15, 2023 · national
Continuity (1)
Related Publication 20240420682A1 · Dec 19, 2024
References Cited (29)
US 6151574A · Lee · 2000 [cited by examiner]
US 9613624B1 · Kramer · 2017 [cited by examiner]
US 9697201B2 · Gao et al. · 2017 [cited by applicant]
US 12190871B1 · Siagian · 2025 [cited by examiner]
US 20180366120A1 · Ushio · 2018 [cited by examiner]
US 20200027444A1 · Prabhavalkar · 2020 [cited by examiner]
US 20200075033A1 · Hijazi · 2020 [cited by examiner]
US 20200357388A1 · Zhao · 2020 [cited by examiner]
US 20210118449A1 · Kim · 2021 [cited by examiner]
US 20210264921A1 · Reece · 2021 [cited by examiner]
US 20220189461A1 · Zhao · 2022 [cited by examiner]
US 20220351720A1 · Alon · 2022 [cited by examiner]
US 20230061505A1 · Oh et al. · 2023 [cited by applicant]
US 20230162737A1 · Kirkpatrick et al. · 2023 [cited by applicant]
US 20230267919A1 · Ponçot · 2023 [cited by examiner]
US 20240062744A1 · Liu · 2024 [cited by examiner]
US 20240087596A1 · Nandwana · 2024 [cited by examiner]
US 20240347042A1 · Weninger · 2024 [cited by examiner]
US 20240347047A1 · Weninger · 2024 [cited by examiner]
US 20240395246A1 · Mcquinn · 2024 [cited by examiner]
US 20240404512A1 · Yu · 2024 [cited by examiner]
JP 6578049B2 · 2019 [cited by applicant]
KR 1020220112596A · 2022 [cited by applicant]
KR 102504445B1 · 2023 [cited by applicant]
Nguyen, et al. “Super-human performance in online low-latency recognition of conversational speech.” arXiv preprint arXiv:2010.03449, Jul. 2021, pp. 1-5. (Year: 2021). [cited by examiner]
Shangguan, Yuan, et al. “Dissecting user-perceived latency of on-device E2E speech recognition.” arXiv preprint arXiv:2104.02207, Aug. 2021, pp. 1-5. (Year: 2021). [cited by examiner]
Amalia Istiqlali Adiba et al., “Towards Immediate Backchannel Generation Using Attention-Based Early Prediction Model”, ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing. [cited by applicant]
Ricardo Falcon-Perez, “Curriculum Learning With Audio Domain Data Augmentation for Sound Event Localization and Detection”, Detection and Classification of Acoustic Scenes and Events 2022. [cited by applicant]
Daniel S. Park et al., “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition”, arXiv:1904.08779v3 [eess.AS] Dec. 3, 2019. [cited by applicant]