IP Library › Granted Patent US 12,609,117
Granted Patent B2
US 12,609,117 · App. 18/392,813 · Granted Apr 21, 2026

Apparatus and method for speech recognition using user prompt

Inventor: Seong Soo Yae (Hwaseong-si, KR)
Assignees: HYUNDAI MOTOR COMPANY; KIA CORPORATION
G10L15/22G10L21/0272G10L25/78G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,609,117
App. No.
18/392,813
Granted
Apr 21, 2026
Kind
B2
Abstract

An apparatus for speech recognition includes a user prompt playback unit configured to change a user prompt that induces utterance for speech recognition into a sound source and to play the sound source. The apparatus further includes a microphone detection signal extraction unit configured to extract a microphone detection signal when a user's speech is received as a speech signal through a microphone. The apparatus also includes a user prompt removal unit configured to remove the user prompt from the microphone detection signal. The apparatus further includes a speech recognition unit configured to recognize a speech based on a value from which the user prompt is removed. The apparatus also includes a response output unit configured to output a response related to the speech based on a result of recognizing the speech.

Claims (53)

1 . An apparatus for speech recognition, the apparatus comprising:

a user prompt playback unit configured to change a user prompt that induces utterance for speech recognition into a sound source and play the sound source;

a microphone detection signal extraction unit configured to extract a microphone detection signal when a speech of a user is received as a speech signal through a microphone;

a user prompt removal unit configured to remove the user prompt from the microphone detection signal;

a speech recognition unit configured to recognize a speech based on a value from which the user prompt is removed;

a response output unit configured to output a response related to the speech based on a result of recognizing the speech;

a selection domain extraction unit configured to extract a domain value (Rdn) and a speech recognition result value from the result of recognizing the speech;

a text-to-speech (TTS) result acquisition unit configured to acquire a result value (Rt) output as text from the speech recognition result value;

a keyword extraction unit configured to extract a keyword from the result value (Rt) and transmit the extracted keyword to a server;

a sound source acquisition unit configured to acquire a sound source (Rdn(x)) corresponding to the keyword extracted from the server;

a mixing unit configured to mix a sound source (Rdn(x)) corresponding to the keyword extracted from the server, the user prompt (Mp), and the result value (Rt); and

a result output unit configured to output a mixed result,

wherein the sound source acquisition unit is further configured to acquire, from a database, a background sound corresponding to a domain value (Rdn) output from the speech recognition result,

wherein the microphone detection signal includes a user prompt (Mp), a noise (En) introduced into a vehicle, and speech (Vs) contents, and

wherein the speech (Vs) contents include speech data identified as beginning-of-speech (BOS) based on an input signal differing from a predetermined sound source pattern being detected and include speech data identified as end-of-speech (EOS) based on only the predetermined sound source pattern being detected.

2 . The apparatus of claim 1 , wherein the noise (En) introduced into the vehicle includes at least one of a wind noise or a road noise.

3 . The apparatus of claim 1 , wherein the sound source corresponding to the domain is pre-stored in the database.

4 . The apparatus of claim 1 , wherein the background sound is acquired by mapping a user-selected sound or a sound stored or recorded using a user terminal to an identifier of the database.

5 . A method for speech recognition, the method comprising:

changing a user prompt that induces utterance for speech recognition into a sound source and playing the sound source;

extracting a microphone detection signal when a speech of a user is received as a speech signal through a microphone;

removing the user prompt from the microphone detection signal;

recognizing a speech based on a value from which the user prompt is removed;

outputting a response related to the speech based on a result of recognizing the speech;

extracting a domain value (Rdn) and a speech recognition result value from the result of recognizing the speech;

acquiring a result value (Rt) output as text from the speech recognition result value;

extracting a keyword from the result value (Rt) output as text and transmitting the extracted keyword to a server;

acquiring a sound source (Rdn(x)) corresponding to the keyword extracted from the server;

mixing the sound source Rdn(x) corresponding to the keyword extracted from the server, the user prompt (Mp), and the result value (Rt) output as text; and

outputting a mixed result,

wherein acquiring the sound source Rdn(x) further includes acquiring, from a database, a background sound corresponding to a domain value (Rdn) output from the speech recognition result,

wherein the microphone detection signal includes a user prompt (Mp), a noise (En) introduced into a vehicle, and speech (Vs) contents, and

wherein the speech (Vs) contents include speech data identified as beginning-of-speech (BOS) based on an input signal differing from a predetermined sound source pattern being detected and include speech data identified as end-of-speech (EOS) based on only the predetermined sound source pattern being detected.

6 . The method of claim 5 , wherein the noise (En) introduced into the vehicle includes at least one of wind noise or road noise.

7 . The method of claim 5 , wherein the sound source corresponding to the domain is pre-stored in the database.

8 . The method of claim 5 , wherein the background sound is acquired by mapping a user-selected sound or a sound stored or recorded using a user terminal to an identifier of the database.

9 . A non-transitory computer-readable recording medium, which stores a computer program including computer-executable instructions configured to be executable by an apparatus for speech recognition including a processor, to cause the processor to execute:

a function including changing a user prompt that induces utterance for speech recognition into a sound source and playing the sound source;

a function including extracting a microphone detection signal when a speech of a user is received as a speech signal through a microphone;

a function including removing the user prompt from the microphone detection signal;

a function including recognizing a speech based on a value from which the user prompt is removed;

a function including outputting a response related to the speech based on a result of recognizing the speech;

a function including extracting a domain value (Rdn) and a speech recognition result value from the result of recognizing the speech;

a function including acquiring a result value (Rt) output as text from the speech recognition result value;

a function including extracting a keyword from the result value (Rt) and transmit the extracted keyword to a server;

a function including acquiring a sound source (Rdn(x)) corresponding to the keyword extracted from the server;

a function including mixing a sound source (Rdn(x)) corresponding to the keyword extracted from the server, the user prompt (Mp), and the result value (Rt); and

a function including outputting a mixed result,

wherein the function of acquiring a sound source (Rdn(x)) further comprises acquiring, from a database, a background sound corresponding to a domain value (Rdn) output from the speech recognition result,

wherein the microphone detection signal includes a user prompt (Mp), a noise (En) introduced into a vehicle, and speech (Vs) contents, and

wherein the speech (Vs) contents include speech data identified as beginning-of-speech (BOS) based on an input signal differing from a predetermined sound source pattern being detected and include speech data identified as end-of-speech (EOS) based on only the predetermined sound source pattern being detected.

10 . The computer-readable recording medium of claim 9 , wherein the noise (En) introduced into the vehicle includes at least one of wind noise or road noise.

11 . The computer-readable recording medium of claim 9 , wherein the background sound is acquired by mapping a user-selected sound or a sound stored or recorded using a user terminal to an identifier of the database.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2023
From: YAE, SEONG SOO
To: HYUNDAI MOTOR COMPANY; KIA CORPORATION
Reel/Frame 065935/0373 →
Priority Claims (1)
KR 10-2023-0093884 · Jul 19, 2023 · national
Continuity (1)
Related Publication 20250029604A1 · Jan 23, 2025
References Cited (20)
US 6246986B1 · Ammicht · 2001 [cited by examiner]
US 9530432B2 · Herbig · 2016 [cited by examiner]
US 11043214B1 · Hedayatnia · 2021 [cited by examiner]
US 11579841B1 · Eich · 2023 [cited by examiner]
US 11646035B1 · Fan · 2023 [cited by examiner]
US 11721330B1 · Pandey · 2023 [cited by examiner]
US 11792570B1 · Govindaraju · 2023 [cited by examiner]
US 20090271189A1 · Agapi · 2009 [cited by examiner]
US 20160111088A1 · Park · 2016 [cited by examiner]
US 20200057487A1 · Sicconi · 2020 [cited by examiner]
US 20210375272A1 · Madwed · 2021 [cited by examiner]
US 20220093101A1 · Krishnan · 2022 [cited by examiner]
US 20220188361A1 · Botros · 2022 [cited by examiner]
US 20240095987A1 · Piramuthu · 2024 [cited by examiner]
US 20240412728A1 · Peterson · 2024 [cited by examiner]
US 20250029604A1 · Yae · 2025 [cited by examiner]
US 20250087208A1 · Lee · 2025 [cited by examiner]
US 20250095645A1 · Lee · 2025 [cited by examiner]
US 20250124915A1 · Chun · 2025 [cited by examiner]
US 20250200293A1 · Liu · 2025 [cited by examiner]