IP Library › Granted Patent US 12,744,040
Granted Patent B2
US 12,744,040 · App. 18/737,673 · Granted Sep 22, 2026

Method for processing misrecognized audio signals, and device therefor

Inventors: Chanhee Choi (Suwon-si, KR); Chansik Bok (Suwon-si, KR); Hyundon Yoon (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L15/22G10L21/02G10L2015/223G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,040
App. No.
18/737,673
Granted
Sep 22, 2026
Kind
B2
Abstract

A method, an electronic device, and a non-transitory computer-readable medium storing instructions may be provided. The method may include receiving an audio signal; based on at least one preset trigger word being included in the received audio signal, determining whether the at least one trigger word included in the audio signal is misrecognized; based on the determining that the at least one preset trigger word is misrecognized, requesting an additional input from a user; and based on the additional input received in response to the request and the audio signal, executing a function corresponding to audio recognition.

Claims (46)

1 . A method of processing a misrecognized audio signal in an electronic device, the method comprising:

receiving an audio signal;

detecting that at least one preset trigger word is included in the received audio signal;

after detecting that the at least one preset trigger word is included in the received audio signal, determining, by a processor of the electronic device, whether the detected trigger word is misrecognized by measuring, by a wake-up word engine executed by the processor, a similarity between the at least one preset trigger word and the received audio signal based on frequency components extracted from the received audio signal,

wherein the wake-up word engine comprises an acoustic model trained with acoustic information for the preset at least one trigger word;

based on determining that the detected trigger word is misrecognized, requesting an additional input from a user; and

based on the additional input from the user received in response to the requesting and the received audio signal, controlling, by the electronic device, whether to execute a function corresponding to audio recognition or to terminate the audio recognition.

2 . The method of claim 1 , wherein the determining of whether the at least one preset trigger word included in the audio signal is misrecognized is based on a history of execution of the function corresponding to the audio recognition within a preset first time.

3 . The method of claim 1 , wherein the determining of whether the at least one preset trigger word included in the audio signal is misrecognized is based on there being no history of execution of the function corresponding to the audio recognition within a preset first time.

4 . The method of claim 1 , wherein the determining of whether the at least one preset trigger word included in the audio signal is misrecognized comprises:

synchronizing the received audio signal and a reference audio signal output from another electronic device; and

determining that the at least one preset trigger word included in the audio signal is misrecognized when a similarity between the synchronized audio signal and the synchronized reference audio signal is greater than or equal to a preset first threshold.

5 . The method of claim 4 , wherein the requesting of the additional input from the user comprises:

adjusting an intensity of the reference audio signal to less than or equal to a preset second threshold; and

requesting the additional input for the at least one preset trigger word from the user.

6 . The method of claim 1 , wherein the determining that the at least one preset trigger word included in the audio signal is misrecognized is based on the received audio signal comprising at least one input signal in addition to the at least one preset trigger word.

7 . The method of claim 6 , wherein the requesting the additional input from the user comprises requesting, from the user, another additional input related to whether to perform the at least one input signal included in the audio signal.

8 . The method of claim 6 , wherein the determining that the at least one preset trigger word is misrecognized comprises:

dividing the audio signal into multiple sections, wherein the multiple sections do not comprise a section corresponding to the at least one preset trigger word; and

determining whether the at least one preset trigger word included in the audio signal is misrecognized based on at least one of: energy values of the multiple sections or zero-crossing rates (ZCRs) of the multiple sections.

9 . The method of claim 1 , wherein the determining that the at least one preset trigger word is misrecognized comprises:

determining that the at least one preset trigger word has the measured similarity that is greater than or equal to a third threshold.

10 . The method of claim 9 , wherein when more than one preset trigger words have a corresponding measured similarity that is greater than or equal to the third threshold, the determining that the at least one preset trigger word is misrecognized comprises:

determining that a word, among the more than one preset trigger words having the corresponding measured similarity greater than or equal to the third threshold, has the corresponding measured similarity that is smaller than a fourth threshold.

11 . An electronic device for processing a misrecognized audio signal, the electronic device comprising:

a memory storing one or more instructions; and

at least one processor configured to execute the one or more instructions,

wherein the at least one processor is configured to:

detect that based on at least one preset trigger word is included in the received audio signal,

after detecting that the at least one preset trigger word is included in a received audio signal, determine whether the detected trigger word is misrecognized by measuring, by a wake-up word engine executed by the at least one processor, a similarity between the at least one preset trigger word and the received audio signal based on frequency components extracted from the received audio signal,

wherein the wake-up word engine comprises an acoustic model trained with acoustic information for the preset at least one trigger word,

based on determining that the detected trigger word is misrecognized, request an additional input from a user; and

based on the additional input from the user received in response to the request and the received audio signal, control whether to execute a function corresponding to audio recognition or to terminate the audio recognition.

12 . The electronic device of claim 11 , wherein the at least one processor is configured to determine that the at least one preset trigger word included in the received audio signal is misrecognized based on a history of execution of the function corresponding to the audio recognition within a preset first time.

13 . The electronic device of claim 11 , further comprising an audio output and a receiver,

wherein the at least one processor is configured to:

synchronize the received audio signal and a reference audio signal output from the audio output, and

determine that the at least one preset trigger word included in the received audio signal is misrecognized when a similarity between the synchronized audio signal and the synchronized reference audio signal is equal to or greater than a preset first threshold.

14 . The electronic device of claim 11 , wherein the at least one processor is configured to determine that the at least one preset trigger word included in the received audio signal is misrecognized based on the received audio signal comprising at least one input signal in addition to the at least one preset trigger word.

15 . A non-transitory computer-readable medium storing instructions, that when executed by a processor, causes the processor to:

receive an audio signal;

detect that at least one preset trigger word is included in the received audio signal;

after detecting that the at least one preset trigger word is included in the received audio signal, determine whether the detected trigger word is misrecognized by measuring, by a wake-up word engine executed by the processor, a similarity between the at least one preset trigger word and the received audio signal based on frequency components extracted from the received audio signal,

wherein the wake-up word engine comprises an acoustic model trained with acoustic information for the preset at least one trigger word;

based on determining that the detected trigger word is misrecognized, request an additional input from a user; and

based on the additional input from the user received in response to the request and the received audio signal, control whether to execute a function corresponding to audio recognition or to terminate the audio recognition.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2024
From: CHOI, CHANHEE; BOK, CHANSIK; YOON, HYUNDON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 067660/0103 →
Priority Claims (1)
KR 10-2021-0176941 · Dec 10, 2021 · national
Continuity (2)
Continuation PCTKR2022018196 · Nov 17, 2022
Related Publication 20240331696A1 · Oct 3, 2024
References Cited (37)
US 9484029B2 · Jung et al. · 2016 [cited by applicant]
US 9691378B1 · Meyers · 2017 [cited by examiner]
US 10056081B2 · Kunitake et al. · 2018 [cited by applicant]
US 10867600B2 · Gruenstein et al. · 2020 [cited by applicant]
US 11164565B2 · Lee · 2021 [cited by applicant]
US 11417327B2 · Choi · 2022 [cited by applicant]
US 11450315B2 · Kim et al. · 2022 [cited by applicant]
US 20090326944A1 · Yano · 2009 [cited by examiner]
US 20180211661A1 · Kudo et al. · 2018 [cited by applicant]
US 20200043492A1 · Kim et al. · 2020 [cited by applicant]
US 20210335347A1 · Kim et al. · 2021 [cited by applicant]
CN 113160802A · 2021 [cited by examiner]
CN 113727157A · 2021 [cited by applicant]
EP 3404655B1 · 2020 [cited by applicant]
JP 2006251298A · 2006 [cited by applicant]
JP 200851895A · 2008 [cited by applicant]
JP 2017117371A · 2017 [cited by applicant]
JP 6351440B2 · 2018 [cited by applicant]
JP 201932387A · 2019 [cited by applicant]
JP 2019032387A · 2019 [cited by examiner]
KR 1020010083494A · 2001 [cited by applicant]
KR 1020120110392A · 2012 [cited by applicant]
KR 20120110392A · 2012 [cited by examiner]
KR 1020140035164A · 2014 [cited by applicant]
KR 1020160097788A · 2016 [cited by applicant]
KR 1020180084392A · 2018 [cited by applicant]
KR 1020190065199A · 2019 [cited by applicant]
KR 102093851B1 · 2020 [cited by applicant]
KR 1020200063521A · 2020 [cited by applicant]
KR 102112564B1 · 2020 [cited by applicant]
KR 1020200141126A · 2020 [cited by applicant]
KR 102246900B1 · 2021 [cited by applicant]
KR 102281590B1 · 2021 [cited by applicant]
International Search Report (PCT/ISA/210) issued Feb. 22, 2023 by the International Searching Authority in the International Patent Application No. PCT/KR2022/018196. [cited by applicant]
Written Opinion (PCT/ISA/237) issued Feb. 22, 2023 by the International Searching Authority in the International Patent Application No. PCT/KR2022/018196. [cited by applicant]
Joe Wang et al., “An Audio-Based Wakeword-Independent Verification System”, Proceedings of Interspeech 2020, Oct. 2020, 5 pages. [cited by applicant]
Communication issued on Jan. 30, 2026 by the Korean Ministry of Intellectual Property (MOIP) in Korean Patent Application No. 10-2021-0176941. [cited by applicant]