IP Library Granted Patent US 12,437,753
Granted Patent B2
US 12,437,753 · App. 17/698,368 · Granted Oct 7, 2025

Method and apparatus with keyword detection

Inventors: Bo Wei (Xi'an, CN); Meirong Yang (Xi'an, CN); Tao Zhang (Xi'an, CN); Xiao Tang (Xi'an, CN); Xing Huang (Xi'an, CN)
Assignee: Samsung Electronics Co., Ltd.
G10L15/08G10L15/02G10L25/30G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,753
App. No.
17/698,368
Filed
Mar 18, 2022
Granted
Oct 7, 2025
Kind
B2
Examiner
HANG, VU B
Art Unit
2654
USPC
704/200
Abstract

A processor-implemented method with keyword detection includes: receiving at least a portion of a voice signal input by a user; extracting a voice feature of the voice signal; inputting an abstract representation sequence of a preset keyword and the voice feature to an end-to-end keyword detection model, and determining a result on whether the preset keyword is present in the voice signal output from the keyword detection model, wherein the keyword detection model predicts whether the preset keyword is present in the voice signal by: determining an abstract representation sequence of the voice signal, based on the voice feature and the abstract representation sequence of the preset keyword; predicting position information of the preset keyword in the voice signal based on the abstract representation sequence of the voice signal; and predicting whether the preset keyword is present in the voice signal, based on the abstract representation sequence of the voice signal and the position information.

Claims (54)

1. A processor-implemented method with keyword detection, the method comprising:

receiving at least a portion of a voice signal input by a user;

extracting a voice feature of the voice signal;

inputting an abstract representation sequence of a preset keyword and the voice feature to an end-to-end keyword detection model; and

determining a result on whether the preset keyword is present in the voice signal output from the keyword detection model,

wherein the keyword detection model predicts whether the preset keyword is present in the voice signal by:

determining an abstract representation sequence of the voice signal, based on the voice feature and the abstract representation sequence of the preset keyword;

predicting position information of the preset keyword in the voice signal based on the abstract representation sequence of the voice signal; and

predicting whether the preset keyword is present in the voice signal, based on the abstract representation sequence of the voice signal and the position information.

2. The method of claim 1 , wherein the keyword detection method is performed in real-time for at least a portion of the voice signal input by the user.

3. The method of claim 1 , wherein the preset keyword comprises either one or both of a keyword defined by the user and a preset keyword preset by a system or an application.

4. The method of claim 1 , wherein the determining of the abstract representation sequence of the voice signal, based on the voice feature and the abstract representation sequence of the preset keyword comprises determining the abstract representation sequence of the voice signal by combining the voice feature with the abstract representation sequence of the preset keyword through an attention mechanism.

5. The method of claim 1 , wherein the predicting of whether the preset keyword is present in the voice signal, based on the abstract representation sequence of the voice signal and the position information comprises:

determining an abstract representation sequence of a portion comprising the preset keyword in the voice signal, based on the abstract representation sequence of the voice signal and the position information; and

predicting whether the preset keyword is present in the voice signal by combining the abstract representation sequence of the preset keyword with the abstract representation sequence of the portion comprising the preset keyword in the voice signal through an attention mechanism.

6. The method of claim 4 , wherein the keyword detection model comprises a voice encoder used for predicting an abstract representation sequence of a voice signal, the voice encoder comprises a plurality of submodules connected in series, and each of the submodules inputs the abstract representation sequence of the preset keyword to a hidden layer abstract representation sequence of the voice signal through the attention mechanism.

7. The method of claim 1 , wherein the abstract representation sequence of the preset keyword is generated by a pre-trained keyword encoder based on a phone sequence of the preset keyword.

8. The method of claim 1 , wherein the keyword detection model is determined by multi-objective joint training, and a multi-objective comprises predicting a phone sequence corresponding to the voice signal, predicting a position of a keyword in the voice signal, and predicting whether the keyword is present in the voice signal.

9. The method of claim 8 , wherein an objective loss function corresponding to an object for predicting the position of the keyword in the voice signal is based on a diagonal pattern of an attention matrix.

10. The method of claim 1 , further comprising either one or both of:

waking up a current electronic device in response to the result output from the keyword detection model indicating that the preset keyword is present in the voice signal; and

outputting the result and the position information.

11. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, configure the one or more processors to perform the method of claim 1 .

12. An apparatus with keyword detection, the apparatus comprising:

one or more processors; and

a memory storing instructions that, when executed by the one or more processors, configure the one or more processors to perform the method of claim 1 .

13. An apparatus with keyword detection, the apparatus comprising:

a receiver configured to receive at least a portion of a voice signal input by a user;

one or more processors configured to:

extract a voice feature of the voice signal; and

input an abstract representation sequence of a preset keyword and the voice feature to an end-to-end keyword detection model; and

determine a result on whether the preset keyword is present in the voice signal output from the keyword detection model,

wherein the keyword detection model predicts whether the preset keyword is present in the voice signal by:

determining an abstract representation sequence of the voice signal, based on the voice feature and the abstract representation sequence of the preset keyword;

predicting position information of the preset keyword in the voice signal based on the abstract representation sequence of the voice signal; and

predicting whether the preset keyword is present in the voice signal, based on the abstract representation sequence of the voice signal and the position information.

14. The apparatus of claim 11 , wherein the keyword detection apparatus processes at least a portion of the voice signal input by the user in real-time.

15. The apparatus of claim 11 , wherein, for the determining of the abstract representation sequence of the voice signal, the one or more processors are configured to determine the abstract representation sequence of the voice signal by combining the voice feature with the abstract representation sequence of the preset keyword through an attention mechanism.

16. The apparatus of claim 11 , wherein the predicting of whether the preset keyword is present in the voice signal based on the abstract representation sequence of the voice signal and the position information comprises:

determining an abstract representation sequence of a portion comprising the preset keyword in the voice signal based on the abstract representation sequence of the voice signal and the position information; and

predicting whether the preset keyword is present in the voice signal by combining the abstract representation sequence of the preset keyword with the abstract representation sequence of the portion comprising the preset keyword in the voice signal through an attention mechanism.

17. A processor-implemented method with keyword detection, the method comprising:

determining text of a keyword;

determining a phone sequence of the text;

determining whether the keyword satisfies a preset condition based on either one or both of the text and the phone sequence;

inputting, in response to the keyword satisfying the preset condition, the phone sequence to a pre-trained keyword encoder;

generating an abstract representation sequence of the keyword from the pre-trained keyword encoder,

wherein the determining of whether the keyword satisfies the preset condition comprises determining whether a number of syllables of the text is greater than or equal to a predetermined number.

18. The method of claim 17 , further comprising:

receiving at least a portion of a voice signal input by a user;

extracting a voice feature of the voice signal; and

inputting the abstract representation sequence of the keyword and the voice feature to an end-to-end keyword detection model; and

determining a result on whether the keyword is present in the voice signal output from the keyword detection model.

19. The method of claim 17 , wherein the text of the keyword is input by a user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2025
From: WEI, BO; YANG, MEIRONG; ZHANG, TAO; TANG, XIAO; HUANG, XING
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 071565/0659 →
Priority Claims (2)
CN 202110291276.X · Mar 18, 2021 · national
KR 10-2021-0182848 · Dec 20, 2021 · national
Continuity (1)
Related Publication 20220301550A1 · Sep 22, 2022
References Cited (27)
US 9672817B2 · Yong · 2017 [cited by examiner]
US 10257314B2 · Agrawal et al. · 2019 [cited by applicant]
US 10984783B2 · Chen · 2021 [cited by examiner]
US 20190318727A1 · Lopez Moreno et al. · 2019 [cited by applicant]
US 20200126556A1 · Mosayyebpour et al. · 2020 [cited by applicant]
US 20200410983A1 · Mohajer et al. · 2020 [cited by applicant]
US 20210056961A1 · Ding et al. · 2021 [cited by applicant]
CN 105679316A · 2016 [cited by applicant]
CN 106782536A · 2017 [cited by applicant]
CN 107665705A · 2018 [cited by applicant]
CN 109065032A · 2018 [cited by applicant]
CN 109147766A · 2019 [cited by applicant]
CN 109545190A · 2019 [cited by applicant]
CN 110119765A · 2019 [cited by applicant]
CN 110288980A · 2019 [cited by applicant]
CN 110334244A · 2019 [cited by applicant]
CN 110767223A · 2020 [cited by applicant]
CN 110827806A · 2020 [cited by applicant]
CN 111009235A · 2020 [cited by applicant]
CN 111144127A · 2020 [cited by applicant]
CN 111508493A · 2020 [cited by applicant]
CN 111933129A · 2020 [cited by applicant]
CN 112151015A · 2020 [cited by applicant]
CN 112309398A · 2021 [cited by applicant]
KR 1020200017139A · 2020 [cited by applicant]
Chinese Office Action issued on Sep. 27, 2023, in counterpart Chinese Patent Application No. 202110291276.X (3 pages in English, 9 pages in Chinese). [cited by applicant]
Zhao, Zeyu, et al. “End-to-End Keyword Search Based on Attention and Energy Scorer for Low Resource Languages.” [cited by applicant]