IP Library › Granted Patent US 12,567,414
Granted Patent B2
US 12,567,414 · App. 18/492,177 · Granted Mar 3, 2026

System and method for detecting a wakeup command for a voice assistant

Inventor: Ranjan Kumar Samal (Bengaluru, IN)
Assignee: Samsung Electronics Co., Ltd.
G10L15/22G10L15/063G10L15/08G10L2015/088G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,567,414
App. No.
18/492,177
Granted
Mar 3, 2026
Kind
B2
Abstract

A method for detecting a wakeup command for a voice assistant is provided. The method includes receiving an audio signal from one or more sources and determining at least one of acoustic parameters or an environmental context of the user. Further, the method includes generating an embedding vector representation associated with the received audio signal and comparing the generated embedding vector representation with one or more prestored embedding vector representations. Furthermore, the method includes detecting the wakeup command in the received audio signal.

Claims (74)

1 . A method for detecting a wakeup command for a voice assistant implemented in a user equipment (UE), the method comprising:

receiving an audio signal from one or more sources, wherein the one or more sources comprises at least one of a user of the UE or one or more environmental elements;

determining at least one of acoustic parameters or an environmental context of the user based on the received audio signal;

generating an embedding vector representation associated with the received audio signal based on the at least one of determined acoustic parameters and determined environmental context by using a machine learning (ML)-based embedding generator model;

comparing the generated embedding vector representation with one or more prestored embedding vector representations; and

detecting the wakeup command in the received audio signal based on determined environmental context and comparison of the generated embedding vector representation with the one or more pre-stored embedding vector representations.

2 . The method of claim 1 ,

wherein the audio signal comprises at least one of a wakeup signal or a noise signal, and

wherein the noise signal is generated by the one or more environmental elements.

3 . The method of claim 1 , wherein the acoustic parameters associated with the user and an environment of the user comprise at least one of a pitch, an intensity, or a magnitude of the received audio signal.

4 . The method of claim 1 ,

wherein determined environmental context includes information associated with one or more objects located in vicinity of the UE and an occurrence of one or more activities in the vicinity of the UE producing the audio signals, and

wherein the information comprises identification information, location information, and status information.

5 . The method of claim 1 , wherein the one or more pre-stored embedding vector representations correspond to vector representations associated with at least one of a combination of a set of true wakeups and one or more environmental contexts of the user, or a combination of a set of false wakeups and the one or more environmental contexts of the user.

6 . The method of claim 5 , further comprising:

sending the one or more prestored embedding vector representations associated with the UE from the UE to one or more other UEs associated with the user via one or more techniques.

7 . The method of claim 1 , further comprising:

determining, via the voice assistant, an occurrence of one of a true wakeup instance or a false wakeup instance for the generated embedding vector representation;

associating the determined one of the true wakeup instance or the false wakeup instance with the generated embedding vector representation, the determined acoustic parameters, and determined environmental context;

storing the generated embedding vector representation in an embedding vector repository upon associating the determined one of the true wakeup instance or the false wakeup instance with the generated embedding vector representation, the determined acoustic parameters, and determined environmental context; and

training an ML-based classification model by using the embedding vector repository upon storing the generated embedding vector representation in the embedding vector repository.

8 . The method of claim 7 , wherein the determining of the occurrence of one of the true wakeup instance or the false wakeup instance comprises:

receiving a set of wake-up commands from the user;

predicting a score for each of the set of wake-up commands;

determining if the predicted score is less than a predefined threshold score;

rejecting the set of wake-up commands upon determining that the predicted score is less than the predefined threshold score;

identifying a pattern associated with the prediction of the score and the rejection of the set of wake-up commands based on the generated score and the predefined threshold score; and

determining a number of attempts by the user for waking-up the voice assistant based on the identified pattern.

9 . The method of claim 8 , wherein the determining of the occurrence of one of the true wakeup instance or the false wakeup instance comprises:

performing a set of acoustic operations on audio signals associated with the received wake-up commands for a predefined time period based on predicted number of attempts, wherein the set of acoustic operations comprise noise suppression, speech boosting, and voice filtering;

monitoring the identified pattern associated with the prediction of the score to evaluate the audio signals upon performing the set of acoustic operations, wherein the audio signals are evaluated in time points at which the number of attempts are predicted;

determining if a wake-up signal is present in the received set of wake-up commands based on a result of evaluation of the audio signals by using a signal detection-based ML model;

generating an embedding vector representation associated with the audio signals upon determining that the wake-up signal is present in the received set of wake-up commands;

associating the true wakeup instance with the generated embedding vector representation; and

storing the generated embedding vector representation in the embedding vector repository upon associating the true wakeup instance with the generated embedding vector representation.

10 . The method of claim 7 , further comprising:

determining that the voice assistant is unable to obtain a natural label for the generated embedding vector representation of the audio signal;

generating a score for the wakeup command in the audio signal based on a number of wakeup attempts, the embedding vector representation, the acoustic parameters, and the environmental context of the user;

associating the generated score with the embedding vector representation, the wakeup command, the acoustic parameters, and the environmental context; and

storing the generated embedding vector representation in the embedding vector repository upon associating the generated score with the embedding vector representation, the wakeup command, the acoustic parameters, and the environmental context.

11 . The method of claim 10 , further comprising:

determining that the voice assistant is unable to determine the occurrence of one of the true wakeup instance or the false wakeup instance for the embedding vector representation of the audio signal;

obtaining an alternate wakeup command from the embedding vector repository based on the embedding vector representation, the acoustic parameters, the environmental context associated with the audio signal and the generated score; and

recommending the obtained alternate wakeup command to the user via one or more modes.

12 . The method of claim 1 , wherein the detecting of the wakeup command in the received audio signal comprises:

classifying the received audio signal into one of a true wakeup instance or a false wakeup instance by using trained ML-based classification model.

13 . A system for detecting a wakeup command for a voice assistant implemented in a user equipment (UE), the system includes one or more processors configured to:

receive an audio signal from one or more sources, wherein the one or more sources comprises at least one of a user of the UE or one or more environmental elements,

determine at least one of acoustic parameters or an environmental context of the user based on the received audio signal,

generate an embedding vector representation associated with the received audio signal based on the at least one of determined acoustic parameters and determined environmental context by using a machine learning (ML)-based embedding generator model,

compare the generated embedding vector representation with one or more prestored embedding vector representations, and

detect the wakeup command in the received audio signal based on determined environmental context and comparison of the generated embedding vector representation with the one or more prestored embedding vector representations.

14 . The system of claim 13 ,

wherein the audio signal comprises at least one of a wakeup signal or a noise signal, and

wherein the noise signal is generated by the one or more environmental elements.

15 . The system of claim 13 , wherein the acoustic parameters associated with the user and an environment of the user comprise at least one of a pitch, an intensity, or a magnitude of the received audio signal.

16 . The system of claim 13 ,

wherein determined environmental context includes information associated with one or more objects located in vicinity of the UE and an occurrence of one or more activities in the vicinity of the UE producing the audio signals, and

wherein the information comprises identification information, location information, and status information.

17 . The system of claim 13 , wherein the one or more pre-stored embedding vector representations correspond to vector representations associated with at least one of a combination of a set of true wakeups and one or more environmental contexts of the user, or a combination of a set of false wakeups and the one or more environmental contexts of the user.

18 . The system of claim 17 , wherein the one or more processors are further configured to:

send the one or more prestored embedding vector representations associated with the UE from the UE to one or more other UEs associated with the user via one or more techniques.

19 . The system of claim 13 , wherein the one or more processors are further configured to:

determine, via the voice assistant, an occurrence of one of a true wakeup instance or a false wakeup instance for the generated embedding vector representation,

associate the determined one of the true wakeup instance or the false wakeup instance with the generated embedding vector representation, the determined acoustic parameters, and determined environmental context,

store the generated embedding vector representation in an embedding vector repository upon associating the determined one of the true wakeup instance or the false wakeup instance with the generated embedding vector representation, the determined acoustic parameters, and determined environmental context, and

train an ML-based classification model by using the embedding vector repository upon storing the generated embedding vector representation in the embedding vector repository.

20 . The system of claim 19 , wherein, in determining the occurrence of one of the true wakeup instance or the false wakeup instance, the one or more processors are further configured to:

receive a set of wake-up commands from the user,

predict a score for each of the set of wake-up commands,

determine if the predicted score is less than a predefined threshold score,

reject the set of wake-up commands upon determining that the predicted score is less than the predefined threshold score,

identify a pattern associated with the prediction of the score and the rejection of the set of wake-up commands based on the generated score and the predefined threshold score, and

determine a number of attempts by the user for waking-up the voice assistant based on the identified pattern.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2023
From: SAMAL, RANJAN KUMAR
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 065311/0336 →
Priority Claims (2)
IN 202241050570 · Sep 5, 2022 · national
IN 2022 41050570 · Aug 10, 2023 · national
Continuity (2)
Continuation PCTKR2023012586 · Aug 24, 2023
Related Publication 20240079007A1 · Mar 7, 2024
References Cited (29)
US 11361763B1 · Maas et al. · 2022 [cited by applicant]
US 11366978B2 · Kim et al. · 2022 [cited by applicant]
US 11694696B2 · Kang et al. · 2023 [cited by applicant]
US 12243511B1 · Joly · 2025 [cited by examiner]
US 12288566B1 · Ganguly · 2025 [cited by examiner]
US 20160267913A1 · Kim et al. · 2016 [cited by applicant]
US 20190311715A1 · Pfeffinger et al. · 2019 [cited by applicant]
US 20200005789A1 · Chae · 2020 [cited by applicant]
US 20200125820A1 · Kim · 2020 [cited by examiner]
US 20200279558A1 · Li et al. · 2020 [cited by applicant]
US 20200312336A1 · Kang · 2020 [cited by examiner]
US 20210043190A1 · Wang · 2021 [cited by examiner]
US 20210043204A1 · Hwang et al. · 2021 [cited by applicant]
US 20210174783A1 · Wieman · 2021 [cited by examiner]
US 20210183367A1 · Sharifi et al. · 2021 [cited by applicant]
US 20210217431A1 · Pearson · 2021 [cited by examiner]
US 20210272573A1 · Yousefi · 2021 [cited by examiner]
US 20210326421A1 · Khoury · 2021 [cited by examiner]
US 20220366914A1 · Moreno · 2022 [cited by examiner]
US 20220399007A1 · Jose · 2022 [cited by examiner]
US 20220406324A1 · Shin · 2022 [cited by examiner]
US 20230038982A1 · Narayanan · 2023 [cited by examiner]
CN 106782536A · 2017 [cited by applicant]
KR 1020200045647A · 2020 [cited by applicant]
KR 1020200119377A · 2020 [cited by applicant]
WO 2018086033A1 · 2018 [cited by applicant]
Indian Office Action dated Jun. 23, 2025, issued in an India Patent Application No. 202241050570. [cited by applicant]
Extended European Search Report dated Jul. 15, 2025, issued in a European Patent Application No. 23863391.1. [cited by applicant]
International Search Report dated Dec. 7, 2023, issued in International Application No. PCT/KR2023/012586. [cited by applicant]