IP Library Granted Patent US 12,688,851
Granted Patent B2
US 12,688,851 · App. 18/605,404 · Granted Jul 21, 2026

Voice based activation detection

Inventors: Cheng Yun Yu (Beijing, CN); Yuan Jie Zhang (Ningbo, CN); Zhi Qiang Kou (Beijing, CN); Gui Lei Yang (Beijing, CN); Han Xu Zheng (Beijing, CN)
Assignee: International Business Machines Corporation
G10L15/22G10L15/063G10L15/08G10L25/63G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,851
App. No.
18/605,404
Granted
Jul 21, 2026
Kind
B2
Abstract

A computer-implemented method, according to one approach, includes: determining whether initial utterances received from a user include predetermined wake-up words. In response to determining that the initial utterances include the predetermined wake-up words, a voice assistant initiates an interaction with the user. In response to determining that the initial utterances in the audio signal stop for a predetermined period following the initial utterances that include the wake-up words, determining whether additional utterances are received from the user. In response to receiving additional utterances from the user, determining whether the additional utterances have an agitated tone. In situations where the additional utterances do not have an agitated tone, the additional utterances are processed. Furthermore, in response to determining that causing the additional utterances to be processed results in action being taken, the initial utterances are labeled as genuine dialog initiation entries.

Claims (68)

1 . A computer-implemented method (CIM), comprising:

in response to receiving an audio signal having initial utterances from a user, determining whether the initial utterances include predetermined wake-up words by:

using a trained artificial intelligence (AI) based model to evaluate a tone of the initial utterances;

in response to determining that the initial utterances include the predetermined wake-up words, causing a voice assistant to initiate an interaction with the user;

in response to receiving additional utterances from the user that do not have an agitated tone, causing the additional utterances to be processed;

in response to determining that causing the additional utterances to be processed results in action being taken, labeling the initial utterances that include the wake-up words as genuine dialog initiation entries;

in response receiving additional utterances from the user that do have an agitated tone, labeling the initial utterances that include the wake-up words as unintentional dialog initiation entries;

causing the additional utterances to be ignored; and

using the labeled unintentional dialog initiation entries to retrain the AI based model.

2 . The CIM of claim 1 , further comprising:

using the labeled genuine dialog initiation entries to retrain the AI based model.

3 . The CIM of claim 1 , further comprising:

in response to determining that the initial utterances in the audio signal do not stop for a predetermined period following the initial utterances that include the wake-up words, labeling the initial utterances that include the wake-up words as unintentional dialog initiation entries; and

using the labeled unintentional dialog initiation entries to retrain the AI based model.

4 . The CIM of claim 1 , further comprising:

in response to determining that additional utterances are not received from the user following a predetermined period, labeling the initial utterances that include the wake-up words as unintentional dialog initiation entries; and

using the labeled unintentional dialog initiation entries to retrain the AI based model.

5 . The CIM of claim 1 , further comprising:

in response to determining that causing the additional utterances to be processed results in no action being taken, labeling the initial utterances that include the wake-up words as unintentional dialog initiation entries; and

using the labeled unintentional dialog initiation entries to retrain the AI based model.

6 . The CIM of claim 1 , wherein the trained AI based model is a trained machine learning model.

7 . The CIM of claim 1 , further comprising:

receiving a mixed audio signal having initial utterances from multiple users;

extracting individual audio signals from the mixed audio signal, each of the extracted individual audio signals having ones of the initial utterances that correspond to a respective one of the multiple users; and

for each of the individual audio signals, determining whether the ones of the initial utterances in a given individual audio signal include the predetermined wake-up words.

8 . A computer program product (CPP), comprising:

a set of one or more computer-readable storage media; and

program instructions, collectively stored in the set of one or more storage media, for causing a processor set to perform the following computer operations:

in response to receiving an audio signal having initial utterances from a user, determining whether the initial utterances include predetermined wake-up words by:

using a trained artificial intelligence (AI) based model to evaluate a tone of the initial utterances;

in response to determining that the initial utterances include the predetermined wake-up words, cause a voice assistant to initiate an interaction with the user;

in response to receiving additional utterances from the user that do not have an agitated tone, cause the additional utterances to be processed;

in response to determining that causing the additional utterances to be processed results in action being taken, label the initial utterances that include the wake-up words as genuine dialog initiation entries;

in response to receiving additional utterances from the user that have an agitated tone, label the initial utterances that include the wake-up words as unintentional dialog initiation entries;

cause the additional utterances to be ignored; and

use the labeled unintentional dialog initiation entries to retrain the AI based model.

9 . The CPP of claim 8 , wherein the program instructions are for causing the processor set to further perform the following computer operations:

use the labeled genuine dialog initiation entries to retrain the AI based model.

10 . The CPP of claim 8 , wherein the program instructions are for causing the processor set to further perform the following computer operations:

in response to determining that the initial utterances in the audio signal do not stop for a predetermined period following the initial utterances that include the wake-up words, label the initial utterances that include the wake-up words as unintentional dialog initiation entries; and

use the labeled unintentional dialog initiation entries to retrain the AI based model.

11 . The CPP of claim 8 , wherein the program instructions are for causing the processor set to further perform the following computer operations:

in response to determining that additional utterances are not received from the user following a predetermined period, label the initial utterances that include the wake-up words as unintentional dialog initiation entries; and

use the labeled unintentional dialog initiation entries to retrain the AI based model.

12 . The CPP of claim 8 , wherein the program instructions are for causing the processor set to further perform the following computer operations:

in response to determining that causing the additional utterances to be processed results in no action being taken, label the initial utterances that include the wake-up words as unintentional dialog initiation entries; and

use the labeled unintentional dialog initiation entries to retrain the AI based model.

13 . The CPP of claim 8 , wherein the trained AI based model is a trained machine learning model.

14 . The CPP of claim 8 , wherein the program instructions are for causing the processor set to further perform the following computer operations:

receive a mixed audio signal having initial utterances from multiple users;

extract individual audio signals from the mixed audio signal, each of the extracted individual audio signals having ones of the initial utterances that correspond to a respective one of the multiple users; and

for each of the individual audio signals, determining whether the ones of the initial utterances in a given individual audio signal include the predetermined wake-up words.

15 . A computer system (CS), comprising:

a processor set;

a set of one or more computer-readable storage media;

program instructions, collectively stored in the set of one or more storage media, for causing the processor set to perform the following computer operations:

in response to receiving an audio signal having initial utterances from a user, determining whether the initial utterances include predetermined wake-up words by:

using a trained artificial intelligence (AI) based model to evaluate a tone of the initial utterances;

in response to determining that the initial utterances include the predetermined wake-up words, cause a voice assistant to initiate an interaction with the user;

in response to receiving additional utterances from the user that do not have an agitated tone, cause the additional utterances to be processed;

in response to determining that causing the additional utterances to be processed results in action being taken, label the initial utterances that include the wake-up words as genuine dialog initiation entries;

in response to receiving additional utterances from the user that have an agitated tone, label the initial utterances that include the wake-up words as unintentional dialog initiation entries;

cause the additional utterances to be ignored; and

use the labeled unintentional dialog initiation entries to retrain the AI based model.

16 . The CS of claim 15 ,

wherein the program instructions are for causing the processor set to further perform the following computer operations:

in response to determining that: the initial utterances in the audio signal do not stop for a predetermined period following the initial utterances that include the wake-up words and additional utterances are not received from the user following the predetermined period, label the initial utterances that include the wake-up words as unintentional dialog initiation entries; and

use the labeled unintentional dialog initiation entries to retrain the AI based model.