IP Library Granted Patent US 12,738,262
Granted Patent B2
US 12,738,262 · App. 18/253,673 · Granted Sep 15, 2026

Enabling training of a machine-learning model for trigger-word detection

Inventor: Tanzia Haque Tanzi (Huddinge, SE)
Assignee: ASSA ABLOY AB
G10L15/063G10L15/08G10L15/1815G10L2015/088G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,262
App. No.
18/253,673
Granted
Sep 15, 2026
Kind
B2
Abstract

It is provided a method for enabling training a machine-learning, ML, model for trigger-word detection, the method being performed in a training data provider ( 1 ). The method comprises: receiving ( 40 ) sound-based data, the sound-based data being based on sounds captured in a space to be monitored; determining ( 42 ) that the sound-based data corresponds to a trigger word, and labelling this sound-based data to correspond to the trigger word; and providing ( 44 ) the labelled sound-based data to train the ML model.

Claims (50)

1 . A method for enabling training a machine-learning, ML, model for trigger-word detection, the method being performed in a training data provider, the method comprising:

receiving sound-based data, the sound-based data being based on sounds captured in a space to be monitored;

determining that the sound-based data corresponds to a trigger word, the determining comprising:

performing speech recognition of the sound-based data;

finding a section of sound-based data that, using the speech recognition, fails to be considered to be the trigger word, but is close to being considered to be the trigger word;

obtaining semantic vector data based on the found section of sound-based data; and

determining that the found section of sound-based data corresponds to the trigger word when a distance, in vector space, between the semantic vector data of the sound-based data and a vector corresponding to the trigger word, is less than a threshold distance;

labelling the sound-based data, including the found section, to correspond to the trigger word; and

providing the labelled sound-based data to train the ML model.

2 . The method according to claim 1 , wherein the sound-based data is in the form of mel-frequency cepstral coefficients, MFCCs.

3 . The method according to claim 1 , further comprising:

discarding sections of the sound-based data that fail to correspond to voice sounds.

4 . The method according to claim 1 , further comprising, after the providing the labelled sound-based data:

discarding all of the sound-based data.

5 . The method according to claim 1 , further comprising:

training a local ML model; and

transmitting at least part of the local ML model to a central location for aggregated learning of a central ML model.

6 . The method according to claim 5 , further comprising:

receiving an updated ML model being based on the central ML model.

7 . The method according to claim 5 , wherein the ML model is the local ML model.

8 . A training data provider for enabling training a machine-learning, ML, model for trigger-word detection, the training data provider comprising:

a processor; and

a memory storing instructions that, when executed by the processor, cause the training data provider to:

receive sound-based data, the sound-based data being based on sounds captured in a space to be monitored;

determine that the sound-based data corresponds to a trigger word, which comprises instructions that, when executed by the processor, cause the training data provider to:

perform speech recognition of the sound-based data;

find a section of sound-based data that, using the speech recognition, fails to be considered to be the trigger word, but is close to being considered to be the trigger word;

obtain semantic vector data based on the found section of sound-based data; and

determine that the found section of sound-based data corresponds to the trigger word when a distance, in vector space, between the semantic vector data of the sound-based data and a vector corresponding to the trigger word, is less than a threshold distance;

label the sound-based data, including the found section, to correspond to the trigger word; and

provide the labelled sound-based data to train the ML model.

9 . The training data provider according to claim 8 , wherein the sound-based data is in the form of mel-frequency cepstral coefficients, MFCCs.

10 . The training data provider according to claim 8 , further comprising instructions that, when executed by the processor, cause the training data provider to:

discard sections of the sound-based data that fail to correspond to voice sounds.

11 . The training data provider according to claim 8 , further comprising instructions that, when executed by the processor, cause the training data provider to, after the instructions to provide the labelled sound-based data, discard all of the sound-based data.

12 . The training data provider according to claim 8 , further comprising instructions that, when executed by the processor, cause the training data provider to:

train a local ML model; and

transmit at least part of the local ML model to a central location for aggregated learning of a central ML model.

13 . The training data provider according to claim 12 , further comprising instructions that, when executed by the processor, cause the training data provider to:

receive an updated ML model being based on the central ML model.

14 . The training data provider according to claim 12 , wherein the ML model is the local ML model.

15 . A non-transitory computer readable storage medium storing a computer program for enabling training a machine-learning, ML, model for trigger-word detection, the computer program comprising computer program code which, when executed on a training data provider, causes the training data provider to:

receive sound-based data, the sound-based data being based on sounds captured in a space to be monitored;

determine that the sound-based data corresponds to a trigger word, wherein the computer program code to determine that the sound-based data corresponds to the trigger word comprises instructions that, when executed by the processor, cause the training data provider to:

perform speech recognition of the sound-based data;

find a section of sound-based data that, using the speech recognition, fails to be considered to be the trigger word, but is close to being considered to be the trigger word;

obtain semantic vector data based on the found section of sound-based data; and

determine that the found section of sound-based data corresponds to the trigger word when a distance, in vector space, between the semantic vector data of the sound-based data and a vector corresponding to the trigger word, is less than a threshold distance;

label the sound-based data, including the found section, to correspond to the trigger word; and

provide the labelled sound-based data to train the ML model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2023
From: TANZI, TANZIA HAQUE
To: ASSA ABLOY AB
Reel/Frame 065047/0866 →
Priority Claims (1)
SE 2051362-8 · Nov 23, 2020 · national
Continuity (1)
Related Publication 20240005909A1 · Jan 4, 2024
References Cited (33)
US 8700392B1 · Hart · 2014 [cited by examiner]
US 10304475B1 · Wang · 2019 [cited by examiner]
US 10388272B1 · Thomson · 2019 [cited by examiner]
US 10468026B1 · Newman · 2019 [cited by examiner]
US 11322138B2 · Li · 2022 [cited by examiner]
US 11355102B1 · Mishchenko · 2022 [cited by examiner]
US 11615784B2 · Gao · 2023 [cited by examiner]
US 11669687B1 · Joshi · 2023 [cited by examiner]
US 20140207457A1 · Biatov · 2014 [cited by examiner]
US 20160189706A1 · Zopf et al. · 2016 [cited by applicant]
US 20160267913A1 · Kim · 2016 [cited by examiner]
US 20170178625A1 · Mamou · 2017 [cited by examiner]
US 20180096681A1 · Ni · 2018 [cited by examiner]
US 20190051299A1 · Ossowski · 2019 [cited by examiner]
US 20190057688A1 · Black · 2019 [cited by applicant]
US 20200143802A1 · Newell · 2020 [cited by applicant]
US 20200349924A1 · Stoimenov · 2020 [cited by examiner]
US 20210158803A1 · Knudson · 2021 [cited by examiner]
US 20210334665A1 · Wang · 2021 [cited by examiner]
US 20220115011A1 · Sharifi · 2022 [cited by examiner]
US 20220165277A1 · Kracun · 2022 [cited by examiner]
US 20230195801A1 · Yang · 2023 [cited by examiner]
US 20240005926A1 · Wood · 2024 [cited by examiner]
CN 107918633 · 2018 [cited by applicant]
CN 108320733A · 2018 [cited by examiner]
CN 110706695 · 2020 [cited by applicant]
WO 2019097276 · 2019 [cited by applicant]
Kepuska, et al. “Wake-up-word speech recognition application for first responder communication enhancement.” Sensors, and Command, Control, Communications, and Intelligence (C3I) Technologies for Homeland Security and H… [cited by examiner]
“Swedish Application Serial No. 20513262-8, Search Report mailed Jun. 17, 2021”, 7 pgs. [cited by applicant]
“International Application Serial No. PCT EP2021 082484, International Search Report mailed Mar. 1, 2022”, 6 pgs. [cited by applicant]
“Serial No. PCT EP2021 082484, Written Opinion mailed Mar. 1, 2022”, 9 pgs. [cited by applicant]
“Swedish Application Serial No. 20513262-8, Office Action mailed Oct. 7, 2022”, 4 pgs. [cited by applicant]
Chung, Yu-An, “Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech”, arxiv.org, Cornell University Library, 201OLIN Library Cornell University Ithaca, NY14853, (2018), 5 pgs. [cited by applicant]