IP Library › Granted Patent US 12,738,271
Granted Patent B2
US 12,738,271 · App. 18/474,968 · Granted Sep 15, 2026

Voice wakeup method and apparatus, storage medium, and system

Inventors: Longshuai Xiao (Beijing, CN); Yinan Zhen (Beijing, CN); Wenjie Li (Beijing, CN); Chao Peng (Beijing, CN); Zhanlei Yang (Beijing, CN)
Assignee: Huawei Technologies Co., Ltd.
G10L15/22G10L15/02G10L15/063G10L15/16G10L2015/025G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,271
App. No.
18/474,968
Filed
Sep 26, 2023
Granted
Sep 15, 2026
Kind
B2
Examiner
HANG, VU B
Art Unit
2654
USPC
704/232
Abstract

A method for voice wakeup is provided, which includes: obtaining original first microphone data; performing first-stage processing based on the first microphone data to obtain first wakeup data, wherein the first-stage processing includes first-stage separation processing and first-stage wakeup processing that are based on a neural network model; performing second-stage processing based on the first microphone data to obtain second wakeup data when the first wakeup data indicates that pre-wakeup succeeds, wherein the second-stage processing includes second-stage separation processing and second-stage wakeup processing that are based on the neural network model; and determining a wakeup result based on the second wakeup data. According to this application, a two-stage separation and wakeup solution is designed, e.g., pre-wakeup determining is performed by using a first-stage separation and wakeup solution, and wakeup determining is performed again after pre-wakeup succeeds, to reduce a false wakeup rate while ensuring a high wakeup rate.

Claims (70)

1 . A voice wakeup method applied to an electronic device, wherein the method comprises:

obtaining first microphone data;

performing first-stage processing based on the first microphone data to obtain first wakeup data, wherein the first-stage processing comprises first-stage separation processing and first-stage wakeup processing that are based on a neural network model;

wherein the first wakeup data comprises a first confidence level, and the first confidence level indicates a probability that the first microphone data comprises a preset wakeup keyword which comprises N wakeup keywords, wherein N is a natural number;

wherein the first confidence level is an (N+1) dimensional vector in which N dimensions correspond to the N wakeup keywords; and based on the first confidence level being greater than a first threshold, the first wakeup data indicates that a pre-wakeup succeeds;

performing second-stage processing according to the first microphone data to obtain second wakeup data, based on the first wakeup data indicating that the pre-wakeup succeeds, wherein the second-stage processing comprises second-stage separation processing and second-stage wakeup processing that are based on the neural network model; and

determining a wakeup result based on the second wakeup data.

2 . The method according to claim 1 , wherein the performing the first-stage processing based on the first microphone data to obtain the first wakeup data comprises:

preprocessing the first microphone data to obtain multi-channel feature data;

invoking, based on the multi-channel feature data, a first-stage separation module that completes training in advance to output first separation data, wherein the first-stage separation module is configured to perform first-stage separation processing; and

invoking, based on the multi-channel feature data and the first separation data, a first-stage wakeup module that completes training in advance to output the first wakeup data, wherein the first-stage wakeup module is configured to perform first-stage wakeup processing.

3 . The method according to claim 2 , wherein the performing the second-stage processing based on the first microphone data to obtain the second wakeup data based on the first wakeup data indicating that the pre-wakeup succeeds comprises:

based on the first wakeup data indicating that the pre-wakeup succeeds, invoking, according to the multi-channel feature data and the first separation data, a second-stage separation module that completes training in advance to output second separation data, wherein the second-stage separation module is configured to perform second-stage separation processing; and

invoking, based on the multi-channel feature data, the first separation data, and the second separation data, a second-stage wakeup module that completes training in advance to output the second wakeup data, wherein the second-stage wakeup module is configured to perform second-stage wakeup processing.

4 . The method according to claim 2 , wherein the first-stage separation module comprises a first-stage multi-feature fusion model and a first-stage separation model, and the invoking, based on the multi-channel feature data, the first-stage separation module that completes training in advance to output the first separation data comprises:

inputting the multi-channel feature data into the first-stage multi-feature fusion model to output first single-channel feature data; and

inputting the first single-channel feature data into the first-stage separation model to output the first separation data.

5 . The method according to claim 2 , wherein the first-stage wakeup module comprises a first wakeup model in a multiple-input single-output form, and the invoking, based on the multi-channel feature data and the first separation data, the first-stage wakeup module that completes training in advance to output the first wakeup data comprises:

inputting the multi-channel feature data and the first separation data into a first-stage wakeup model to output the first wakeup data.

6 . The method according to claim 2 , wherein the first-stage wakeup module comprises a first wakeup model in a multiple-input multiple-output form and a first post-processing module, and the invoking, based on the multi-channel feature data and the first separation data, the first-stage wakeup module that completes training in advance to output the first wakeup data comprises:

inputting the multi-channel feature data and the first separation data into the first wakeup model to output phoneme sequence information respectively corresponding to a plurality of pieces of sound source data; and

inputting the phoneme sequence information respectively corresponding to the plurality of pieces of sound source data into the first post-processing module to output the first wakeup data, wherein the first wakeup data comprises second confidence levels respectively corresponding to the plurality of pieces of sound source data, and the second confidence levels indicate an acoustic feature similarity between the sound source data and a preset wakeup keyword.

7 . A voice wakeup apparatus, wherein the apparatus comprises:

a processor; and

a memory, configured to store instructions to be executed by the processor, wherein

upon executing the instructions, the processor is configured to implement the following:

obtaining first microphone data;

performing first-stage processing based on the first microphone data to obtain first wakeup data, wherein the first-stage processing comprises first-stage separation processing and first-stage wakeup processing that are based on a neural network model;

wherein the first wakeup data comprises a first confidence level, and the first confidence level indicates a probability that the first microphone data comprises a preset wakeup keyword which comprises N wakeup keywords, wherein N is a natural number;

wherein the first confidence level is an (N+1) dimensional vector in which N dimensions correspond to the N wakeup keywords; and based on the first confidence level being greater than a first threshold, the first wakeup data indicates that a pre-wakeup succeeds;

performing second-stage processing according to the first microphone data to obtain second wakeup data, based on the first wakeup data indicating that the pre-wakeup succeeds, wherein the second-stage processing comprises second-stage separation processing and second-stage wakeup processing that are based on the neural network model; and

determining a wakeup result based on the second wakeup data.

8 . The apparatus according to claim 7 , wherein upon executing the instructions, the processor is further configured to implement the following:

preprocessing the first microphone data to obtain multi-channel feature data;

performing first-stage separation processing based on the multi-channel feature data to output first separation data; and

performing first-stage wakeup processing based on the multi-channel feature data and the first separation data to output the first wakeup data.

9 . The apparatus according to claim 8 , wherein upon executing the instructions, the processor is further configured to implement the following:

based on the first wakeup data indicating that the pre-wakeup succeeds, performing second-stage separation processing based on the multi-channel feature data and the first separation data to output second separation data; and

performing second-stage wakeup processing based on the multi-channel feature data, the first separation data, and the second separation data to output the second wakeup data.

10 . The apparatus according to claim 9 , wherein

the first-stage separation processing is streaming sound source separation processing, and the first-stage wakeup processing is streaming sound source wakeup processing; and/or

the second-stage separation processing is offline sound source separation processing, and the second-stage wakeup processing is offline sound source wakeup processing.

11 . The apparatus according to claim 9 , further comprising:

at least one wakeup model in a multiple-input single-output form or a multiple-input multiple-output form.

12 . The apparatus according to claim 9 , wherein a dual-path conformer network structure is used.

13 . The apparatus according to claim 8 , further comprising: a first-stage multi-feature fusion model and a first-stage separation model, wherein the first-stage separation model is further configured to:

input the multi-channel feature data into the first-stage multi-feature fusion model to output first single-channel feature data; and

input the first single-channel feature data into the first-stage separation model to output the first separation data.

14 . The apparatus according to claim 8 , wherein a first-stage wakeup module comprises a first wakeup model in a multiple-input single-output form, wherein upon executing the instructions, the processor is further configured to implement the following:

inputting the multi-channel feature data and the first separation data into the first-stage wakeup module to output the first wakeup data.

15 . The apparatus according to claim 8 , further comprising: a first wakeup model in a multiple-input multiple-output form and a first post-processing device, wherein upon executing the instructions, the processor is further configured to implement the following:

inputting the multi-channel feature data and the first separation data into the first wakeup model to output phoneme sequence information respectively corresponding to a plurality of pieces of sound source data; and

inputting the phoneme sequence information respectively corresponding to the plurality of pieces of sound source data into the first post-processing device to output the first wakeup data, wherein the first wakeup data comprises second confidence levels respectively corresponding to the plurality of pieces of sound source data, and the second confidence level indicates an acoustic feature similarity between the sound source data and a preset wakeup keyword.

16 . The apparatus according to claim 9 , further comprising a second wakeup model in a multiple-input single-output form, wherein upon executing the instructions, the processor is further configured to implement the following:

inputting the multi-channel feature data, the first separation data, and the second separation data into a second-stage wakeup model to output the second wakeup data, wherein the second wakeup data comprises a third confidence level, and the third confidence level indicates a probability that the first microphone data comprises the preset wakeup keyword.

17 . The apparatus according to claim 9 , further comprising a second wakeup model in a multiple-input multiple-output form and a second post-processing device, wherein upon executing the instructions, the processor is further configured to implement the following:

inputting the multi-channel feature data, the first separation data, and the second separation data into a second-stage wakeup model to output phoneme sequence information respectively corresponding to a plurality of pieces of sound source data; and

inputting the phoneme sequence information respectively corresponding to the plurality of pieces of sound source data into the second post-processing device to output the second wakeup data, wherein the second wakeup data comprises fourth confidence levels respectively corresponding to the plurality of pieces of sound source data, and the fourth confidence level indicates an acoustic feature similarity between the sound source data and the preset wakeup keyword.

18 . A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer program instructions, and upon the computer program instructions being executed by a processor, the processor is configured to implement the following:

obtaining first microphone data;

performing first-stage processing based on the first microphone data to obtain first wakeup data, wherein the first-stage processing comprises first-stage separation processing and first-stage wakeup processing that are based on a neural network model;

wherein the first wakeup data comprises a first confidence level, and the first confidence level indicates a probability that the first microphone data comprises a preset wakeup keyword which comprises N wakeup keywords, wherein N is a natural number;

wherein the first confidence level is an (N+1) dimensional vector in which N dimensions correspond to the N wakeup keywords; and based on the first confidence level being greater than a first threshold, the first wakeup data indicates that a pre-wakeup succeeds;

performing second-stage processing according to the first microphone data to obtain second wakeup data, based on the first wakeup data indicating that the pre-wakeup succeeds, wherein the second-stage processing comprises second-stage separation processing and second-stage wakeup processing that are based on the neural network model; and

determining a wakeup result based on the second wakeup data.

19 . A voice wakeup system, wherein the voice wakeup system is configured to perform the method according to claim 1 .

20 . The non-transitory computer-readable storage medium according to claim 18 , wherein the performing the first-stage processing based on the first microphone data to obtain the first wakeup data comprises:

preprocessing the first microphone data to obtain multi-channel feature data;

invoking, based on the multi-channel feature data, a first-stage separation module that completes training in advance to output first separation data, wherein the first-stage separation module is configured to perform first-stage separation processing; and

invoking, based on the multi-channel feature data and the first separation data, a first-stage wakeup module that completes training in advance to output the first wakeup data, wherein the first-stage wakeup module is configured to perform first-stage wakeup processing.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2025
From: XIAO, LONGSHUAI; ZHEN, YINAN; LI, WENJIE; PENG, CHAO; YANG, ZHANLEI
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 072144/0466 →
Priority Claims (1)
CN 202110348176.6 · Mar 31, 2021 · national
Continuity (2)
Continuation PCTCN2022083055 · Mar 25, 2022
Related Publication 20240029736A1 · Jan 25, 2024
References Cited (12)
US 11564036B1 · Kamath Koteshwara · 2023 [cited by examiner]
US 20180254040A1 · Droppo · 2018 [cited by examiner]
US 20200012887A1 · Li · 2020 [cited by examiner]
US 20200349927A1 · Stoimenov · 2020 [cited by examiner]
US 20220189471A1 · Sharifi · 2022 [cited by examiner]
CN 110277093A · 2019 [cited by applicant]
CN 110875045A · 2020 [cited by applicant]
CN 111223488A · 2020 [cited by applicant]
CN 112289311A · 2021 [cited by applicant]
Ji et al., “Integration of Multi-Look Beamformers for Multi-Channel Keyword Spotting,” ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, pp. 7464-7468,… [cited by applicant]
Wang et al., “Wake Word Detection with Alignment-Free Lattice-Free MMI,” arXiv:2005.08347v3 [eess.AS], total 5 pages (Jul. 28, 2020). [cited by applicant]
You et al., “Two-stage Strategy for Small-footprint Wake-up-word Speech Recognition System,” 2020 International Joint Conference of Neural Networks (IJCNN), IEEE, XP033831820, total 6 pages (Jul. 19, 2020). [cited by applicant]