IP Library Granted Patent US 12,555,575
Granted Patent B2
US 12,555,575 · App. 17/915,465 · Granted Feb 17, 2026

Wakeup indicator monitoring method, apparatus and electronic device

Inventors: Xu Li (Beijing, CN); Zeming Chen (Beijing, CN)
Assignee: Beijing Baidu Netcom Science Technology Co., Ltd.
G10L15/22G10L15/02G10L15/08G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,575
App. No.
17/915,465
Granted
Feb 17, 2026
Kind
B2
Abstract

The present wakeup indicator monitoring method, apparatus and electronic device includes: acquiring M pieces of audio data of a device to be monitored; determining a first wakeup confidence of each piece of M pieces of audio data, wherein the first wakeup confidence indicates a probability that the audio data contains a first wakeup word for waking up the device to be monitored; acquiring a first audio data with a first wakeup confidence in a target zone in M pieces of audio data, wherein the wakeup confidence in the target zone indicates that the audio data contains a wakeup word for waking up an audio device; and determining the ratio of the first audio data to M pieces of audio data as a wakeup rate of a device to be monitored, where the wakeup indicator of the device to be monitored includes the wakeup rate.

Claims (66)

1 . A wakeup indicator monitoring method, applied to a voice interaction device to be monitored and used for audio testing of the voice interaction device, the method is performed by a wakeup indicator monitoring device and comprises:

acquiring, by the wakeup indicator monitoring device, M pieces of audio data of the voice interaction device to be monitored, wherein M is a positive integer greater than 1;

determining, by the wakeup indicator monitoring device, a first wakeup confidence for each piece of the M pieces of audio data, the first wakeup confidence indicating a probability that audio data contains a first wakeup word for waking up the device to be monitored;

acquiring, by the wakeup indicator monitoring device, a first audio data with a first wakeup confidence in a target zone in the M pieces of audio data, wherein the wakeup confidence indicates that the audio data in the target zone comprises a wakeup word for waking up an audio device; and

determining, by the wakeup indicator monitoring device, the ratio of the first audio data to the M pieces of audio data as a wakeup rate of the device to be monitored, where a wakeup indicator of the device to be monitored comprises the wakeup rate;

wherein prior to obtaining M pieces of audio data for a device to be monitored, the method further comprises:

acquiring, by the wakeup indicator monitoring device, P pieces of audio data of N audio devices and an annotation result of the P pieces of audio data, wherein the annotation result indicates whether the audio data contains a second wakeup word for waking up the audio devices, N is a positive integer, and P is a positive integer greater than 1;

determining, by the wakeup indicator monitoring device, a second wakeup confidence of each piece of the P pieces of audio data; and

counting, by the wakeup indicator monitoring device, an zone where a second wakeup confidence of second audio data of which the ratio is greater than a pre-set threshold value in the P pieces of audio data is located, and obtaining the target zone, wherein the second audio data is the audio data with an annotation result that characterizes it contains the audio data of the second wakeup word;

wherein the P pieces of audio data are obtained from audio log data of the N audio devices, the audio log data comprising a plurality of audio data, and the acquiring P pieces of audio data of N audio devices comprises:

classifying, by the wakeup indicator monitoring device, each piece of audio data in the audio log data by L dimensions respectively so as to obtain L pieces of classification characteristic information about each piece of audio data in the audio log data, L being a positive integer;

determining, by the wakeup indicator monitoring device, audio feature information for each dimension based on the classified feature information of the audio log data;

respectively sampling, by the wakeup indicator monitoring device, in the audio log data based on audio feature information about each dimension so as to obtain an audio sampling result of the L dimensions; and

generating, by the wakeup indicator monitoring device, the P pieces of audio data comprising the audio sampling results of the L dimensions.

2 . The method of claim 1 , wherein the determining a first wakeup confidence for each piece of the M pieces of audio data comprises:

performing feature extraction on target audio data to obtain audio features of the target audio data, wherein the target audio data is any one of the M pieces of audio data; and

scoring the target audio data based on the audio features to obtain a first wakeup confidence of the target audio data.

3 . A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method of claim 1 .

4 . A wakeup indicator monitoring method, applied to a voice interaction device to be monitored and used for audio testing of the voice interaction device, the method is performed by a wakeup indicator monitoring device and comprises:

acquiring, by the wakeup indicator monitoring device, M pieces of audio data of the voice interaction device to be monitored, wherein M is a positive integer greater than 1;

determining, by the wakeup indicator monitoring device, a first wakeup confidence for each piece of the M pieces of audio data, the first wakeup confidence indicating a probability that audio data contains a first wakeup word for waking up the device to be monitored;

acquiring, by the wakeup indicator monitoring device, first audio data with a first wakeup confidence in a target zone in the M pieces of audio data, wherein the wakeup confidence in the target zone characterizes that the audio data does not contain a wakeup word for waking up an audio device; and

determining, by the wakeup indicator monitoring device, the ratio of the first audio data to the M pieces of audio data as a false wakeup rate of the device to be monitored, where a wakeup indicator of the device to be monitored comprises the false wakeup rate;

wherein prior to obtaining M pieces of audio data for a device to be monitored, the method further comprises:

acquiring, by the wakeup indicator monitoring device, P pieces of audio data of N audio devices and an annotation result of the P pieces of audio data, wherein the annotation result indicates whether the audio data contains a second wakeup word for waking up the audio devices, N is a positive integer, and P is a positive integer greater than 1;

determining, by the wakeup indicator monitoring device, a second wakeup confidence of each piece of the P pieces of audio data; and

counting, by the wakeup indicator monitoring device, an zone where a second wakeup confidence of second audio data of which the ratio is greater than a pre-set threshold value in the P pieces of audio data is located, and obtaining the target zone, wherein the second audio data is the audio data with an annotation result that characterizes it contains the audio data of the second wakeup word;

wherein the P pieces of audio data are obtained from audio log data of the N audio devices, the audio log data comprising a plurality of audio data, and the acquiring P pieces of audio data of N audio devices comprises:

classifying, by the wakeup indicator monitoring device, each piece of audio data in the audio log data by L dimensions respectively so as to obtain L pieces of classification characteristic information about each piece of audio data in the audio log data, L being a positive integer;

determining, by the wakeup indicator monitoring device, audio feature information for each dimension based on the classified feature information of the audio log data;

respectively sampling, by the wakeup indicator monitoring device, in the audio log data based on audio feature information about each dimension so as to obtain an audio sampling result of the L dimensions; and

generating, by the wakeup indicator monitoring device, the P pieces of audio data comprising the audio sampling results of the L dimensions.

5 . An electronic device comprising:

at least one processor; and

a memory communicatively coupled to the at least one processor; wherein,

the memory stores instructions are executable by the at least one processor to enable the at least one processor to perform the method of claim 4 .

6 . A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method of claim 4 .

7 . An electronic device comprising:

at least one processor; and

a memory communicatively coupled to the at least one processor; wherein,

the memory stores instructions are executable by the at least one processor to enable the at least one processor to perform a wakeup indicator monitoring method, the method is applied to a voice interaction device to be monitored and used for audio testing of the voice interaction device, the method is performed by a wakeup indicator monitoring device and comprises:

acquiring M pieces of audio data of a voice interaction device to be monitored, wherein M is a positive integer greater than 1;

determining a first wakeup confidence for each piece of the M pieces of audio data, the first wakeup confidence indicating a probability that audio data contains a first wakeup word for waking up the device to be monitored;

acquiring a first audio data with a first wakeup confidence in a target zone in the M pieces of audio data, wherein the wakeup confidence indicates that the audio data in the target zone comprises a wakeup word for waking up an audio device; and

determining the ratio of the first audio data to the M pieces of audio data as a wakeup rate of the device to be monitored, where a wakeup indicator of the device to be monitored comprises the wakeup rate;

wherein prior to obtaining M pieces of audio data for a device to be monitored, the method further comprises:

acquiring P pieces of audio data of N audio devices and an annotation result of the P pieces of audio data, wherein the annotation result indicates whether the audio data contains a second wakeup word for waking up the audio devices, N is a positive integer, and P is a positive integer greater than 1;

determining a second wakeup confidence of each piece of the P pieces of audio data; and

counting an zone where a second wakeup confidence of second audio data of which the ratio is greater than a pre-set threshold value in the P pieces of audio data is located, and obtaining the target zone, wherein the second audio data is the audio data with an annotation result that characterizes it contains the audio data of the second wakeup word;

wherein the P pieces of audio data are obtained from audio log data of the N audio devices, the audio log data comprising a plurality of audio data, and the acquiring P pieces of audio data of N audio devices comprises:

classifying each piece of audio data in the audio log data by L dimensions respectively so as to obtain L pieces of classification characteristic information about each piece of audio data in the audio log data, L being a positive integer;

determining audio feature information for each dimension based on the classified feature information of the audio log data;

respectively sampling in the audio log data based on audio feature information about each dimension so as to obtain an audio sampling result of the L dimensions; and

generating the P pieces of audio data comprising the audio sampling results of the L dimensions.

8 . The electronic device of claim 7 , wherein the at least one processor is configured to perform:

prior to obtaining M pieces of audio data for a device to be monitored, acquiring P pieces of audio data of N audio devices and an annotation result of the P pieces of audio data, wherein the annotation result indicates whether the audio data contains a second wakeup word for waking up the audio devices, N is a positive integer, and P is a positive integer greater than 1;

determining a second wakeup confidence of each piece of the P pieces of audio data; and

counting an zone where a second wakeup confidence of second audio data of which the ratio is greater than a pre-set threshold value in the P pieces of audio data is located, and obtaining the target zone, wherein the second audio data is the audio data with an annotation result that characterizes it contains the audio data of the second wakeup word.

9 . The electronic device of claim 8 , wherein the P pieces of audio data are obtained from audio log data of the N audio devices, the audio log data comprising a plurality of audio data, and the acquiring P pieces of audio data of N audio devices comprises:

classifying each piece of audio data in the audio log data by L dimensions respectively so as to obtain L pieces of classification characteristic information about each piece of audio data in the audio log data, L being a positive integer;

determining audio feature information for each dimension based on the classified feature information of the audio log data;

respectively sampling in the audio log data based on audio feature information about each dimension so as to obtain an audio sampling result of the L dimensions; and

generating the P pieces of audio data comprising the audio sampling results of the L dimensions.

10 . The electronic device of claim 7 , wherein the determining a first wakeup confidence for each piece of the M pieces of audio data comprises:

performing feature extraction on target audio data to obtain audio features of the target audio data, wherein the target audio data is any one of the M pieces of audio data; and

scoring the target audio data based on the audio features to obtain a first wakeup confidence of the target audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2022
From: LI, XU; CHEN, ZEMING
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 061678/0672 →
Priority Claims (1)
CN 202011577341.7 · Dec 28, 2020 · national
Continuity (1)
Related Publication 20230130399A1 · Apr 27, 2023
References Cited (56)
US 9275637B1 · Salvador · 2016 [cited by examiner]
US 11295741B2 · Mont-Reynaud · 2022 [cited by examiner]
US 11521599B1 · Jose · 2022 [cited by examiner]
US 20150154953A1 · Bapat · 2015 [cited by examiner]
US 20160077792A1 · Bansal · 2016 [cited by examiner]
US 20200105256A1 · Fainberg · 2020 [cited by examiner]
US 20200160858A1 · Gomes · 2020 [cited by examiner]
US 20200193971A1 · Feinauer et al. · 2020 [cited by applicant]
US 20200365152A1 · Han · 2020 [cited by examiner]
US 20210183370A1 · Saadoun · 2021 [cited by examiner]
US 20220013111A1 · Chen et al. · 2022 [cited by applicant]
US 20220223150A1 · Li · 2022 [cited by examiner]
CN 106782536A · 2017 [cited by applicant]
CN 107767861A · 2018 [cited by applicant]
CN 107871506A · 2018 [cited by applicant]
CN 107990908A · 2018 [cited by applicant]
CN 108536668A · 2018 [cited by examiner]
CN 108694940A · 2018 [cited by applicant]
CN 108735203A · 2018 [cited by applicant]
CN 108847219A · 2018 [cited by applicant]
CN 109817219A · 2019 [cited by applicant]
CN 110164474A · 2019 [cited by applicant]
CN 110428810A · 2019 [cited by applicant]
CN 110473539A · 2019 [cited by applicant]
CN 110570840A · 2019 [cited by applicant]
CN 110838289A · 2020 [cited by examiner]
CN 111081241A · 2020 [cited by examiner]
CN 111739512A · 2020 [cited by applicant]
CN 111767083A · 2020 [cited by applicant]
CN 111880856A · 2020 [cited by applicant]
KR 20190034964A · 2019 [cited by applicant]
Ma, L., Zhang, H., Zhao, P., & Su, T. (2020). Competitive Wakeup Scheme for Distributed Devices. arXiv preprint arXiv:2005.09242. (Year: 2020). [cited by examiner]
Extended European Search Report corresponding to European Patent Application No. 21912809.7, dated Oct. 12, 2023. (10 Pages). [cited by applicant]
English Translation of CN107767861A. (23 Pages). [cited by applicant]
Japanese Office Action corresponding to Japanese Patent Application No. 2022-514849, dated Apr. 18, 2023. (4 Pages). [cited by applicant]
Machine English Translation Japanese Office Action corresponding to Japanese Patent Application No. 2022-514849, dated Apr. 18, 2023. (4 Pages). [cited by applicant]
Machine English Translation of CN111767083A. (25 Pages). [cited by applicant]
English Translation of International Search Report corresponding to International Patent Application No. PCT/CN2021/092100, dated Sep. 28, 2021. (6 pages). [cited by applicant]
International Search Report corresponding to International Patent Application No. PCT/CN2021/092100, dated Sep. 28, 2021. (9 pages). [cited by applicant]
Chinese Office Action corresponding to Chinese Patent Application 202011577341.7, dated Jan. 6, 2022. (9 Pages). [cited by applicant]
Machine English Translation Chinese Office Action corresponding to Chinese Patent Application 202011577341.7, dated Jan. 6, 2022. (5 Pages). [cited by applicant]
Machine English Translation of CN110473539A. (22 Pages). [cited by applicant]
Machine English Translation of CN109817219A. (15 Pages). [cited by applicant]
Machine English Translation of CN108847219A. (18 Pages). [cited by applicant]
Machine English Translation of CN108735203A. (17 Pages). [cited by applicant]
Machine English Translation of CN108694940A. (32 Pages). [cited by applicant]
Machine English Translation of CN107990908A. (22 Pages). [cited by applicant]
Machine English Translation of CN107871506A. (22 Pages). [cited by applicant]
Machine English Translation of KR20190034964A. (23 Pages). [cited by applicant]
Machine English Translation of CN111880856A. (34 Pages). [cited by applicant]
Machine English Translation of CN111081241A. (20 Pages). [cited by applicant]
Machine English Translation of CN110570840A. (32 Pages). [cited by applicant]
Machine English Translation of CN110428810A. (35 Pages). [cited by applicant]
Machine English Translation of CN111739512A. (25 Pages). [cited by applicant]
Machine English Translation of CN110164474A. (21 Pages). [cited by applicant]
Machine English Translation of CN106782536A. (21 Pages). [cited by applicant]