IP Library Granted Patent US 12,555,568
Granted Patent B2
US 12,555,568 · App. 18/086,249 · Granted Feb 17, 2026

Device control method and apparatus, readable storage medium and chip

Inventors: Xiuyun Zhang (Beijing, CN); Chunxing Cai (Beijing, CN)
Assignee: Beijing Xiaomi Mobile Software Co., Ltd.
G10L15/08G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,568
App. No.
18/086,249
Granted
Feb 17, 2026
Kind
B2
Abstract

A device control method, includes: in response to determining a target voice module is awakened, a collected first voice information is sent to a first server, the first voice information is configured to convert the first voice information into a first control instruction by the first server and send the first control instruction to a second server, and the first voice information contains information for controlling a second target device; the first control instruction fed back by the second server is received; and the second target device is controlled by using the target voice module or a processor according to the first control instruction, and the target voice module and the processor are equipped in a first target device.

Claims (77)

1 . A method for controlling a device that includes a target voice module, the method comprising:

sending a first voice information to a first server in response to determining the target voice module is awakened, the first voice information is used to convert the first voice information into a first control instruction by the first server and send the first control instruction to a second server, and the first voice information contains information for controlling a second target device;

receiving the first control instruction fed back by the second server;

controlling the second target device by using the target voice module or a processor according to the first control instruction, and the target voice module and the processor are equipped in a first target device;

determining whether the target voice module is capable of processing the first control instruction;

performing, according to the first control instruction, audio function control on a target device in response to determining the target voice module is capable of processing the first control instruction; and

performing, according to the first control instruction, non-audio function control on the target device through a processor of the first target device in response to determining the target voice module is not capable of processing the first control instruction;

wherein the target voice module is awakened by:

receiving, by a third server, a second voice information sent by a third target device; and

awakening, by the third server, the target voice module from a plurality of voice modules according to priorities of the plurality of voice modules and distances between the plurality of voice modules and a sound source outputting the second voice information

wherein the target device is the first target device in which the target voice module is located, or the target device is the second target device that supports IoT control and differs from the first target device.

2 . The method according to claim 1 , wherein the performing, according to the first control instruction, non-audio function control on target device through a processor of the first target device comprises:

sending the first control instruction to the processor in response to determining the first voice information is a first preset voice information, and the first control instruction is used to instruct the processor to control the second target device according to the first control instruction;

wherein the first preset voice information is voice information irrelevant to audio playing.

3 . The method according to claim 1 , wherein the performing, according to the first control instruction, the audio function control on a target device comprises:

controlling the second target device according to the first control instruction in response to determining the first voice information is a second preset voice information;

wherein the second preset voice information is voice information related to audio playing.

4 . The method according to claim 1 , further comprising:

receiving a second control instruction fed back by the first server in response to determining the first voice information is a third preset voice information; and

controlling the second target device according to the second control instruction;

wherein the third preset voice information is voice information related to voice interaction.

5 . The method according to claim 2 , wherein the method further comprises:

receiving a configuration file sent by a fourth server in response to determining the first target device establishes a connection with the fourth server, and the configuration file is used to update an awakening function of the target voice module.

6 . The method according to claim 1 , wherein the method further comprises:

registering first device information of the first target device to the second server and a fourth server simultaneously, and the second server and the fourth server establishing connection with the target voice module of the first target device according to the first device information.

7 . The method according to claim 1 , wherein the method further comprises:

receiving an upgrade file sent by the second server; and

upgrading the target voice module and the processor according to the upgraded file.

8 . A non-transitory computer-readable storage medium storing computer program instructions for controlling a device, wherein the computer program instructions when executed by a processor cause the processor to execute a method comprising:

sending a first voice information to a first server in response to determining a target voice module is awakened, the first voice information is used to convert the first voice information into a first control instruction by the first server and send the first control instruction to a second server, and the first voice information contains information for controlling a second target device;

receiving the first control instruction fed back by the second server; and

controlling the second target device by using the target voice module or the processor according to the first control instruction, and the target voice module and the processor are equipped in a first target device;

wherein the target voice module is awakened by:

receiving, by a third server, a second voice information sent by a third target device; and

awakening, by the third server, the target voice module from a plurality of voice modules according to priorities of the plurality of voice modules and distances between the plurality of voice modules and a sound source outputting the second voice information

performing, according to the first control instruction, audio function control on a target device in response to determining the target voice module is capable of processing the first control instruction

performing, according to the first control instruction, non-audio function control on target device through a processor of the first target device in response to determining the target voice module is not capable of processing the first control instruction;

wherein the target device is the first target device in which the target voice module is located, or the target device is the second target device that supports IoT control and differs from the first target device.

9 . The non-transitory computer-readable storage medium according to claim 8 , wherein performing, according to the first control instruction, non-audio function control on target device through a processor of the first target device comprises:

send the first control instruction to the processor of the first target device in response to determining the first voice information is a first preset voice information, and the first control instruction is used to instruct the processor of the first target device to control the second target device according to the first control instruction;

wherein the first preset voice information is voice information irrelevant to audio playing.

10 . The non-transitory computer-readable storage medium according to claim 8 , wherein the performing, according to the first control instruction, the audio function control on a target device comprises:

controlling the second target device according to the first control instruction in response to determining the first voice information is a second preset voice information;

wherein the second preset voice information is voice information related to audio playing.

11 . The non-transitory computer-readable storage medium according to claim 8 , wherein the method further comprises

receiving a second control instruction fed back by the first server in response to determining the first voice information is a third preset voice information; and

controlling the second target device according to the second control instruction;

wherein the third preset voice information is voice information related to voice interaction.

12 . The non-transitory computer-readable storage medium according to claim 9 , wherein the method further comprises:

receiving a configuration file sent by a fourth server in response to determining the first target device establishes a connection with the fourth server, and the configuration file is used to update an awakening function of the target voice module.

13 . The non-transitory computer-readable storage medium according to claim 8 , wherein the method further comprises:

registering first device information of the first target device to the second server and a fourth server simultaneously, and the second server and the fourth server establishing connection with the target voice module of the first target device according to the first device information.

14 . The non-transitory computer-readable storage medium according to claim 8 , wherein the method further comprises:

receiving an upgrade file sent by the second server; and

upgrading the target voice module and the processor according to the upgraded file.

15 . A chip, comprising:

a processor; and

an interface that is communicatively coupled to the processor, wherein the processor is configured to:

receive a second voice information sent by a third target device through a third server; and

awaken, by the third server, a target voice module from a plurality of voice modules according to priorities of the plurality of voice modules and distances between the plurality of voice modules and a sound source outputting the second voice information;

wherein the target voice module is awakened by:

receiving, by the third server, the second voice information sent by the third target device;

awakening, by the third server, the target voice module from the plurality of voice modules according to priorities of the plurality of voice modules and the distances between the plurality of voice modules and the sound source outputting the second voice information;

performing, according to a first control instruction, audio function control on a target device in response to determining the target voice module is capable of processing the first control instruction;

performing, according to the first control instruction, non-audio function control on target device through a processor of a first target device in response to determining the target voice module is not capable of processing the first control instruction;

wherein the target device is the first target device in which the target voice module is located, or the target device is a second target device that supports IoT control and differs from the first target device.

16 . The chip according to claim 15 , wherein the

performing, according to the first control instruction non-audio function control on target device through a processor of the first target device comprises:

sending the first control instruction to the processor of the first target device in response to determining a first voice information is a first preset voice information, and the first control instruction is used to instruct the processor of the first target device to control the second target device according to the first control instruction;

wherein the first preset voice information is voice information irrelevant to audio playing.

17 . The chip according to claim 15 , wherein the performing, according to the first control instruction, the audio function control on a target device comprises:

controlling the second target device according to the first control instruction in response to determining a first voice information is a second preset voice information;

wherein the second preset voice information is voice information related to audio playing.

18 . The chip according to claim 15 , wherein the processor is further configured to:

receive a second control instruction fed back by a first server in response to determining the first voice information is a third preset voice information; and

controlling the second target device according to the second control instruction;

wherein the third preset voice information is voice information related to voice interaction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2022
From: ZHANG, XIUYUN; CAI, CHUNXING
To: BEIJING XIAOMI MOBILE SOFTWARE CO., LTD.
Reel/Frame 062175/0580 →
Priority Claims (1)
CN 202211193908.X · Sep 28, 2022 · national
Continuity (1)
Related Publication 20240105164A1 · Mar 28, 2024
References Cited (27)
US 10971173B2 · Kothari · 2021 [cited by examiner]
US 11393478B2 · Bates · 2022 [cited by examiner]
US 11514109B2 · Sharifi · 2022 [cited by examiner]
US 11822857B2 · Aiken · 2023 [cited by examiner]
US 11942085B1 · Mutagi · 2024 [cited by examiner]
US 12118995B1 · Blanksteen · 2024 [cited by examiner]
US 20190295542A1 · Huang et al. · 2019 [cited by applicant]
US 20200159491A1 · Mutagi · 2020 [cited by examiner]
US 20200194004A1 · Bates · 2020 [cited by examiner]
US 20210375281A1 · Gao · 2021 [cited by applicant]
US 20220013121A1 · Ni · 2022 [cited by examiner]
US 20220028379A1 · Carbune · 2022 [cited by examiner]
US 20220139573A1 · Sharifi · 2022 [cited by examiner]
US 20230362026A1 · Bajaj · 2023 [cited by examiner]
CN 104965448A · 2015 [cited by applicant]
CN 108668153A · 2018 [cited by applicant]
CN 110070863A · 2019 [cited by applicant]
CN 111261151A · 2020 [cited by applicant]
CN 111722824A · 2020 [cited by applicant]
CN 111724784A · 2020 [cited by applicant]
CN 112151013A · 2020 [cited by applicant]
CN 113470634A · 2021 [cited by applicant]
CN 114023303A · 2022 [cited by applicant]
CN 115083401A · 2022 [cited by applicant]
KR 1020190050761A · 2019 [cited by applicant]
Chinese Office Action issued on Jul. 9, 2024 for Chinese Patent Application No. 202211193908. [cited by applicant]
Extended European Search Report issued on Aug. 16, 2023 for European Patent Application No. 22217062.3. [cited by applicant]