IP Library › Granted Patent US 10,621,437
Granted Patent B2
US 10,621,437 · App. 15/926,788 · Granted Apr 14, 2020

Method, apparatus for controlling a smart device and computer storge medium

Inventor: Qiao Ren (Beijing, CN)
Assignee: BEIJING XIAOMI MOBILE SOFTWARE CO., LTD.
G06K9/00671G06F3/0482G06F3/0484G06F3/0488G06K9/00744G06K9/6202
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,621,437
App. No.
15/926,788
Granted
Apr 14, 2020
Kind
B2
Abstract

The disclosure relates to a method for controlling a smart device, an apparatus, and non-transitory computer-readable medium. The method includes acquiring a video stream captured by a smart camera that is bound to the user account, wherein the video stream includes multi-frame video that includes a plurality of one-frame video images; performing pattern recognition on each of the plurality of one-frame video images, wherein the pattern recognition is configured to determine an area that includes at least one smart device in at least one of the plurality of one-frame video images; determining, based on the pattern recognition, a target area that includes the smart device in a first one-frame video image of the plurality of one-frame video images; displaying the first one-frame video image including the target area on a touch screen; detecting, via the touch screen, a control operation within the target area of the first one-frame video image; and controlling the smart device located in the target area based on the control operation.

Claims (92)

1. A method for controlling a smart device that is bound to a user account, comprising:

acquiring a video stream captured by a smart camera that is bound to the user account, wherein the video stream includes multi-frame video that includes a plurality of one-frame video images;

performing pattern recognition on each of the plurality of one-frame video images, wherein the pattern recognition is configured to determine an area that includes at least one smart device in at least one of the plurality of one-frame video images;

determining, based on the pattern recognition, a target area that includes the smart device in a first one-frame video image of the plurality of one-frame video images;

displaying the first one-frame video image including the target area on a touch screen;

detecting, via the touch screen, a control operation within the target area of the first one-frame video image; and

controlling the smart device located in the target area based on the control operation,

wherein the method further comprises:

determining, based on the pattern recognition, a plurality of image areas in the first one-frame video image;

performing feature extraction on each of the plurality of image areas;

obtaining a plurality of first feature vectors based on the feature extraction; and

determining a plurality of smart devices included in the first one-frame video image and corresponding ones of the plurality of image areas where each of the smart devices is located based on the plurality of first feature vectors and a plurality of second feature vectors, wherein there is a one-to-one correspondence between the plurality of second feature vectors and a plurality of smart devices bound to the user account.

2. The method of claim 1 , further comprising:

for each second feature vector of the plurality of second feature vectors, determining an Euclidean distance between the second feature vector and each first feature vector of the plurality of first feature vectors, to obtain a plurality of Euclidean distances;

determining that the first one-frame video image includes the smart device corresponding to the second feature vector when a minimum Euclidean distance among the plurality of Euclidean distances is less than a distance threshold; and

determining a first image area of the plurality of image areas that is associated with the minimum Euclidian distance as the target area that includes the smart device corresponding to the second feature vector in the first one-frame video image.

3. The method of claim 2 , further comprising:

before determining that the first one-frame video image includes the smart device corresponding to the second feature vector, displaying identity confirmation information of the smart device, wherein the identity confirmation information of the smart device includes a device identification of the smart device corresponding to the second feature vector; and

determining that the first one-frame video image includes the smart device corresponding to the second feature vector when a confirmation command for the identity confirmation information of the smart device is received.

4. The method of claim 3 , further comprising:

before determining the plurality of smart devices included in the first one-frame video image, acquiring an image of each smart device of the plurality of smart devices bound to the user account;

performing feature extraction on each of the images of the plurality of smart devices to obtain feature vectors of the plurality of smart devices; and

storing the feature vectors of the plurality of smart devices as the second feature vectors.

5. The method of claim 2 , further comprising:

before determining plurality of smart devices included in the first one-frame video image, acquiring an image of each smart device of the plurality of smart devices bound to the user account;

performing feature extraction on each of the images of the plurality of smart devices to obtain feature vectors of the plurality of smart devices; and

storing the feature vectors of the plurality of smart devices as the second feature vectors.

6. The method of claim 1 , further comprising:

before determining the plurality of smart devices included in the first one-frame video image, acquiring an image of each smart device of the plurality of smart devices bound to the user account;

performing feature extraction on each of the images of the plurality of smart devices to obtain feature vectors of the plurality of smart devices; and

storing the feature vectors of the plurality of smart devices as the second feature vectors.

7. The method of claim 1 , wherein controlling the smart device located in the target area based on the control operation comprises:

displaying a control interface of the smart device located in the target area, wherein the control interface includes a plurality of control options;

receiving a selection operation that is configured to select one of the control options; and

controlling the smart device based on the selected one of the control options.

8. An apparatus for controlling a smart device that is bound to a user account, comprising:

a processor; and

a storage configured to store executable instructions executed by the processor;

wherein the processor is configured to:

acquire a video stream captured by a smart camera that is bound to the user account, wherein the video stream includes multi-frame video that includes a plurality of one-frame video images;

perform pattern recognition on each of the plurality of one-frame video images, wherein the pattern recognition is configured to determine an area that includes at least one smart device in at least one of the plurality of one-frame video images;

determine, based on the pattern recognition, a target area that includes the smart device in a first one-frame video image of the plurality of one-frame video images;

display the first one-frame video image including the target area on a touch screen;

detect, via the touch screen, a control operation within the target area of the first one-frame video image; and

control the smart device located in the target area based on the control operation,

wherein the processor is further configured to:

determine, based on the pattern recognition, a plurality of image areas in the first one-frame video image;

perform feature extraction on each of the plurality of image areas;

obtain a plurality of first feature vectors based on the feature extraction; and

determine a plurality of smart devices included in the first one-frame video image and corresponding ones of the plurality of image areas where each of the smart devices is located based on the plurality of first feature vectors and a plurality of second feature vectors, wherein there is a one-to-one correspondence between the plurality of second feature vectors and a plurality of smart devices bound to the user account.

9. The apparatus of claim 8 , wherein the processor is further configured to:

for each second feature vector of the plurality of second feature vectors, determine an Euclidean distance between the second feature vector and each first feature vector of the plurality of first feature vectors, to obtain a plurality of Euclidean distances;

determine that the first one-frame video image includes the smart device corresponding to the second feature vector when a minimum Euclidean distance among the plurality of Euclidean distances is less than a preset distance threshold; and

determine a first image area of the plurality of image areas that is associated with the minimum Euclidian distance as the target area that includes the smart device corresponding to the second feature vector in the first one-frame video image.

10. The apparatus of claim 9 , wherein the processor is further configured to:

display identity confirmation information of the smart device, wherein the identity confirmation information of the smart device includes a device identification of the smart device corresponding to the second feature vector; and

determine that the first one-frame video image includes the smart device corresponding to the second feature vector when a confirmation command for the identity confirmation information of the smart device is received.

11. The apparatus of claim 8 , wherein the processor is further configured to:

acquire an image of each smart device of the plurality of smart devices bound to the user account;

perform feature extraction one each of the images of the plurality of smart devices to obtain feature vectors of the plurality of smart devices; and

store the feature vectors of the plurality of smart devices as the second feature vectors.

12. The apparatus of claim 8 , wherein the processor is further configured to:

display a control interface of the smart device located in the target area, wherein the control interface includes a plurality of control options;

receive a selection operation that is configured to select one of the control options; and

control the smart device based on the selected one of the control options.

13. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors of a computing device, cause the computing device to:

acquire a video stream captured by a smart camera that is bound to the user account, wherein the video stream includes multi-frame video that includes a plurality of one-frame video images;

perform pattern recognition on each of the plurality of one-frame video images, wherein the pattern recognition is configured to determine an area that includes at least one smart device in at least one of the plurality of one-frame video images;

determine, based on the pattern recognition, a target area that includes the smart device in a first one-frame video image of the plurality of one-frame video images;

display the first one-frame video image including the target area on a touch screen;

detect, via the touch screen, a control operation within the target area of the first one-frame video image; and

control the smart device located in the target area based on the control operation,

wherein the instructions further cause the computing device to:

determine, based on the pattern recognition, a plurality of image areas in the first one-frame video image;

perform feature extraction on each of the plurality of image areas;

obtain a plurality of first feature vectors based on the feature extraction; and

determine a plurality of smart devices included in the first one-frame video image and corresponding ones of the plurality of image areas where each of the smart devices is located based on the plurality of first feature vectors and a plurality of second feature vectors, wherein there is a one-to-one correspondence between the plurality of second feature vectors and a plurality of smart devices bound to the user account.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the instructions cause the computing device to:

for each second feature vector of the plurality of second feature vectors, determine an Euclidean distance between the second feature vector and each first feature vector of the plurality of first feature vectors, to obtain a plurality of Euclidean distances;

determine that the first one-frame video image includes the smart device corresponding to the second feature vector when a minimum Euclidean distance among the plurality of Euclidean distances is less than a preset distance threshold; and

determine a first image area of the plurality of image areas that is associated with the minimum Euclidian distance as the target area that includes the smart device corresponding to the second feature vector in the first one-frame video image.

15. The non-transitory computer-readable storage medium of claim 14 , wherein the instructions cause the computing device to:

display identity confirmation information of the smart device, wherein the identity confirmation information of the smart device includes a device identification of the smart device corresponding to the second feature vector; and

determine that the first one-frame video image includes the smart device corresponding to the second feature vector when a confirmation command for the identity confirmation information of the smart device is received.

16. The non-transitory computer-readable storage medium of claim 13 , wherein the instructions cause the computing device to:

acquire an image of each smart device of the plurality of smart devices bound to the user account;

perform feature extraction one each of the images of the plurality of smart devices to obtain feature vectors of the plurality of smart devices; and

store the feature vectors of the plurality of smart devices as the second feature vectors.

17. The non-transitory computer-readable storage medium of claim 13 , wherein the instructions cause the computing device to:

display a control interface of the smart device located in the target area, wherein the control interface includes a plurality of control options;

receive a selection operation that is configured to select one of the control options; and

control the smart device based on the selected one of the control options.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2018
From: REN, QIAO
To: BEIJING XIAOMI MOBILE SOFTWARE CO., LTD.
Reel/Frame 045293/0192 →
Priority Claims (1)
CN 2017 1 0169381 · Mar 21, 2017 · national
Continuity (1)
Related Publication 20180276474A1 · Sep 27, 2018