IP Library Granted Patent US 11,321,583
Granted Patent B2
US 11,321,583 · App. 16/576,234 · Granted May 3, 2022

Image annotating method and electronic device

Inventors: Shiguo Lian (Guandong, CN); Zhaoxiang Liu (Guandong, CN); Ning Wang (Guandong, CN); Yibing Nan (Guandong, CN)
Assignee: CLOUDMINDS ROBOTICS CO., LTD.
G06K9/6253G06F3/16G06K9/6261G06K9/6263G06V10/22G10L15/08G10L15/22G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,321,583
App. No.
16/576,234
Granted
May 3, 2022
Kind
B2
Abstract

An image annotating method includes: acquiring an image collected at a terminal; acquiring voice information associated with the image; annotating the image according to the voice information; and storing an annotated result of the image.

Claims (87)

1. An image annotating method, comprising:

acquiring an image collected at a terminal;

acquiring voice information associated with the image;

extracting key features in the voice information;

annotating the image according to the key features; and

storing an annotated result of the image;

wherein prior to acquiring voice information associated with the image, the method comprises:

automatically annotating the image by image recognition;

wherein the annotating the image according to the voice information comprises:

annotating the image according to the voice information when automatic annotating fails.

2. The method according to claim 1 , wherein the image comprises a plurality of objects to be annotated, prior to acquiring voice information associated with the image, the method further comprising:

extracting region information of the objects to be annotated in the image by a region extraction algorithm;

performing sub-region division on the objects to be annotated in the image according to the region information;

sending a result of the sub-region division or an image performed with the sub-region division;

wherein the acquiring voice information associated with the image comprises: acquiring voice information associated with a sub-region in the image.

3. The method according to claim 2 , wherein after the sending the result of the sub-region division or the image performed with the sub-region division, the method further comprises:

acquiring an image obtained after the result of the sub-region division or the image performed with the sub-region division is subjected to adjustment operation at the terminal;

the annotating the image according to the voice information comprises: annotating the image subjected to the adjustment operation according to the voice information.

4. The method according to claim 1 , wherein the annotating the image according to the voice information; and storing the annotated result of the image comprises:

extracting keywords in the voice information based on voice recognition, wherein the keywords correspond to sub-regions; and

establishing a mapping relation table between the keywords and the sub-regions;

annotating the sub-regions according to the mapping relation table; and

storing the annotated result.

5. The method according to claim 1 , wherein prior to acquiring voice information associated with the image, the method comprises:

automatically annotating the image by image recognition;

displaying an automatically annotated result at the terminal after automatically annotating the image;

wherein the annotating the image according to the voice information comprises:

storing the automatically annotated result when the voice information indicates that the automatically annotated result is correct; and/or annotating the image according to the voice information when the voice information indicates that the automatically annotated result is incorrect.

6. An electronic device, comprising:

at least one processor; and,

a memory communicatively connected to the at least one processor;

wherein the memory stores an instruction program executable by the at least one processor, and the instruction program is executed by the at least one processor to cause the at least one processor to perform the steps of:

acquiring an image collected at a terminal;

acquiring voice information associated with the image;

extracting key features in the voice information;

annotating the image according to the key features; and

storing an annotated result of the image;

wherein prior to acquiring voice information associated with the image, the instruction program is executed by the at least one processor to cause the at least one processor to perform the steps of:

automatically annotating the image by image recognition;

wherein the annotating the image according to the voice information comprises:

annotating the image according to the voice information when automatic annotating fails.

7. The electronic device according to claim 6 , wherein prior to acquiring voice information associated with the image, the instruction program is executed by the at least one processor to cause the at least one processor to perform the steps of:

automatically annotating the image by image recognition;

displaying an automatically annotated result at the terminal after automatically annotating the image;

wherein the annotating the image according to the voice information comprises:

storing the automatically annotated result when the voice information indicates that the automatically annotated result is correct; and/or annotating the image according to the voice information when the voice information indicates that the automatically annotated result is incorrect.

8. The electronic device according to claim 6 , wherein the image comprises a plurality of objects to be annotated, and wherein prior to acquiring voice information associated with the image, the instruction program is executed by the at least one processor to cause the at least one processor to perform the steps of:

extracting region information of the objects to be annotated in the image by a region extraction algorithm;

performing sub-region division on the objects to be annotated in the image according to the region information;

sending a result of the sub-region division or an image performed with the sub-region division;

wherein the acquiring voice information associated with the image comprises: acquiring voice information associated with a sub-region in the image.

9. The electronic device according to claim 8 , wherein after the sending the result of the sub-region division or the image performed with the sub-region division, the instruction program is executed by the at least one processor to cause the at least one processor to perform the steps of:

acquiring an image obtained after the result of the sub-region division or the image performed with the sub-region division is subjected to adjustment operation at the terminal;

the annotating the image according to the voice information comprises: annotating the image subjected to the adjustment operation according to the voice information.

10. The electronic device according to claim 6 , wherein the annotating the image according to the voice information; and storing the annotated result of the image comprises:

extracting keywords in the voice information based on voice recognition, wherein the keywords correspond to sub-regions; and

establishing a mapping relation table between the keywords and the sub-regions;

annotating the sub-regions according to the mapping relation table; and

storing the annotated result.

11. A non-volatile computer readable storage medium, wherein the computer readable storage medium stores computer executable instructions configured to cause a computer to perform the steps of:

acquiring an image collected at a terminal;

acquiring voice information associated with the image;

extracting key features in the voice information;

annotating the image according to the key features; and

storing an annotated result of the image;

wherein prior to acquiring voice information associated with the image, the computer executable instructions configured to cause the computer to perform the steps of:

automatically annotating the image by image recognition;

wherein the annotating the image according to the voice information comprises:

annotating the image according to the voice information when automatic annotating fails.

12. The non-volatile computer readable storage medium according to claim 11 , wherein the annotating the image according to the voice information; and storing the annotated result of the image comprises:

extracting keywords in the voice information based on voice recognition, wherein the keywords correspond to sub-regions; and

establishing a mapping relation table between the keywords and the sub-regions;

annotating the sub-regions according to the mapping relation table; and

storing the annotated result.

13. The non-volatile computer readable storage medium according to claim 11 , wherein prior to acquiring voice information associated with the image, the computer executable instructions configured to cause the computer to perform the steps of:

automatically annotating the image by image recognition;

displaying an automatically annotated result at the terminal after automatically annotating the image;

wherein the annotating the image according to the voice information comprises:

storing the automatically annotated result when the voice information indicates that the automatically annotated result is correct; and/or annotating the image according to the voice information when the voice information indicates that the automatically annotated result is incorrect.

14. The non-volatile computer readable storage medium according to claim 11 , wherein the image comprises a plurality of objects to be annotated, prior to acquiring voice information associated with the image, the computer executable instructions configured to cause the computer to perform the steps of:

extracting region information of the objects to be annotated in the image by a region extraction algorithm;

performing sub-region division on the objects to be annotated in the image according to the region information;

sending a result of the sub-region division or an image performed with the sub-region division;

wherein the acquiring voice information associated with the image comprises: acquiring voice information associated with a sub-region in the image.

15. The non-volatile computer readable storage medium according to claim 14 , wherein after the sending the result of the sub-region division or the image performed with the sub-region division, the computer executable instructions configured to cause the computer to perform the steps of:

acquiring an image obtained after the result of the sub-region division or the image performed with the sub-region division is subjected to adjustment operation at the terminal;

the annotating the image according to the voice information comprises: annotating the image subjected to the adjustment operation according to the voice information.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2026
From: DATAA NEW TECHNOLOGY CO., LTD.
To: CHONGQING XINGJIE SHUXING TECHNOLOGY PARTNERSHIP ENTERPRISE (LIMITED PARTNERSHIP)
Reel/Frame 074153/0658 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2025
From: CLOUDMINDS ROBOTICS CO., LTD.
To: DATAA NEW TECHNOLOGY CO., LTD.
Reel/Frame 072052/0055 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2021
From: CLOUDMINDS (SHENZHEN) ROBOTICS SYSTEMS CO., LTD.
To: CLOUDMINDS ROBOTICS CO., LTD.
Reel/Frame 055620/0238 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2019
From: LIAN, SHIGUO; LIU, ZHAOXIANG; WANG, NING; NAN, YIBING
To: CLOUDMINDS (SHENZHEN) ROBOTICS SYSTEMS CO., LTD.
Reel/Frame 050446/0899 →
Continuity (2)
Continuation PCTCN2017077253 · Mar 20, 2017
Related Publication 20200012888A1 · Jan 9, 2020