IP Library › Granted Patent US 12,430,895
Granted Patent B2
US 12,430,895 · App. 17/951,271 · Granted Sep 30, 2025

Method and apparatus for updating object recognition model

Inventors: Jun Yue (Shenzhen, CN); Li Qian (Shenzhen, CN); Songcen Xu (Shenzhen, CN); Bin Shao (Shenzhen, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06V10/7747G06V10/26G06V10/464G06V10/761G06V10/764G06V10/809
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,895
App. No.
17/951,271
Filed
Sep 23, 2022
Granted
Sep 30, 2025
Kind
B2
Art Unit
2667
USPC
382/103
Abstract

A method and apparatus for updating an object recognition model in the field of artificial intelligence are disclosed. According to the method, a target image and first voice information of a user are obtained. The first voice information indicates a first category of a target object in the target image. A feature library of a first object recognition model is updated based on the target image and the first voice information. The updated first object recognition model includes a feature of the target object and a first label indicating the first category, and the feature of the target object corresponds to the first label. A recognition rate of an object recognition model can be improved more easily according to the technical solution provided in this application.

Claims (71)

1. A method for updating an object recognition model, comprising:

obtaining a target image captured by a photographing device;

obtaining first voice information based on a voice instruction captured by a voice device, wherein the first voice information indicates a first category of a target object in the target image;

generating a first label indicating the first category of the target object; and

updating a first object recognition model to form a updated first object recognition model based on the target image and the first voice information, to add the first label into the first object recognition model, wherein the updated first object recognition model comprises a feature of the target object and the first label, and there is a correspondence between the feature of the target object and the first label,

wherein the updating a first object recognition model based on the target image and the first voice information comprises:

determining, based on a similarity between the first label and each of at least one category of label, that the first label is a first category of label in the at least one category of label, wherein a similarity between the first label and the first category of label is greater than a similarity between the first label and another category of label in the at least one category of label;

determining, based on the target image by using second object recognition model, a probability that a category label of the target object is the first category of label; and

when the probability is greater than or equal to a probability threshold, adding the feature of the target object and the first label to the first object recognition model.

2. The method according to claim 1 , wherein the determining, based on a similarity between the first label and each of at least one category of label, that the first label is a first category of label in the at least one category of label comprises:

determining, based on a similarity between a semantic feature of the first label and a semantic feature of each of the at least one category of label, that the first label is the first category of label in the at least one category of label; and

that a similarity between the first label and the first category of label is greater than a similarity between the first label and another category of label in the at least one category of label comprises:

a distance between the semantic feature of the first label and a semantic feature of the first category of label is less than a distance between the semantic feature of the first label and a semantic feature of the another category of label.

3. The method according to claim 1 , wherein the target image comprises a first object, the target object is an object, in the target image, that is located in a direction indicated by the first object and that is closest to the first object among objects located in the direction indicated by the first object, and the first object comprises an eyeball or a finger.

4. The method according to claim 3 , wherein the updating a first object recognition model based on the target image and the first voice information comprises:

determining a bounding box of the first object in the target image;

determining, based on an image in the bounding box, the direction indicated by the first object;

performing visual saliency detection on the target image, to obtain a plurality of salient regions in the target image;

determining a target salient region among the plurality of salient regions based on the direction indicated by the first object, wherein the target salient region is in the direction indicated by the first object and is closest to the bounding box of the first object among salient regions in the direction indicated by the first object; and

updating the first object recognition model based on the target salient region, wherein an object in the target salient region comprises the target object.

5. The method according to claim 4 , wherein the determining, based on an image in the bounding box, the direction indicated by the first object comprises:

classifying the image in the bounding box by using a classification model, to obtain a target category of the first object; and

determining, based on the target category of the first object, the direction indicated by the first object.

6. An apparatus for updating an object recognition model, comprising at least one processor coupled to a memory;

the memory is configured to store instructions that, when executed by the at least one processor, enable the apparatus to perform operations comprising:

obtaining a target image captured by a photographing device;

obtaining first voice information based on a voice instruction captured by a voice device, wherein the first voice information indicates a first category of a target object in the target image;

generating a first label indicating the first category of the target object; and

updating a first object recognition model to form a updated first object recognition model based on the target image and the first voice information, to add the first label into the first object recognition model, wherein the updated first object recognition model comprises a feature of the target object and the first label, and there is a correspondence between the feature of the target object and the first label,

wherein the updating a first object recognition model based on the target image and the first voice information comprises:

determining, based on a similarity between the first label and each of at least one category of label, that the first label is a first category of label in the at least one category of label, wherein a similarity between the first label and the first category of label is greater than a similarity between the first label and another category of label in the at least one category of label;

determining, based on the target image by using a second object recognition model, a probability that a category label of the target object is the first category of label; and

when the probability is greater than or equal to a probability threshold, adding the feature of the target object and the first label to the first object recognition model.

7. The apparatus according to claim 6 , wherein the determining, based on a similarity between the first label and each of at least one category of label, that the first label is a first category of label in the at least one category of label comprises:

determining, based on a similarity between a semantic feature of the first label and a semantic feature of each of the at least one category of label, that the first label is the first category of label in the at least one category of label; and

that a similarity between the first label and the first category of label is greater than a similarity between the first label and another category of label in the at least one category of label comprises:

a distance between the semantic feature of the first label and a semantic feature of the first category of label is less than a distance between the semantic feature of the first label and a semantic feature of the another category of label.

8. The apparatus according to claim 6 , wherein the target image comprises a first object, the target object is an object, in the target image, that is located in a direction indicated by the first object and that is closest to the first object among objects located in the direction indicated by the first object, and the first object comprises an eyeball or a finger.

9. The apparatus according to claim 8 , wherein the updating a first object recognition model based on the target image and the first voice information comprises:

determining a bounding box of the first object in the target image;

determining, based on an image in the bounding box, the direction indicated by the first object;

performing visual saliency detection on the target image, to obtain a plurality of salient regions in the target image;

determining a target salient region among the plurality of salient regions based on the direction indicated by the first object, wherein the target salient region is in the direction indicated by the first object and is closest to the bounding box of the first object among salient regions in the direction indicated by the first object; and

updating the first object recognition model based on the target salient region, wherein an object in the target salient region comprises the target object.

10. The apparatus according to claim 9 , wherein the determining, based on an image in the bounding box, the direction indicated by the first object comprises:

classifying the image in the bounding box by using a classification model, to obtain a target category of the first object; and

determining, based on the target category of the first object, the direction indicated by the first object.

11. A non-transitory computer-readable medium storing information comprising instructions that, when executed by at least one processor, enable the at least one processor to perform operations comprising:

obtaining a target image captured by a photographing device;

obtaining first voice information based on a voice instruction captured by a voice device, wherein the first voice information indicates a first category of a target object in the target image;

generating a first label indicating the first category of the target object; and

updating a first object recognition model to form a updated first object recognition model based on the target image and the first voice information, to add the first label into the first object recognition model, wherein the updated first object recognition model comprises a feature of the target object and the first label, and there is a correspondence between the feature of the target object and the first label,

wherein the updating a first object recognition model based on the target image and the first voice information comprises:

determining, based on a similarity between the first label and each of at least one category of label, that the first label is a first category of label in the at least one category of label, wherein a similarity between the first label and the first category of label is greater than a similarity between the first label and another category of label in the at least one category of label;

determining, based on the target image by using a second object recognition model, a probability that a category label of the target object is the first category of label; and

when the probability is greater than or equal to a probability threshold, adding the feature of the target object and the first label to the first object recognition model.

12. The non-transitory computer-readable medium according to claim 11 , wherein the determining, based on a similarity between the first label and each of at least one category of label, that the first label is a first category of label in the at least one category of label comprises:

determining, based on a similarity between a semantic feature of the first label and a semantic feature of each of the at least one category of label, that the first label is the first category of label in the at least one category of label; and

that a similarity between the first label and the first category of label is greater than a similarity between the first label and another category of label in the at least one category of label comprises:

a distance between the semantic feature of the first label and a semantic feature of the first category of label is less than a distance between the semantic feature of the first label and a semantic feature of the another category of label.

13. The non-transitory computer-readable medium according to claim 11 , wherein the target image comprises a first object, the target object is an object, in the target image, that is located in a direction indicated by the first object and that is closest to the first object among objects located in the direction indicated by the first object, and the first object comprises an eyeball or a finger.

14. The non-transitory computer-readable medium according to claim 13 , wherein the updating a first object recognition model based on the target image and the first voice information comprises:

determining a bounding box of the first object in the target image;

determining, based on an image in the bounding box, the direction indicated by the first object;

performing visual saliency detection on the target image, to obtain a plurality of salient regions in the target image;

determining a target salient region among the plurality of salient regions based on the direction indicated by the first object, wherein the target salient region is in the direction indicated by the first object and is closest to the bounding box of the first object among salient regions in the direction indicated by the first object; and

updating the first object recognition model based on the target salient region, wherein an object in the target salient region comprises the target object.

15. The non-transitory computer-readable medium according to claim 14 , wherein the determining, based on an image in the bounding box, the direction indicated by the first object comprises:

classifying the image in the bounding box by using a classification model, to obtain a target category of the first object; and

determining, based on the target category of the first object, the direction indicated by the first object.

16. The method according to claim 1 , wherein the first voice information is obtained, using a natural language understanding method, based on the voice instruction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 2, 2022
From: YUE, JUN; QIAN, LI; XU, SONGCEN; SHAO, BIN
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 061950/0144 →
Priority Claims (1)
CN 202010215064.9 · Mar 24, 2020 · national
Continuity (2)
Continuation PCTCN2021082003 · Mar 22, 2021
Related Publication 20230020965A1 · Jan 19, 2023
References Cited (19)
US 8463608B2 · Dow et al. · 2013 [cited by applicant]
US 10381022B1 · Chaudhuri et al. · 2019 [cited by applicant]
US 10410627B2 · Cohen et al. · 2019 [cited by applicant]
US 11017317B2 · Rajkumar · 2021 [cited by examiner]
US 20140289323A1 · Kutaragi · 2014 [cited by examiner]
US 20180286397A1 · Nakano · 2018 [cited by examiner]
US 20200042796A1 · Kim · 2020 [cited by examiner]
US 20200175384A1 · Zhang · 2020 [cited by examiner]
CN 101943982A · 2011 [cited by applicant]
CN 107223246A · 2017 [cited by examiner]
CN 107808149A · 2018 [cited by applicant]
CN 108288208A · 2018 [cited by applicant]
CN 109800805A · 2019 [cited by applicant]
CN 109978881A · 2019 [cited by applicant]
CN 110070107B · 2020 [cited by applicant]
Holzapfel H et al: “A dialogue approach to learning object descriptions and semantic categories”, Robotics and Autonomous Systems, Elsevier BV, Amsterdam, NL, vol. 56, No. 11, Nov. 30, 2008, XP025561295, 10 pages. [cited by applicant]
Extended European Search Report issued in EP21775238.5, dated Jun. 19, 2023, 9 pages. [cited by applicant]
International Search Report and Written Opinion issued in PCT/CN2021/082003, dated Jun. 25, 2021, 9 pages. [cited by applicant]
Office Action issued in CN202010215064.9, dated Nov. 30, 2024, 9 pages. [cited by applicant]