IP Library › Granted Patent US 12,127,726
Granted Patent B2
US 12,127,726 · App. 17/231,958 · Granted Oct 29, 2024

System and method for robust image-query understanding based on contextual features

Inventors: Yu Wang (Bellevue, WA); Yilin Shen (Santa Clara, CA); Hongxia Jin (San Jose, CA)
Assignee: Samsung Electronics Co., Ltd.
A47L9/2805A47L11/40G06F18/214G06V10/768G06V20/10A47L2201/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,127,726
App. No.
17/231,958
Granted
Oct 29, 2024
Kind
B2
Abstract

A method includes obtaining, using at least one processor of an electronic device, an image-query understanding model. The method also includes obtaining, using the at least one processor, an image and a user query associated with the image, where the image includes a target image area and the user query includes a target phrase. The method further includes retraining, using the at least one processor, the image-query understanding model using a correlation between the target image area and the target phrase to obtain a retrained image-query understanding model.

Claims (66)

1. A method comprising:

obtaining, using at least one processor of a robot, an image-query understanding model;

obtaining, using the at least one processor, an image and a user query associated with the image, wherein the image comprises a target image area and the user query comprises a target phrase, and wherein the target image area is marked on the image by a user and the target phrase is identified within the user query by the user during operation of the robot; and

retraining, using the at least one processor, the image-query understanding model using a correlation between the target image area and the target phrase to obtain a retrained image-query understanding model;

wherein retraining the image-query understanding model comprises determining (i) one or more query-level contextual features, (ii) one or more question features, (iii) one or more image-level contextual features, and (iv) one or more post-processed image features based on at least one weighted attention function.

2. The method of claim 1 , wherein the image-query understanding model comprises:

a question contextual feature extraction (Q-CFE) module configured to determine the one or more query-level contextual features and the one or more question features; and

an image contextual feature extraction (I-CFE) module configured to determine the one or more image-level contextual features and the one or more post-processed image features.

3. The method of claim 2 , wherein the image-query understanding model further comprises a weighted contextual feature question-image understanding (WCUQIU) module configured to process (i) the one or more query-level contextual features and the one or more question features from the Q-CFE module and (ii) the one or more image-level contextual features and the one or more post-processed image features from the I-CFE module.

4. The method of claim 1 , wherein retraining the image-query understanding model comprises:

determining the correlation between the target image area and the target phrase by obtaining an inner product between a first vector representing the target image area and a second vector representing the target phrase.

5. The method of claim 4 , wherein retraining the image-query understanding model further comprises:

establishing multiple weights indicating an importance of projections of multiple vectors to each other, the multiple vectors including the first vector and the second vector.

6. The method of claim 1 , wherein:

the robot comprises a cleaning robot; and

the image-query understanding model corresponds to operations of the cleaning robot.

7. The method of claim 6 , wherein the retrained image-query understanding model is configured to be used by the cleaning robot in an inference mode that does not include use of a target phrase.

8. An electronic device for controlling a robot, the electronic device comprising:

at least one memory configured to store instructions; and

at least one processing device configured when executing the instructions to:

obtain an image-query understanding model;

obtain an image and a user query associated with the image, wherein the image comprises a target image area and the user query comprises a target phrase, and wherein the target image area is marked on the image by a user and the target phrase is identified within the user query by the user during operation of the robot; and

retrain the image-query understanding model using a correlation between the target image area and the target phrase to obtain a retrained image-query understanding model;

wherein, to retrain the image-query understanding model, the at least one processing device is configured to determine (i) one or more query-level contextual features, (ii) one or more question features, (iii) one or more image-level contextual features, and (iv) one or more post-processed image features based on at least one weighted attention function.

9. The electronic device of claim 8 , wherein:

the at least one processing device is configured to execute a question contextual feature extraction (Q-CFE) module to determine the one or more query-level contextual features and the one or more question features; and

the at least one processing device is configured to execute an image contextual feature extraction (I-CFE) module to determine the one or more image-level contextual features and the one or more post-processed image features.

10. The electronic device of claim 9 , wherein the at least one processing device is further configured to execute a weighted contextual feature question-image understanding (WCUQIU) module to process (i) the one or more query-level contextual features and the one or more question features from the Q-CFE module and (ii) the one or more image-level contextual features and the one or more post-processed image features from the I-CFE module.

11. The electronic device of claim 8 , wherein, to retrain the image-query understanding model, the at least one processing device is configured to determine the correlation between the target image area and the target phrase by obtaining an inner product between a first vector representing the target image area and a second vector representing the target phrase.

12. The electronic device of claim 11 , wherein, to retrain the image-query understanding model, the at least one processing device is further configured to establish multiple weights indicating an importance of projections of multiple vectors to each other, the multiple vectors including the first vector and the second vector.

13. The electronic device of claim 8 , wherein:

the robot comprises a cleaning robot; and

the image-query understanding model corresponds to operations of the cleaning robot.

14. The electronic device of claim 13 , wherein the retrained image-query understanding model is configured to be used by the cleaning robot in an inference mode that does not include use of a target phrase.

15. A non-transitory machine-readable medium containing instructions that when executed cause at least one processor of a robot to:

obtain an image-query understanding model;

obtain an image and a user query associated with the image, wherein the image comprises a target image area and the user query comprises a target phrase, and wherein the target image area is marked on the image by a user and the target phrase is identified within the user query by the user during operation of the robot; and

retrain the image-query understanding model using a correlation between the target image area and the target phrase to obtain a retrained image-query understanding model;

wherein the instructions that when executed cause the at least one processor to retrain the image-query understanding model comprise:

instructions that when executed cause the at least one processor to determine (i) one or more query-level contextual features, (ii) one or more question features, (iii) one or more image-level contextual features, and (iv) one or more post-processed image features based on at least one weighted attention function.

16. The non-transitory machine-readable medium of claim 15 , wherein the image-query understanding model comprises:

a question contextual feature extraction (Q-CFE) module configured to determine the one or more query-level contextual features and the one or more question features; and

an image contextual feature extraction (I-CFE) module configured to the one or more image-level contextual features and the one or more post-processed image features.

17. The non-transitory machine-readable medium of claim 16 , wherein the image-query understanding model further comprises a weighted contextual feature question-image understanding (WCUQIU) module configured to process (i) the one or more query-level contextual features and the one or more question features from the Q-CFE module and (ii) the one or more image-level contextual features and the one or more post-processed image features from the I-CFE module.

18. The non-transitory machine-readable medium of claim 15 , wherein the instructions that when executed cause the at least one processor to retrain the image-query understanding model comprise:

instructions that when executed cause the at least one processor to determine the correlation between the target image area and the target phrase by obtaining an inner product between a first vector representing the target image area and a second vector representing the target phrase.

19. The non-transitory machine-readable medium of claim 18 , wherein the instructions that when executed cause the at least one processor to retrain the image-query understanding model comprise:

instructions that when executed cause the at least one processor to establish multiple weights indicating an importance of projections of multiple vectors to each other, the multiple vectors including the first vector and the second vector.

20. The non-transitory machine-readable medium of claim 15 , wherein:

the robot comprises a cleaning robot; and

the image-query understanding model corresponds to operations of the cleaning robot.

21. A method comprising:

training, using at least one processor of a robot, an image-query understanding model based on a correlation between a target image area of a first image and a target phrase of a user query, wherein the target image area is marked on the first image by a user and the target phrase is identified within the user query by the user during operation of the robot;

obtaining, using the at least one processor, a user instruction to perform an operation;

obtaining, using the at least one processor, a second image showing a first area to perform the operation;

determining, using the at least one processor and the trained image-query understanding model, a second area that is a subset of the first area based on the user instruction; and

performing, using the robot, the operation for the second area according to the user instruction;

wherein the image-query understanding model comprises (i) one or more query-level contextual features, (ii) one or more question features, (iii) one or more image-level contextual features, and (iv) one or more post-processed image features determined based on the target image area and the target phrase using weights learned during the training.

22. The method of claim 21 , wherein the second area includes at least one feature that is different from a feature of the target image area of the first image.

23. The method of claim 22 , wherein the image-query understanding model comprises:

a question contextual feature extraction (Q-CFE) module associated with the one or more query-level contextual features and the one or more question features; and

an image contextual feature extraction (I-CFE) module associated with the one or more image-level contextual features and the one or more post-processed image features.

24. The method of claim 23 , wherein the image-query understanding model further comprises a weighted contextual feature question-image understanding (WCUQIU) module configured to process (i) the one or more query-level contextual features and the one or more question features from the Q-CFE module and (ii) the one or more image-level contextual features and the one or more post-processed image features from the I-CFE module.

25. The method of claim 21 , wherein:

the robot comprises a cleaning robot; and

the operation comprises a cleaning operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2021
From: WANG, YU; SHEN, YILIN; JIN, HONGXIA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 055936/0264 →
Continuity (2)
Provisional Application 63017887 · Apr 30, 2020
Related Publication 20210342624A1 · Nov 4, 2021
Cited By (1)
US 12,210,835