IP Library › Granted Patent US 11,694,021
Granted Patent B2
US 11,694,021 · App. 16/850,473 · Granted Jul 4, 2023

Apparatus for generating annotated image information using multimodal input data, apparatus for training an artificial intelligence model using annotated image information, and methods thereof

Inventor: Federico Fancellu (Toronto, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06F40/169G06F3/0481G06F18/214G06F40/205G06F40/30G06V10/945G06V30/19147G06V30/19173G06V30/274
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,694,021
App. No.
16/850,473
Granted
Jul 4, 2023
Kind
B2
Abstract

A method for providing a user interface (UI) for generating training data for an artificial intelligence (AI) model may include providing, for display via the UI, image information that depicts an object, a set of operations of the object, and a process associated with the set of operations. The method may include providing, for display via the UI, text information that describes the object, the set of operations of the object, and the process associated with the set of operations. The method may include receiving, via the UI, a user input that associates respective image information of the image information with corresponding text information of the text information. The method may include generating association information that associates the respective image information with the corresponding text information, based on the user input. The method may include generating discourse and semantic information from the text information associated to the image information.

Claims (57)

1. A device for providing a user interface (UI) for generating training data for an artificial intelligence (AI) model, the device comprising:

a memory configured to store instructions; and

a processor configured to execute the instructions to:

provide, for display via the UI, image information that depicts an object, a set of operations of the object, and a process associated with the set of operations;

provide, for display via the UI, text information that describes the object, the set of operations of the object, and the process associated with the set of operations;

receive, via the UI, a first touch input in association with the image information, and a second touch input in association with the text information to associate respective image information of the image information with corresponding text information of the text information; and

generate association information that associates the respective image information with the corresponding text information, based on the first touch input and the second touch input,

wherein the image information is associated with an instructional video regarding the object, and wherein the text information corresponds to at least one of a product manual associated with the object, or captions associated with the instructional video.

2. The device of claim 1 , wherein the processor is further configured to:

receive, via the UI, discourse parsing information that identifies a relation between image information of the association information; and

generate annotated image information based on the discourse parsing information.

3. The device of claim 2 , wherein the processor is further configured to:

input the annotated image information into the AI model as training data for the AI model to permit the AI model to identify the relation between the image information of the association information.

4. The device of claim 1 , wherein the processor is further configured to:

receive, via the UI, semantics parsing information that provides the text information of the association information in a machine-understandable format; and

generate annotated image information based on the semantics parsing information.

5. The device of claim 4 , wherein the processor is further configured to:

input the annotated image information into the AI model as training data for the AI model to permit the AI model to convert the text information of the association information to the machine-understandable format.

6. The device of claim 1 , wherein the processor is further configured to:

input the association information into the AI model as training data for the AI model to permit the AI model to associate the respective image information with the corresponding text information.

7. A method for providing a user interface (UI) for generating training data for an artificial intelligence (AI) model, the method comprising:

providing, for display via the UI, image information that depicts an object, a set of operations of the object, and a process associated with the set of operations;

providing, for display via the UI, text information that describes the object, the set of operations of the object, and the process associated with the set of operations;

receiving, via the UI, a first touch input in association with the image information, and a second touch input in association with the text information to associate respective image information of the image information with corresponding text information of the text information; and

generating association information that associates the respective image information with the corresponding text information, based on the first touch input and the second touch input,

wherein the image information is associated with an instructional video regarding the object, and wherein the text information corresponds to at least one of a product manual associated with the object, or captions associated with the instructional video.

8. The method of claim 7 , further comprising:

receiving, via the UI, discourse parsing information that identifies a relation between image information of the association information; and

generating annotated image information based on the discourse parsing information.

9. The method of claim 8 , further comprising:

inputting the annotated image information into the AI model as training data for the AI model to permit the AI model to identify the relation between the image information of the association information.

10. The method of claim 7 , further comprising:

receiving, via the UI, semantics parsing information that provides the text information of the association information in a machine-understandable format; and

generating annotated image information based on the semantics parsing information.

11. The method of claim 10 , further comprising:

inputting the annotated image information into the AI model as training data for the AI model to permit the AI model to convert the text information of the association information to the machine-understandable format.

12. The method of claim 7 , further comprising:

inputting the association information into the AI model as training data for the AI model to permit the AI model to associate the respective image information with the corresponding text information.

13. A non-transitory computer-readable medium storing instructions, the instructions comprising:

one or more instructions that, when executed by one or more processors of a device for training an artificial intelligence (AI) model, cause the one or more processors to:

provide, for display via the UI, image information that depicts an object, a set of operations of the object, and a process associated with the set of operations;

provide, for display via the UI, text information that describes the object, the set of operations of the object, and the process associated with the set of operations;

receive, via the UI, a first touch input in association with the image information, and a second touch input in association with the text information to associate respective image information of the image information with corresponding text information of the text information; and

generate association information that associates the respective image information with the corresponding text information, based on the first touch input and the second touch input,

wherein the image information is associated with an instructional video regarding the object, and

wherein the text information corresponds to at least one of a product manual associated with the object or captions associated with the instructional video.

14. The non-transitory computer-readable medium of claim 13 , wherein the one or more instructions further cause the one or more processors to:

receive, via the UI, discourse parsing information that identifies a relation between image information of the association information; and

generate annotated image information based on the discourse parsing information.

15. The non-transitory computer-readable medium of claim 14 , wherein the one or more instructions further cause the one or more processors to:

receive, via the UI, semantics parsing information that provides the text information of the association information in a machine-understandable format; and

input the annotated image information and the semantics parsing information into the AI model as training data for the AI model to permit the AI model to identify the relation between the image information of the association information, and to permit the AI model to convert the text information of the association information to the machine-understandable format.

16. The non-transitory computer-readable medium of claim 13 , wherein the one or more instructions further cause the one or more processors to:

receive, via the UI, semantics parsing information that provides the text information of the association information in a machine-understandable format; and

generate annotated image information based on the semantics parsing information.

17. The non-transitory computer-readable medium of claim 13 , wherein the one or more instructions further cause the one or more processors to:

input the association information into the AI model as training data for the AI model to permit the AI model to associate the respective image information with the corresponding text information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2020
From: FANCELLU, FEDERICO
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 052418/0083 →
Continuity (1)
Related Publication 20210326643A1 · Oct 21, 2021