IP Library Granted Patent US 10,664,060
Granted Patent B2
US 10,664,060 · App. 16/044,335 · Granted May 26, 2020

Multimodal input-based interaction method and device

Inventors: Chunyuan Liao (Shanghai, CN); Rongxing Tang (Shanghai, CN); Mei Huang (Shanghai, CN)
Assignee: HISCENE INFORMATION TECHNOLOGY CO., LTD.
G06F3/017G06F3/01G06F3/012G06F3/167G06K9/00355G06K9/6254G06K9/6273G06K9/6289G10L15/265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,664,060
App. No.
16/044,335
Granted
May 26, 2020
Kind
B2
Abstract

An object of the present disclosure is to provide a method for interacting based on multimodal inputs, which enables a higher approximation to user natural interaction, comprising: acquiring a plurality of input information from at least one of a plurality of input modules; performing comprehensive logic analysis of the plurality of input information so as to generate an operation command, wherein the operation command has operation elements, the operation elements at least including an operation object, an operation action, and an operation parameter; and performing a corresponding operation on the operation object based on the operation command.

Claims (34)

1. A method for a smart eyewear apparatus to interact based on multimodal inputs, comprising:

acquiring a plurality of input information from at least one of a plurality of input modules, the plurality of input modules including: an image input module, a voice input module, a touch input module, and a sensing input module, the plurality of input information including at least any one of: real scene information, virtual scene information, gesture information, voice information, touch information, and sensing information;

performing, using corresponding processing modules, recognition pre-processing to the plurality of input information of the input modules respectively to generate a plurality of structured data, wherein the processing modules include a scene image recognition module, a gesture recognition module, a voice recognition module, a touch recognition module, or a sensing recognition module; and

determining element types corresponding to the structured data; performing logic matching and/or arbitration selection to the structured data of a same element type to determine element information of the operation element corresponding to the element type; and if a combination of element information of the operation elements corresponding to the determined different element types complies with executing business logic, generating the operation command based on the element information of the corresponding operation element, wherein the operation command has operation elements, the operation elements at least including an operation object, an operation action, and an operation parameter; and

performing a corresponding operation on the operation object based on the operation command.

2. The method according to claim 1 , wherein performing, using corresponding processing modules, recognition pre-processing to the plurality of input information of the input modules respectively to generate a plurality of structured data, comprises at least any one of:

recognizing, using the scene image recognition module, the virtual scene information and/or the real scene information inputted by the image input module to obtain structured data of a set of operable objects;

recognizing, using the gesture recognition module, the gesture information inputted by the image input module to obtain a structured data of a set of operable objects and/or structured data of a set of operable actions;

recognizing, using the touch recognition module, the touch information inputted by the touch input module to obtain at least any one of the following structured data: structured data of a position of a cursor on a screen, structured data of a set of operable actions, and structured data of input parameters; and

recognizing, using the voice recognition module, the voice information inputted by the voice input module to obtain at least any one of the following structured data: structured data of a set of operable objects, structured data of a set of operable actions, and structured data of input parameters.

3. The method according to claim 1 , wherein performing logic matching and/or arbitration selection to the structured data of a same element type to determine element information of the operation element corresponding to the element type, comprises:

performing logic matching to the structured data of the same element type to determine at least one to-be-selected element information;

performing arbitration selection to the to-be-selected element information to select one of them as selected element information; and

determining the element information of the operation element corresponding to the element type based on the selected element information.

4. The method according to claim 3 , further comprises:

re-performing arbitration selection to the remaining to-be-selected element information so as to reselect one of them as selected element information when a combination of the element information of the determined operation elements corresponding to the different element types does not comply with executing business logic; and

clearing the element information of the operation elements corresponding to all operation types when the duration of reselection exceeds an overtime or all of the combination of the element information determined for the to-be-selected element information does not comply with executing business logic.

5. The method according to claim 3 , wherein performing arbitration selection to the to-be-selected element information to select one of them as selected element information, comprises:

performing contention selection based on time orders and/or priority rankings of the to-be-selected element information; when the time orders and priority rankings of the to-be-selected element information are both identical, performing random selection to select one of them as the selected element information.

6. The method according to claim 3 , wherein determining the element information of the operation element corresponding to the element type based on the selected element information, comprises:

determining whether there currently exists element information of the operation element corresponding to the element type;

in the case of existence, determining whether the priority of the selected element information is higher than the existing element information; and

if yes, replacing the existing element information with the selected element information and determining the selected element information as the element information of the operation element corresponding to the element type.

7. The method according to claim 1 , further comprises:

performing logic matching and arbitration selection to all of the structured data using a machine learning method so as to determine element information of the operation element corresponding to each of the element types, wherein the machine learning method includes at least one of: a decision tree method, a random forest method, and a convolutional neural network method.

8. The method according to claim 1 , further comprises:

creating a deep learning neural network architecture model; and

inputting raw data of the input information into the deep learning neural network architecture model so as to be subjected to fusion processing and model operation, thereby generating an operation command.

9. The method according to claim 8 , wherein the deep learning neural network architecture model is a convolutional neural network architecture model.

10. The method according to claim 1 , further comprises:

transmitting the plurality of input information to a split-mount control device to perform comprehensive logic analysis so as to generate the operation command, wherein the split-mount control device is physically separated from a body of the smart eyewear apparatus and is in communication connection with the smart eyewear apparatus in a wired or wireless manner.

11. The method according to claim 1 , further comprising:

acquiring relevant information of an operation command to be set by the user and updating the operation command based on the relevant information of the to-be-set operation command.

12. A non-transitory computer readable storage medium, including computer code, which, when being executed, causes a method according to claim 1 to be executed.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2018
From: LIAO, CHUNYUAN; TANG, RONGXING; HUANG, MEI
To: HISCENE INFORMATION TECHNOLOGY CO., LTD
Reel/Frame 046448/0051 →
Priority Claims (1)
CN 2016 1 0049586 · Jan 25, 2016 · national
Continuity (2)
Continuation PCTCN2017078225 · Mar 25, 2017
Related Publication 20180329512A1 · Nov 15, 2018