IP Library › Granted Patent US 12,333,831
Granted Patent B2
US 12,333,831 · App. 17/566,995 · Granted Jun 17, 2025

Image based command classification and task engine for a computing system

Inventors: Bharath Cheluvaraju (Bangalore, IN); Saheel Ram Godhane (Warud, IN); Nishan Hassan (Delhi, IN); Sheetal Shamsher Sethi (Hyderabad, IN)
Assignee: Microsoft Technology Licensing, LLC
G06V30/10G06F3/048G06F3/0481G06F3/0482G06N3/0455G06V10/40G06V10/82G06V30/413G06N3/048G06N3/09G06N7/01G06Q10/10G06V10/454G06V20/20G06V20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,831
App. No.
17/566,995
Filed
Dec 31, 2021
Granted
Jun 17, 2025
Kind
B2
Art Unit
2662
USPC
382/190
Abstract

Provided are methods, systems, and computer storage media for determining a command (e.g., intent) of an image based on image data features. A task associated with the determined command is generated based on a portion of the image data features. Task entities corresponding to the task are determined. The task and the corresponding task entities are generated and configured for use in a computer productivity application. Accordingly, present embodiments provide for generating command-specific tasks and task entities that may be integratable for use in a computer productivity application.

Claims (52)

1. A computerized system, the computerized system comprising:

at least one computer processor; and

computer memory storing computer-readable instructions that, when used by the at least one computer processor, cause the at least one computer processor to perform operations comprising:

receiving an image comprising at least an indication of non-alphanumeric-character objects;

extracting image data for the image based on at least the non-alphanumeric-character objects to produce a plurality of image data features;

determining, by a machine learning model, an image context based on a first image data feature of the plurality of image data features, wherein the image context defines a relationship between alphanumeric characters depicted in the image and the non-alphanumeric-character objects;

determining a command associated with the image based on a second image data feature in the plurality of image data features;

determining at least one task corresponding to the command based on at least one of: the command, the first image data feature, the image context, and the second image data feature;

based on the at least one task and the plurality of image data features, determining at least one task entity; and

generating the at least one task, comprising the at least one task entity and configured for use in a computer productivity application.

2. The computerized system of claim 1 , wherein the command is determined based on the machine learning model that is trained using a plurality of images comprising the alphanumeric characters and the non-alphanumeric-character objects.

3. The computerized system of claim 1 , wherein the image further comprises an indication of the alphanumeric characters, and wherein the plurality of image data features comprise visual features associated with text corresponding to the alphanumeric characters and spatial features associated with coordinates associated with and relating the alphanumeric characters and the non-alphanumeric-character objects.

4. The computerized system of claim 1 , wherein the image further comprises an indication of the alphanumeric characters, and the operations further comprising determining a corresponding coordinate set for each alphanumeric character and for each non-alphanumeric-character object, wherein the command is determined based on the second image data feature indicative of a comparison of each corresponding coordinate set to one another.

5. The computerized system of claim 1 , wherein the operations further comprise communicating the at least one task to an application layer of a computing device, wherein communicating the at least one task to the application layer comprises integrating the at least one task or the at least one task entity into a software application.

6. The computerized system of claim 5 , wherein the command comprises a calendar event, the at least one task entity comprises a calendar event description, the at least one task comprises creating the calendar event integratable with the software application, wherein communicating the at least one task to the application layer comprises adding the calendar event to an electronic calendar associated with the software application.

7. The computerized system of claim 1 , wherein the operations further comprise:

determining a second command of the image based on extracted text and extracted coordinates being applied to the machine learning model;

determining a second task based on at least one of the second command;

determining a second task entity associated with the second task based on the extracted text, the extracted coordinates, or both; and

generating the second task, comprising the second task entity.

8. The computerized system of claim 1 , wherein determining the command of the image comprises determining that the image comprises at least one of: a recipe, a scheduled event, a list of items.

9. The computerized system of claim 1 , wherein the image comprises a digital video or computer animation, and wherein the plurality of image data features comprise time-sensitive spatial image features and visual image features.

10. At least one computer-storage media having computer-executable instructions embodied thereon that, when executed by a computing system having a processor and memory, cause the processor to:

receive an image comprising alphanumeric characters and non-alphanumeric-character objects;

extract image data for the image based on the alphanumeric characters and the non-alphanumeric-character objects to produce a plurality of image data features;

determining a relationship between the alphanumeric characters and the non-alphanumeric-character objects based on a first image data feature of the plurality of image data features;

determine a command associated with the image based on a second image data feature in the plurality of image data features;

determine at least one task corresponding to the command based on at least one of: the command, the first image data feature, the second image data feature, and the relationship between the alphanumeric characters and the non-alphanumeric-character objects;

based on the at least one task and the plurality of image data features, determine at least one task entity; and

generate the at least one task, comprising the at least one task entity and configured for use in a computer productivity application.

11. The computer-storage media of claim 10 , wherein the plurality of image data features comprise visual features associated with text corresponding to the alphanumeric characters and spatial features associated with coordinates associated with and relating the alphanumeric characters and the non-alphanumeric-character objects.

12. The computer-storage media of claim 10 , wherein the instructions further cause the processor to determine a corresponding coordinate set for each alphanumeric character and for each non-alphanumeric-character object, wherein the command is determined based on the second image data feature indicative of a comparison of each corresponding coordinate set to one another.

13. The computer-storage media of claim 10 , wherein the instructions further cause the processor to communicate the at least one task to an application layer of a computing device, wherein communicating the at least one task to the application layer comprises integrating the at least one task or the at least one task entity into a software application.

14. The computer-storage media of claim 13 , wherein the command comprises a calendar event, the at least one task entity comprises an event description, the at least one task comprises the calendar event integratable with the software application, wherein communicating the at least one task to the application layer comprises appending the calendar event to the software application.

15. The computer-storage media of claim 10 , wherein the instructions further cause the processor to:

determine a second command of the image based on extracted text and extracted coordinates being applied to a machine learning model;

determine a second entity associated with the second command based on the extracted text, the extracted coordinates, or both; and

generate a second task, comprising the second entity, based on the second command.

16. A computer-implemented method, comprising:

receiving an image comprising alphanumeric characters and non-alphanumeric-character objects;

extracting, using a machine learning model, (i) text from the alphanumeric characters and (ii) spatial data from the alphanumeric characters and the non-alphanumeric-character objects, wherein (i) the text and (ii) the spatial data are associated with the machine learning model that is trained based on visual features associated with the text and based on spatial features associated with the spatial data;

determining, by the machine learning model, an image context based on the visual features and the spatial features, wherein the image context defines a relationship between the alphanumeric characters and the non-alphanumeric-character objects depicted in the image;

determining a command of the image based on the extracted text and the extracted spatial data being applied to the machine learning model;

determining a task associated with the command based on at least one of: the extracted text, the extracted spatial data, the command, and the image context;

determining at least one task entity corresponding to the task based on at least one of the extracted text, the extracted spatial data, the task, or the command; and

generating the task, comprising the at least one task entity.

17. The computer-implemented method of claim 16 , further comprising:

determining a second command of the image based on the extracted text and the extracted spatial data being applied to the machine learning model;

determining a second task entity associated with the second command based on the extracted text, the extracted spatial data, or both; and

generating a second task comprising the second task entity.

18. The computer-implemented method of claim 16 , further comprising communicating the task to an application layer of a computing device, wherein communicating the task to the application layer comprises integrating, into a software application, at least one of the task or the at least one task entity.

19. The computer-implemented method of claim 18 , wherein the command comprises a calendar event, the at least one task entity comprises an event date, the task comprises creating the calendar event integratable with the software application, wherein communicating the task to the application layer comprises adding the calendar event to an electronic calendar associated with the software application.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2022
From: CHELUVARAJU, BHARATH; HASSAN, NISHAN; GODHANE, SAHEEL RAM; SETHI, SHEETAL SHAMSHER
To: MICROSOFT TECHNOLOGY LICENSING, LLC,
Reel/Frame 059892/0016 →
Continuity (1)
Related Publication 20230215198A1 · Jul 6, 2023
References Cited (40)
US 9911034B2 · Chulinin · 2018 [cited by applicant]
US 9923991B2 · Lai et al. · 2018 [cited by applicant]
US 10223586B1 · Leibovitz · 2019 [cited by examiner]
US 10860854B2 · Anorga et al. · 2020 [cited by applicant]
US 10963505B2 · Adlersberg et al. · 2021 [cited by applicant]
US 11087086B2 · An et al. · 2021 [cited by applicant]
US 20120304060A1 · Kompalli · 2012 [cited by examiner]
US 20170064035A1 · Lai · 2017 [cited by examiner]
US 20180336415A1 · Anorga · 2018 [cited by examiner]
US 20200082345A1 · MacBeth · 2020 [cited by examiner]
US 20210192202A1 · Tripuraneni · 2021 [cited by examiner]
CN 112101165A · 2020 [cited by applicant]
WO WO2017205305A1 · 2017 [cited by examiner]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/045957”, Mailed Date: Feb. 7, 2023, 13 Pages. [cited by applicant]
“ATIS DataSet from MS CNTK”, Retrieved From: http://web.archive.org/web/20210126225547/https://www.kaggle.com/siddhadev/atis-dataset-from-ms-cntk, Jan. 26, 2021, 2 Pages. [cited by applicant]
“CONLL Corpora”, Retrieved From: http://web.archive.org/web/20210124200202/https://www.kaggle.com/www.kaggle.com/nltkdata/conll-corpora, Jan. 24, 2007, 2 Pages. [cited by applicant]
“CRF Layer on the Top of BILSTM-1”, Retrieved From: https://createmomo.github.io/2017/09/12/CRF_Layer_on_the_Top_of_BiLSTM_1/, Sep. 12, 2017, 7 Pages. [cited by applicant]
“How to use OCR”, Retrieved From: https://help.thecookbookapp.com/hc/en-gb/articles/360002691375-How-to-use-OCR, Feb. 25, 2021, 5 Pages. [cited by applicant]
“Language Understanding (LUIS) Documentation”, Retrieved From: https://docs.microsoft.com/en-us/azure/cognitive-services/luis/, Jan. 26, 2021, 545 Pages. [cited by applicant]
“Machine Learning for Mobile Developers”, Retrieved From: http://web.archive.org/web/20210918011811/https:/developers.google.com/ml-kit, Sep. 18, 2021, 6 Pages. [cited by applicant]
“Named Entity Recognition: Concept, Tools and Tutorial”, Retrieved From: https://monkeylearn.com/blog/named-entity-recognition/, Oct. 21, 2021, 13 Pages. [cited by applicant]
“Natural Language Toolkit”, Retrieved From: https://www.nltk.org/, Retrieved Date: Nov. 17, 2018, 2 Pages. [cited by applicant]
“Scikit-Learn: Machine Learning in Python”, Retrieved From: https://scikit-learn.org/stable/, Jan. 3, 2021, 2 Pages. [cited by applicant]
“Sklearn Feature Extraction Text.TfidfVectorizer”, Retrieved From: https://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.TfidfVectorizer.html, Jan. 20, 2021, 7 Pages. [cited by applicant]
“Sklearn Model Selection GridSearchCV”, Retrieved From: https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GridSearchCV.html?highlight=gridsearchcv#sklearn.model_selection.GridSearchCV, Dec. 5, 20… [cited by applicant]
“Sklearn Multioutput ClassifierChain”, Retrieved From: https://scikit-learn.org/stable/modules/generated/sklearn.multioutput.ClassifierChain.html?highlight=classifierchain#/sklearn.multioutput.ClassifierChain, 4 Pages. [cited by applicant]
“sklearn.ensemble. RandomForestClassifier”, Retrieved From: http://web.archive.org/web/20210308073211/https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestClassifier.html, Mar. 8, 2021, 7 Pages. [cited by applicant]
“Sklearn.Svm.SVC”, Retrieved From: https://scikit-learn.org/stable/modules/generated/sklearn.svm.SVC.html?highlight=svc#sklearn.svm.SVC, Jan. 22, 2021, 8 Pages. [cited by applicant]
“The Easiest Way to Organize Your Recipes”, Retrieved From: https://recipekeeperonline.com/, Jan. 19, 2021, 13 Pages. [cited by applicant]
“Viterbi Algorithm”, Retrieved From: Viterbi Algorithm, 2016, 14 Pages. [cited by applicant]
Chen, et al., “BERT for Joint Intent Classification and Slot Filling”, In Repository of arXiv:1902.10909v1, Feb. 28, 2019, 6 Pages. [cited by applicant]
Li, et al., “Dice Loss for Data-imbalanced NLP Tasks”, Retrieved From: In Repository of arXiv:1911.02855v3, Aug. 29, 2020, 12 Pages. [cited by applicant]
Lin, et al., “Focal Loss for Dense Object Detection”, In Repository of arXiv:1708.02002v2, Feb. 7, 2018, 10 Pages. [cited by applicant]
Marcus, Stacy, “Create a Task by Taking a Photo”, Retrieved From: https://help.workframe.com/en/articles/2020470-create-a-task-by-taking-a-photo, Oct. 11, 2021, 4 Pages. [cited by applicant]
Rosebrock, Adrian, “OCR a document, form, or invoice with Tesseract, OpenCV, and Python”, Retrieved From: https://www.pyimagesearch.com/2020/09/07/ocr-a-document-form-or-invoice-with-tesseract-opencv-and-python/, Sep. 7… [cited by applicant]
Sean, Endicott, “You Can Now Create a Task in Microsoft to Do by Sharing an Image on Android”, Retrieved From: https://www.windowscentral.com/you-can-now-create-task-microsoft-do-android-sharing-image, May 6, 2021, 8 Pa… [cited by applicant]
Wisley, Connie, “6 Best Receipt OCR Software for Easy Receipt Filing and Tracking”, Retrieved From: https://www.cisdem.com/resource/receipt-ocr.html, Feb. 28, 2020, 9 Pages. [cited by applicant]
Xu, Yiheng, “LayoutLM: Pre-training of Text and Layout for Document Image Understanding”, In Repository of arXiv:1912.13318v5, Jun. 16, 2020, 9 Pages. [cited by applicant]
Zelic, Filip, “A comprehensive guide to OCR with Tesseract, OpenCV and Python”, Retrieved From: https://hanonets.com/blog/ocr-with-tesseract/, Jan. 8, 2021, 42 Pages. [cited by applicant]
“The Cookbook App”, Retrieved From: https://thecookbookapp.com/, Jan. 27, 2021, 4 Pages. [cited by applicant]