IP Library Granted Patent US 12,469,292
Granted Patent B2
US 12,469,292 · App. 18/205,698 · Granted Nov 11, 2025

System and method for implementing a multimodal assistant using large language models

Inventors: Viral Chawda (Frisco, TX); Charles Scott McAllister (Silver Spring, MD)
Assignee: KPMG LLP
G06V20/50G06F3/0484G06F16/532G06F40/40G06V10/26G06V10/945G06V20/63G06V30/19
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,469,292
App. No.
18/205,698
Granted
Nov 11, 2025
Kind
B2
Abstract

An embodiment of the present invention is directed to a multimodal assistant for mechanics and technicians in military warehouses using large language models. An exemplary system provides a conversational assistant to mechanics and technicians working in challenging environments, such as shop or warehouse floors with heavy industrial equipment and componentry. The innovative system may utilize multimodal large language models as well as image segmentation techniques to efficiently retrieve relevant information from asset reference materials, documentation and other sources of instructional information. The system may also use multimodal large language models to extract information from the retrieved documents and/or other data sources to provide guidance on performing discrete tasks related to the asset, such as routine maintenance, part replacement, services and/or other related actions.

Claims (41)

1 . A computer-implemented system for providing a conversational assistant using multimodal large language models and image segmentation, the system comprising:

an interactive user interface that is configured to receive one or more inputs;

a database interface that communicates with a database that stores and manages asset data; and

a processor executing on a mobile device and coupled to the interface and the database interface, the processor further configured to perform the steps of:

receiving, via the interactive user interface, a user query and a scan of an image associated with an asset;

detecting, via a computer vision detection model, a text from the scan wherein the text is imprinted on the asset and the asset is a customized asset without a conventional serial number;

performing, via a prediction model, text recognition of the text and identifying one or more predicted texts with corresponding confidence levels;

displaying, via the interactive user interface executing on the mobile device, the one or more predicted texts comprising an asset identifier;

identifying, via an asset lookup agent applying a segmentation model, a plurality of objects associated with the asset identifier wherein each object of the plurality of objects is fed to a lookup model to identify each object against an internal asset database;

for a selected object, identifying, via the interactive user interface, a corresponding service action;

retrieving, via a reference lookup agent, a set of reference data associated with the selected object;

based on the identified service action and the selected object, generating, via an advice agent, a set of instructions and corresponding set of tools; and

providing, via the interactive user interface, a response to the user query wherein the response comprises the set of instructions and the corresponding set of tools.

2 . The system of claim 1 , wherein the asset lookup agent relies on pre-processing of the user query for consistency in one or more agent prompts.

3 . The system of claim 1 , wherein the corresponding service action comprises: repair, replace and maintenance.

4 . The system of claim 1 , wherein the set of reference data comprises one or more documents from a document database that communicates with a client inventory management system via an integration API.

5 . The system of claim 1 , wherein the set of reference data comprises instructional information that relate directly and indirectly to the asset.

6 . The system of claim 1 , wherein the response comprises an estimated time required for the identified service action.

7 . The system of claim 1 , wherein the set of instructions comprises multimedia content comprising a combination of: audio, video, images and text.

8 . The system of claim 1 , wherein the response comprises an inventory status for at least one tool of the corresponding set of tools.

9 . The system of claim 1 , wherein at least one of the asset lookup agent, the reference lookup agent and the advice agent applies a multimodal large language model.

10 . The system of claim 1 , wherein the asset comprises military machinery.

11 . A computer-implemented method for providing a conversational assistant using multimodal large language models and image segmentation, the method comprising the steps of:

receiving, via an interactive user interface, a user query and a scan of an image associated with an asset;

detecting, via a computer vision detection model, a text from the scan wherein the text is imprinted on the asset and the asset is a customized asset without a conventional serial number;

performing, via a prediction model, text recognition of the text and identifying one or more predicted texts with corresponding confidence levels;

displaying, via the interactive user interface executing on the mobile device, the one or more predicted texts comprising an asset identifier;

identifying, via an asset lookup agent applying a segmentation model, a plurality of objects associated with the asset identifier wherein each object of the plurality of objects is fed to a lookup model to identify each object against an internal asset database;

for a selected object, identifying, via the interactive user interface, a corresponding service action;

retrieving, via a reference lookup agent, a set of reference data associated with the selected object;

based on the identified service action and the selected object, generating, via an advice agent, a set of instructions and corresponding set of tools; and

providing, via the interactive user interface, a response to the user query wherein the response comprises the set of instructions and the corresponding set of tools.

12 . The method of claim 11 , wherein the asset lookup agent relies on pre-processing of the user query for consistency in one or more agent prompts.

13 . The method of claim 11 , wherein the corresponding service action comprises: repair, replace and maintenance.

14 . The method of claim 11 , wherein the set of reference data comprises one or more documents from a document database that communicates with a client inventory management system via an integration API.

15 . The method of claim 11 , wherein the set of reference data comprises instructional information that relate directly and indirectly to the asset.

16 . The method of claim 11 , wherein the response comprises an estimated time required for the identified service action.

17 . The method of claim 11 , wherein the set of instructions comprises multimedia content comprising a combination of: audio, video, images and text.

18 . The method of claim 11 , wherein the response comprises an inventory status for at least one tool of the corresponding set of tools.

19 . The method of claim 11 , wherein at least one of the asset lookup agent, the reference lookup agent and the advice agent applies a multimodal large language model.

20 . The method of claim 11 , wherein the asset comprises military machinery.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2023
From: CHAWDA, VIRAL; MCALLISTER, CHARLES SCOTT
To: KPMG LLP
Reel/Frame 063852/0616 →
Continuity (3)
Continuation In Part 17659286 · Apr 14, 2022
Provisional Application 63265167 · Dec 9, 2021
Related Publication 20230326212A1 · Oct 12, 2023
References Cited (22)
US 7106905B2 · Simske · 2006 [cited by applicant]
US 10992606B1 · Mitchell · 2021 [cited by applicant]
US 11049063B2 · Ali et al. · 2021 [cited by applicant]
US 11269883B2 · Reed · 2022 [cited by applicant]
US 11295783B2 · Shen et al. · 2022 [cited by applicant]
US 11550299B2 · Cella · 2023 [cited by examiner]
US 11580348B2 · Volkerink et al. · 2023 [cited by applicant]
US 11682025B2 · McLaney · 2023 [cited by examiner]
US 12001917B2 · Brebner · 2024 [cited by examiner]
US 20090146832A1 · Ebert et al. · 2009 [cited by applicant]
US 20160069644A1 · Bell · 2016 [cited by applicant]
US 20180137349A1 · Such et al. · 2018 [cited by applicant]
US 20200311658A1 · Gundel et al. · 2020 [cited by applicant]
US 20200356254A1 · Missig et al. · 2020 [cited by applicant]
US 20210056499A1 · Jacobus et al. · 2021 [cited by applicant]
US 20210182773A1 · Padmanabhan · 2021 [cited by applicant]
US 20210248693A1 · Riland et al. · 2021 [cited by applicant]
US 20220230020A1 · Saeugling · 2022 [cited by applicant]
US 20230127525A1 · Rai · 2023 [cited by examiner]
Y. Baek et al, Character Region Awareness for Text Detection, CVPR paper, 2019, pp. 9365-9374. [cited by applicant]
J. Baek et al, What is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis, Clova AI Research, arXiv:1904.01906v4 [cs.CV] Dec. 18, 2019, pp. 1-19. [cited by applicant]
International Searching Authority, PCT Notification of the International Search Report and Written Opinion, International Application No. PCT/US22/51902, Mar. 7, 2023, pp. 1-15. [cited by applicant]