IP Library › Granted Patent US 12,602,410
Granted Patent B2
US 12,602,410 · App. 18/631,590 · Granted Apr 14, 2026

Multimodal context selection for large language model based resolutions addressing technical issues

Inventors: Ravi Shukla (Bengaluru, IN); Gaurav Bhattacharjee (Bangalore, IN); Ramakanth Kanagovi (Hyderabad, IN)
Assignee: Dell Products L.P.
G06F16/3329G06F16/383
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,410
App. No.
18/631,590
Granted
Apr 14, 2026
Kind
B2
Abstract

A method for technical issue resolution. The method includes: receiving, from a user, a text query concerning a technical issue; obtaining query-related context relevant to the text query; and processing, through a large language model (LLM), the text query and the query-related context to produce a multimodal query response used by the user to address the technical issue. More specifically, embodiments described herein utilize text topic and zero shot classification models to translate multimodal technical documentation (e.g., including text and images) into topic relevant metadata; and process queries, pertaining to technical issues, using a multimodal LLM provided with query-related text and image context derived from said topic relevant metadata.

Claims (65)

1 . A method for technical issue resolution, the method comprising:

extracting a text portion of a multimodal document;

partitioning the text portion into text portion chunks;

processing, using a text topic model, the text portion chunks to discover a plurality of text topics associated with the multimodal document and to obtain a normalized topic weight between each text portion chunk and each text topic;

selecting, for each text topic, a text portion chunks subset comprising a plurality of text portion chunks each mapped to the text topic and to the normalized topic weight at least meeting a normalized topic weight threshold; and

combining, for each text topic, the text portion chunks subset to obtain a topic-related text representative of the text topic;

receiving, from a user and following combining the text portion chunks subset, a text query concerning a technical issue;

obtaining query-related context relevant to the text query, wherein obtaining the query-related context, comprises:

identifying a query-related text topic of the topic-related text for the text query;

selecting at least one query-related text topic chunk each representing a text portion chunk mapped to the query-related text topic, wherein the text portion chunk is further mapped to the normalized topic weight at least equal to the normalized topic weight threshold;

combining the at least one query-related text topic chunk to obtain query-related text context;

selecting at least one query-related image subset each representing an image portion subset mapped to the query-related text topic, wherein the image portion subset is further mapped to an image-topic similarity score at least equal to an image-topic similarity score threshold; and

combining the at least one query-related image subset and the query-related text context to obtain the query-related context; and

processing, through a large language model (LLM), the text query and the query-related context to produce a multimodal query response used by the user to address the technical issue.

2 . The method of claim 1 , wherein the text portion chunk and the image portion subset belong to a same multimodal document.

3 . The method of claim 1 , wherein the text portion chunk and the image portion subset belong to different multimodal documents.

4 . The method of claim 1 , wherein identifying the query-related text topic, comprises:

obtaining topic-related texts representative of a plurality of text topics;

processing, using a zero shot text model, the text query and the topic-related texts to produce a query-topic vector comprising a query-topic similarity score for each topic-related text; and

identifying a text topic mapped to a topic-related text mapped to a highest query-topic similarity score as the query-related text topic.

5 . The method of claim 1 , wherein the multimodal query response comprises text content and image content.

6 . The method of claim 1 , the method further comprising:

extracting, from the multimodal document, an image portion comprising at least one image; and

processing, using a zero shot image model, the at least one image and the plurality of text topics to obtain an image-topic similarity score between each image and each text topic.

7 . A non-transitory computer readable medium (CRM) comprising computer readable program code, which when executed by a computer processor, enables the computer processor to perform a method for technical issue resolution, the method comprising:

extracting a text portion of a multimodal document;

partitioning the text portion into text portion chunks;

processing, using a text topic model, the text portion chunks to discover a plurality of text topics associated with the multimodal document and to obtain a normalized topic weight between each text portion chunk and each text topic;

selecting, for each text topic, a text portion chunks subset comprising a plurality of text portion chunks each mapped to the text topic and to the normalized topic weight at least meeting a normalized topic weight threshold; and

combining, for each text topic, the text portion chunks subset to obtain a topic-related text representative of the text topic;

receiving, from a user and following combining the text portion chunks subset, a text query concerning a technical issue;

obtaining query-related context relevant to the text query, wherein obtaining the query-related context, comprises:

identifying a query-related text topic of the topic-related text for the text query;

selecting at least one query-related text topic chunk each representing a text portion chunk mapped to the query-related text topic, wherein the text portion chunk is further mapped to the normalized topic weight at least equal to the normalized topic weight threshold;

combining the at least one query-related text topic chunk to obtain query-related text context;

selecting at least one query-related image subset each representing an image portion subset mapped to the query-related text topic, wherein the image portion subset is further mapped to an image-topic similarity score at least equal to an image-topic similarity score threshold; and

combining the at least one query-related image subset and the query-related text context to obtain the query-related context; and

processing, through a large language model (LLM), the text query and the query-related context to produce a multimodal query response used by the user to address the technical issue.

8 . The non-transitory CRM of claim 7 , wherein the text portion chunk and the image portion subset belong to a same multimodal document.

9 . The non-transitory CRM of claim 7 , wherein the text portion chunk and the image portion subset belong to different multimodal documents.

10 . The non-transitory CRM of claim 7 , wherein identifying the query-related text topic, comprises:

obtaining topic-related texts representative of a plurality of text topics;

processing, using a zero shot text model, the text query and the topic-related texts to produce a query-topic vector comprising a query-topic similarity score for each topic-related text; and

identifying a text topic mapped to a topic-related text mapped to a highest query-topic similarity score as the query-related text topic.

11 . The non-transitory CRM of claim 7 , wherein the multimodal query response comprises text content and image content.

12 . The non-transitory CRM of claim 7 , the method further comprising:

extracting, from the multimodal document, an image portion comprising at least one image; and

processing, using a zero shot image model, the at least one image and the plurality of text topics to obtain an image-topic similarity score between each image and each text topic.

13 . A system, comprising:

a large language model (LLM); and

a multimodal context selector operatively connected to the LLM, and comprising a computer processor configured to perform a method for technical issue resolution, the method comprising:

extracting a text portion of a multimodal document;

partitioning the text portion into text portion chunks;

processing, using a text topic model, the text portion chunks to discover a plurality of text topics associated with the multimodal document and to obtain a normalized topic weight between each text portion chunk and each text topic;

selecting, for each text topic, a text portion chunks subset comprising a plurality of text portion chunks each mapped to the text topic and to the normalized topic weight at least meeting a normalized topic weight threshold; and

combining, for each text topic, the text portion chunks subset to obtain a topic-related text representative of the text topic;

receiving, from a user and following combining the text portion chunks subset, a text query concerning a technical issue;

obtaining query-related context relevant to the text query, wherein obtaining the query-related context, comprises:

identifying a query-related text topic of the topic-related text for the text query;

selecting at least one query-related text topic chunk each representing a text portion chunk mapped to the query-related text topic, wherein the text portion chunk is further mapped to the normalized topic weight at least equal to the normalized topic weight threshold;

combining the at least one query-related text topic chunk to obtain query-related text context;

selecting at least one query-related image subset each representing an image portion subset mapped to the query-related text topic, wherein the image portion subset is further mapped to an image-topic similarity score at least equal to an image-topic similarity score threshold; and

combining the at least one query-related image subset and the query-related text context to obtain the query-related context; and

processing, through the LLM, the text query and the query-related context to produce a multimodal query response used by the user to address the technical issue.

14 . The system of claim 13 , wherein the multimodal query response comprises text content and image content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2024
From: SHUKLA, RAVI; BHATTACHARJEE, GAURAV; KANAGOVI, RAMAKANTH
To: DELL PRODUCTS L.P.
Reel/Frame 067105/0526 →
Continuity (1)
Related Publication 20250321986A1 · Oct 16, 2025
References Cited (8)
US 11972223B1 · DeFoor · 2024 [cited by examiner]
US 12050599B1 · Garcia-Sanchez · 2024 [cited by examiner]
US 20120163707A1 · Baker · 2012 [cited by examiner]
US 20220405315A1 · Mahindru · 2022 [cited by examiner]
US 20230281225A1 · Ross · 2023 [cited by examiner]
US 20240248963A1 · Parham · 2024 [cited by examiner]
US 20240281472A1 · LaRhette · 2024 [cited by examiner]
US 20240370736A1 · Chang · 2024 [cited by examiner]