IP Library Granted Patent US 12,436,995
Granted Patent B1
US 12,436,995 · App. 18/758,095 · Granted Oct 7, 2025

Interactive user assistance system with rich multimodal content experience

Inventors: Min Gong (Shanghai, CN); Zijia Wang (London, GB); Deepaganesh Paulraj (Bangalore, IN)
Assignee: Dell Products L.P.
G06F16/73G06F16/9558
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,436,995
App. No.
18/758,095
Filed
Jun 28, 2024
Granted
Oct 7, 2025
Kind
B1
Art Unit
2159
USPC
707/706
Abstract

Methods and systems for managing an interactive user assistance system that provides a rich multimodal content experience for a user are disclosed. In particular, hyperlinks may be included in responses generated for the user's questions to the system. These hyperlinks may take the user to one or more videos or video clips containing content that would assist the user in better resolving the user's questions. The videos or video clips may be stored in a hierarchically managed video clip pool that provides efficient retrieval and association with related content for enhancing the user's access to such multimodal content.

Claims (55)

1. A method for managing an interactive user assistance system that provides a rich multimodal content experience for a user, the method comprising:

obtaining an input from the user;

generating a response to the input using one or more machine learning models;

identifying one or more referenceable terms within the response using the response and a video hierarchy and semantic representation database, the video hierarchy and semantic representation database that comprises a semantic vector space in which a plurality of semantic representations of video clips of videos are stored and a hierarchically managed video clip pool, the hierarchically managed video clip pool comprises, for a first video among the videos:

first video clips clipped from the first video;

a first video hierarchy, of one or more video hierarchies that indicate a hierarchical association between each of the videos and video clips of the videos, created for the first video and the first video clips;

a semantic representation, of the plurality of semantic representations, generated for each of the first video and the first video clips,

wherein the semantic representation of each of the first video and the first video clips are stored into the semantic vector space, the semantic vector space and information associated with the first video hierarchy are stored into the video hierarchy and semantic representation database as part of the hierarchically managed video clip pool, and the semantic vector space comprises areas associated with each of the one or more referenceable terms, and the semantic representation of each of the first video and the first video clips are stored in respective ones of the areas of the semantic vector space based on a level of similarity between the semantic representation and the one or more referenceable terms;

replacing, using the level of similarity, each of the one or more referenceable terms within the response with a hyperlink to a video, among the videos stored in the hierarchically managed video clip pool, to obtain an enhanced response; and

providing the enhanced response to the user.

2. The method of claim 1 , wherein the one or more machine learning models comprise a large language model (LLM) utilizing a retrieval-augmented generation (RAG) framework.

3. The method of claim 1 , wherein the semantic representation of each of the first video and the first video clips are generated using multimodal content embedding techniques.

4. The method of claim 1 , wherein

the video associated with the hyperlink in the enhanced response is one of the first video or the first video clips, and

upon detecting that the user has accessed the first video or the first video clips using the hyperlink, providing the user with information regarding the first video hierarchy to grant the user access to any ones of the first video and the first video clips within the first video hierarchy.

5. The method of claim 1 , wherein the input comprises a question and the video comprises content for assisting the user resolve the question.

6. The method of claim 5 , wherein the question is associated with an error of a data processing system, and the content comprises step-by-step tutorials for resolving the error.

7. A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations for managing an interactive user assistance system that provides a rich multimodal content experience for a user, the operations comprising:

obtaining an input from the user;

generating a response to the input using one or more machine learning models;

identifying one or more referenceable terms within the response using the response and a video hierarchy and semantic representation database, the video hierarchy and semantic representation database that comprises a semantic vector space in which a plurality of semantic representations of video clips of videos are stored and a hierarchically managed video clip pool, the hierarchically managed video clip pool comprises, for a first video among the videos:

first video clips clipped from the first video;

a first video hierarchy, of one or more video hierarchies that indicate a hierarchical association between each of the videos and video clips of the videos, created for the first video and the first video clips;

a semantic representation, of the plurality of semantic representations, generated for each of the first video and the first video clips,

wherein the semantic representation of each of the first video and the first video clips are stored into the semantic vector space, the semantic vector space and information associated with the first video hierarchy are stored into the video hierarchy and semantic representation database as part of the hierarchically managed video clip pool, and the semantic vector space comprises areas associated with each of the one or more referenceable terms, and the semantic representation of each of the first video and the first video clips are stored in respective ones of the areas of the semantic vector space based on a level of similarity between the semantic representation and the one or more referenceable terms;

replacing, using the level of similarity, each of the one or more referenceable terms within the response with a hyperlink to a video, among the videos stored in the hierarchically managed video clip pool, to obtain an enhanced response; and

providing the enhanced response to the user.

8. The non-transitory machine-readable medium of claim 7 , wherein the one or more machine learning models comprise a large language model (LLM).

9. The non-transitory machine-readable medium of claim 8 , wherein the LLM utilizes a retrieval-augmented generation (RAG) framework.

10. The non-transitory machine-readable medium of claim 7 , wherein the semantic representation of each of the first video and the first video clips are generated using multimodal content embedding techniques.

11. The non-transitory machine-readable medium of claim 7 , wherein

the video associated with the hyperlink in the enhanced response is one of the first video or the first video clips, and

upon detecting that the user has accessed the first video or the first video clips using the hyperlink, providing the user with information regarding the first video hierarchy to grant the user access to any ones of the first video and the first video clips within the first video hierarchy.

12. The non-transitory machine-readable medium of claim 7 , wherein the input comprises a question and the video comprises content for assisting the user resolve the question.

13. The non-transitory machine-readable medium of claim 12 , wherein the question is associated with an error of a data processing system, and the content comprises step-by-step tutorials for resolving the error.

14. A user assistance manager, comprising:

a processor; and

a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations for managing an interactive user assistance system that provides a rich multimodal content experience for a user, the operations comprising:

obtaining an input from the user;

generating a response to the input using one or more machine learning models;

identifying one or more referenceable terms within the response using the response and a video hierarchy and semantic representation database, the video hierarchy and semantic representation database that comprises a semantic vector space in which a plurality of semantic representations of video clips of videos are stored and a hierarchically managed video clip pool, the hierarchically managed video clip pool comprises, for a first video among the videos:

first video clips clipped from the first video;

a first video hierarchy, of one or more video hierarchies that indicate a hierarchical association between each of the videos and video clips of the videos, created for the first video and the first video clips;

a semantic representation, of the plurality of semantic representations, generated for each of the first video and the first video clips,

wherein the semantic representation of each of the first video and the first video clips are stored into the semantic vector space, the semantic vector space and information associated with the first video hierarchy are stored into the video hierarchy and semantic representation database as part of the hierarchically managed video clip pool, and the semantic vector space comprises areas associated with each of the one or more referenceable terms, and the semantic representation of each of the first video and the first video clips are stored in respective ones of the areas of the semantic vector space based on a level of similarity between the semantic representation and the one or more referenceable terms;

replacing, using the level of similarity, each of the one or more referenceable terms within the response with a hyperlink to a video, among the videos stored in the hierarchically managed video clip pool, to obtain an enhanced response; and

providing the enhanced response to the user.

15. The user assistance manager of claim 14 , wherein the semantic representation of each of the first video and the first video clips are generated using multimodal content embedding techniques.

16. The user assistance manager of claim 14 , wherein the one or more machine learning models comprise a large language model (LLM).

17. The user assistance manager of claim 16 , wherein the LLM utilizes a retrieval-augmented generation (RAG) framework.

18. The user assistance manager of claim 14 , wherein

the video associated with the hyperlink in the enhanced response is one of the first video or the first video clips, and

upon detecting that the user has accessed the first video or the first video clips using the hyperlink, providing the user with information regarding the first video hierarchy to grant the user access to any ones of the first video and the first video clips within the first video hierarchy.

19. The user assistance manager of claim 14 , wherein the input comprises a question and the video comprises content for assisting the user resolve the question.

20. The user assistance manager of claim 19 , wherein the question is associated with an error of a data processing system, and the content comprises step-by-step tutorials for resolving the error.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2024
From: GONG, MIN; WANG, ZIJIA; PAULRAJ, DEEPAGANESH
To: DELL PRODUCTS L.P.
Reel/Frame 067884/0741 →
References Cited (53)
US 8538897B2 · Han et al. · 2013 [cited by applicant]
US 9852646B2 · Kozloski · 2017 [cited by examiner]
US 10217058B2 · Gamon · 2019 [cited by examiner]
US 10572329B2 · Harutyunyan et al. · 2020 [cited by applicant]
US 10616314B1 · Plenderleith et al. · 2020 [cited by applicant]
US 10740793B1 · Sussman et al. · 2020 [cited by applicant]
US 10776196B2 · Ohana et al. · 2020 [cited by applicant]
US 10853867B1 · Bulusu et al. · 2020 [cited by applicant]
US 11513930B2 · Chan et al. · 2022 [cited by applicant]
US 11720940B2 · Lakshminarayan et al. · 2023 [cited by applicant]
US 11734102B1 · Wang et al. · 2023 [cited by applicant]
US 11748185B2 · Xu et al. · 2023 [cited by applicant]
US 11909836B2 · Wulf et al. · 2024 [cited by applicant]
US 12061970B1 · Lo et al. · 2024 [cited by applicant]
US 20040125124A1 · Kim · 2004 [cited by examiner]
US 20090113248A1 · Bock et al. · 2009 [cited by applicant]
US 20090216910A1 · Duchesneau · 2009 [cited by applicant]
US 20100257058A1 · Karidi et al. · 2010 [cited by applicant]
US 20100318856A1 · Yoshida · 2010 [cited by applicant]
US 20130041748A1 · Hsiao et al. · 2013 [cited by applicant]
US 20130198240A1 · Ameri-Yahia et al. · 2013 [cited by applicant]
US 20140310222A1 · Davlos et al. · 2014 [cited by applicant]
US 20150161241A1 · Haggar · 2015 [cited by examiner]
US 20150227838A1 · Wang et al. · 2015 [cited by applicant]
US 20150288557A1 · Gates et al. · 2015 [cited by applicant]
US 20180205645A1 · Bays · 2018 [cited by applicant]
US 20190095313A1 · Xu et al. · 2019 [cited by applicant]
US 20190129785A1 · Liu et al. · 2019 [cited by applicant]
US 20200026590A1 · Lopez et al. · 2020 [cited by applicant]
US 20210027205A1 · Sevakula et al. · 2021 [cited by applicant]
US 20210241141A1 · Dugger et al. · 2021 [cited by applicant]
US 20210287109A1 · Cmielowski et al. · 2021 [cited by applicant]
US 20220100187A1 · Isik et al. · 2022 [cited by applicant]
US 20220283890A1 · Chopra et al. · 2022 [cited by applicant]
US 20220358005A1 · Saha et al. · 2022 [cited by applicant]
US 20220417078A1 · Matsuo et al. · 2022 [cited by applicant]
US 20230016199A1 · Jividen et al. · 2023 [cited by applicant]
US 20240028955A1 · Harutyunyan et al. · 2024 [cited by applicant]
US 20240168835A1 · Wang et al. · 2024 [cited by applicant]
US 20250086211A1 · Bolcer et al. · 2025 [cited by applicant]
CN 108280168A · 2018 [cited by applicant]
CN 111476371A · 2020 [cited by applicant]
CN 112541806A · 2021 [cited by applicant]
EP 4235505A1 · 2023 [cited by applicant]
Kevin Dela Rosa, Video Enriched Retrieval Augmented Generation Using Aligned Video Captions, May 27, 2024 [retrieved online Mar. 6, 2025]. Retrieved from the Internet: https://doi.org/10.48550/arXiv.2405.17706 (Year: 20… [cited by examiner]
Zhao, Wayne Xin, et al., “A Survey of Large Language Models,” arXiv preprint arXiv:2303.18223 (2023) (97 Pages). [cited by applicant]
Kaddour, Jean, et al., “Challenges and Applications of Large Language Models,” arXiv preprint arXiv:2307.10169 (2023) (72 Pages). [cited by applicant]
Naveed, Humza, et al., “A Comprehensive Overview of Large Language Models,” arXiv preprint arXiv:2307.06435 (2023) (35 Pages). [cited by applicant]
Boffa, Matteo, et al., “LogPrécis: Unleashing Language Models for Automated Shell Log Analysi,” arXiv preprint arXiv:2307.08309 (2023) (17 Pages). [cited by applicant]
Chen, Yinfang, et al., “Empowering Practical Root Cause Analysis by Large Language Models for Cloud Incidents,” arXiv preprint arXiv:2305.15778 (2023) (15 Pages). [cited by applicant]
Lee, Yukyung, et al., “LAnoBERT : System Log Anomaly Detection based on BERT Masked Language Model,” Applied Soft Computing 146 (2023): 110689 (18 Pages). [cited by applicant]
Pfeiffer, Jonas, et al. “Adapterfusion: Non-destructive task composition for transfer learning.” arXiv preprint arXiv:2005.00247 (2020) (17 Pages). [cited by applicant]
Houlsby, Neil, et al. “Parameter-efficient transfer learning for NLP.” International conference on machine learning. PMLR, 2019 (13 Pages). [cited by applicant]