IP Library Granted Patent US 12,443,607
Granted Patent B1
US 12,443,607 · App. 18/804,648 · Granted Oct 14, 2025

LLM-based recommender system for data catalog

Inventors: Jing Guo (Xi'an, CN); Ming Yan (Xi'an, CN); Lianjie Qin (Xi'an, CN); Jingtao Li (Xi'an, CN)
Assignee: SAP SE
G06F16/24575G06F16/248
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,443,607
App. No.
18/804,648
Granted
Oct 14, 2025
Kind
B1
Abstract

LLMs use probabilistic methods to create coherent responses, sometimes going beyond the training data. This can result in “LLM hallucination.” Asset metadata is used in constructing prompts for the LLM, enhancing the LLM's understanding of available assets. Including a limited set of candidate assets in an LLM prompt template as part of a chain-of-thought prompt can solve the hallucination issue. Multiple rounds of interaction with the LLM may be used, allowing for more dynamic and responsive user engagement. The LLM-based recommender system for data catalogs may comprise asset metadata, user data, and a recommender model. The output of the LLM-based recommender system is a recommended asset list. The prompts may explicitly instruct the LLM to limit elements of the list to the candidate set of assets. The LLM's advanced context understanding and reasoning abilities enable it to deliver accurate and interpretable personalized recommendations.

Claims (48)

1. A system for recommending data assets, the system comprising:

a memory that stores instructions; and

one or more processors coupled to the memory and configured to execute the instructions to perform operations comprising:

determining, based on a first asset of a plurality of data assets, a candidate asset set;

generating, based on the candidate asset set and metadata for the candidate asset set, a refined candidate set;

generating, based on the refined candidate set, a prompt for a large language model (LLM); and

receiving, from the LLM and in response to the prompt, a structured list of recommended data assets, the recommended data assets being a subset of the candidate asset set.

2. The system of claim 1 , wherein the determining of the candidate asset set comprises using a collaborative filtering algorithm.

3. The system of claim 1 , wherein the generating of the LLM prompt is further based on a user role.

4. The system of claim 3 , wherein the generating of the LLM prompt is further based on a user interest.

5. The system of claim 1 , wherein the generating of the LLM prompt is further based on a user query.

6. The system of claim 1 , wherein the LLM prompt further comprises chain-of-thought instructions.

7. The system of claim 1 , wherein the operations further comprise:

receiving, via a user interface, a user query;

providing the user query to the LLM; and

receiving, from the LLM and in response to the user query, a revised structured list of recommended data assets.

8. The system of claim 1 , wherein the operations further comprise:

determining, based on a user role and a user area of interest, a second candidate asset set;

generating, based on the second candidate asset set and metadata for the second candidate asset set, a second refined candidate set;

generating, based on the second refined candidate set, a second prompt for the LLM; and

receiving, from the LLM, a second structured list of second recommended data assets, the second recommended data assets being a subset of the second candidate asset set.

9. A non-transitory computer-readable medium that stores instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

determining, based on a first asset of a plurality of data assets, a candidate asset set;

generating, based on the candidate asset set and metadata for the candidate asset set, a refined candidate set;

generating, based on the refined candidate set, a prompt for a large language model (LLM); and

receiving, from the LLM and in response to the prompt, a structured list of recommended data assets, the recommended data assets being a subset of the candidate asset set.

10. The non-transitory computer-readable medium of claim 9 , wherein the determining of the candidate asset set comprises using a collaborative filtering algorithm.

11. The non-transitory computer-readable medium of claim 9 , wherein the generating of the LLM prompt is further based on a user role.

12. The non-transitory computer-readable medium of claim 11 , wherein the generating of the LLM prompt is further based on a user interest.

13. The non-transitory computer-readable medium of claim 9 , wherein the generating of the LLM prompt is further based on a user query.

14. The non-transitory computer-readable medium of claim 9 , wherein the LLM prompt further comprises chain-of-thought instructions.

15. The non-transitory computer-readable medium of claim 9 , wherein the operations further comprise:

receiving, via a user interface, a user query;

providing the user query to the LLM; and

receiving, from the LLM and in response to the user query, a revised structured list of recommended data assets.

16. The non-transitory computer-readable medium of claim 9 , wherein the operations further comprise:

determining, based on a user role and a user area of interest, a second candidate asset set;

generating, based on the second candidate asset set and metadata for the second candidate asset set, a second refined candidate set;

generating, based on the second refined candidate set, a second prompt for the LLM; and

receiving, from the LLM, a second structured list of second recommended data assets, the second recommended data assets being a subset of the second candidate asset set.

17. A method comprising:

determining, by one or more processors and based on a first asset of a plurality of data assets, a candidate asset set;

generating, based on the candidate asset set and metadata for the candidate asset set, a refined candidate set;

generating, based on the refined candidate set, a prompt for a large language model (LLM); and

receiving, from the LLM and in response to the prompt, a structured list of recommended data assets, the recommended data assets being a subset of the candidate asset set.

18. The method of claim 17 , wherein the determining of the candidate asset set comprises using a collaborative filtering algorithm.

19. The method of claim 17 , wherein the generating of the LLM prompt is further based on a user role.

20. The method of claim 19 , wherein the generating of the LLM prompt is further based on a user interest.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2024
From: GUO, JING; YAN, MING; QIN, LIANJIE; LI, JINGTAO
To: SAP SE
Reel/Frame 068285/0264 →
References Cited (13)
US 7305436B2 · Willis · 2007 [cited by examiner]
US 7370276B2 · Willis · 2008 [cited by examiner]
US 7885902B1 · Shoemaker · 2011 [cited by examiner]
US 12135740B1 · Yu · 2024 [cited by examiner]
US 12306828B1 · Chakraborty · 2025 [cited by examiner]
US 20220374329A1 · Savir · 2022 [cited by examiner]
US 20240265193A1 · Schafer · 2024 [cited by examiner]
US 20240296287A1 · Jia · 2024 [cited by examiner]
US 20240330411A1 · Wu · 2024 [cited by examiner]
US 20240372876A1 · Shachar · 2024 [cited by examiner]
US 20240403373A1 · Chao · 2024 [cited by examiner]
US 20250045336A1 · Gonsalves · 2025 [cited by examiner]
US 20250173514A1 · Greene · 2025 [cited by examiner]
Cited By (1)
US 12,681,925