IP Library › Granted Patent US 12,242,468
Granted Patent B1
US 12,242,468 · App. 18/797,294 · Granted Mar 4, 2025

Generative machine learning with retriever having reconfigurable sequence of rankers

Inventors: Eliot P. Brenner (Westfield, NJ); Koustuv Dasgupta (Scarsdale, NY); Dinesh Gupta (Princeton, NJ); Manjunath G. Hegde (Bangalore, IN); Amy Francesca Pajak (London, GB); Goncalo Nuno Ventura de Melo (London, GB); Abdallah Mohamed Abdo Mohamed Bashir (London, GB)
Assignee: Goldman Sachs & Co. LLC
G06F16/242G06F16/24522G06F16/24578
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,468
App. No.
18/797,294
Granted
Mar 4, 2025
Kind
B1
Abstract

A method includes obtaining an input query at a retriever model, where the retriever model includes a reconfigurable sequence of one or more rankers selected from among a plurality of rankers. Each ranker is configured to identify a specified number of information chunks relevant to the input query. The method also includes providing one or more of the information chunks from the retriever model to a generative model. The method further includes using the generative model to create a response to the input query, where the response is based on the one or more information chunks. The plurality of rankers includes a bi-encoder, a cross-encoder, and a large language model (LLM)-ranker.

Claims (53)

1. A method comprising:

obtaining an input query at a retriever model, the retriever model comprising a reconfigurable sequence of one or more rankers selected from among a plurality of rankers, each ranker configured to identify a specified number of information chunks relevant to the input query;

providing one or more of the information chunks from the retriever model to a generative model;

using the generative model to create a response to the input query, the response based on the one or more information chunks; and

tuning the retriever model by determining the specified number of information chunks to be identified by each ranker in the reconfigurable sequence;

wherein the plurality of rankers comprises a bi-encoder, a cross-encoder, and a large language model (LLM)-ranker; and

wherein the specified number of information chunks to be identified by each ranker in the reconfigurable sequence is determined using a grid search.

2. The method of claim 1 , wherein the retriever model is configured to process different fields of information in the information chunks differently.

3. The method of claim 1 , wherein the retriever model is configured to dynamically select a size of the one or more information chunks provided to the generative model.

4. The method of claim 1 , further comprising:

providing the response to a user device associated with a user.

5. The method of claim 1 , wherein the generative model comprises a large language model.

6. The method of claim 1 , wherein the grid search used to determine the specified number of information chunks to be identified by each ranker in the reconfigurable sequence comprises:

selecting values for the specified number of information chunks to be identified by each ranker in the reconfigurable sequence using a pre-defined grid;

using the retriever model and the generative model to process a validation set that includes annotated document chunks based on the selected values;

evaluating an accuracy of a system that includes the retriever model and the generative model based on the processing of the validation set; and

simultaneously tuning the values for the specified number of information chunks to be identified by each ranker in the reconfigurable sequence based on results of the evaluating.

7. The method of claim 6 , wherein outputs generated by the retriever model and the generative model based on the validation set are compared to annotations of the annotated document chunks, the annotations treated as ground truths.

8. An apparatus comprising:

at least one processing device configured to:

provide an input query to a retriever model, the retriever model comprising a reconfigurable sequence of one or more rankers selected from among a plurality of rankers, each ranker configured to identify a specified number of information chunks relevant to the input query;

provide one or more of the information chunks from the retriever model to a generative model;

use the generative model to create a response to the input query, the response based on the one or more information chunks; and

determine the specified number of information chunks to be identified by each ranker in the reconfigurable sequence in order to tune the retriever model;

wherein the plurality of rankers comprises a bi-encoder, a cross-encoder, and a large language model (LLM)-ranker; and

wherein the at least one processing device is configured to determine the specified number of information chunks to be identified by each ranker in the reconfigurable sequence using a grid search.

9. The apparatus of claim 8 , wherein the retriever model is configured to process different fields of information in the information chunks differently.

10. The apparatus of claim 8 , wherein the retriever model is configured to dynamically select a size of the one or more information chunks provided to the generative model.

11. The apparatus of claim 8 , wherein the at least one processing device is further configured to provide the response to a user device associated with a user.

12. The apparatus of claim 8 , wherein the generative model comprises a large language model.

13. The apparatus of claim 8 , wherein, to determine the specified number of information chunks to be identified by each ranker in the reconfigurable sequence using the grid search, the at least one processing device configured to:

select values for the specified number of information chunks to be identified by each ranker in the reconfigurable sequence using a pre-defined grid;

use the retriever model and the generative model to process a validation set that includes annotated document chunks based on the selected values;

evaluate an accuracy of a system that includes the retriever model and the generative model based on the processing of the validation set; and

simultaneously tune the values for the specified number of information chunks to be identified by each ranker in the reconfigurable sequence based on results of the evaluating.

14. The apparatus of claim 13 , wherein the at least one processing device is configured to compare outputs generated by the retriever model and the generative model based on the validation set to annotations of the annotated document chunks, the annotations treated as ground truths.

15. A non-transitory computer readable medium containing instructions that when executed cause at least one processor to:

obtain an input query at a retriever model, the retriever model comprising a reconfigurable sequence of one or more rankers selected from among a plurality of rankers, each ranker configured to identify a specified number of information chunks relevant to the input query;

provide one or more of the information chunks from the retriever model to a generative model;

use the generative model to create a response to the input query, the response based on the one or more information chunks; and

determine the specified number of information chunks to be identified by each ranker in the reconfigurable sequence to tune the retriever model;

wherein the plurality of rankers comprises a bi-encoder, a cross-encoder, and a large language model (LLM)-ranker; and

wherein the instructions when executed cause the at least one processor to determine the specified number of information chunks to be identified by each ranker in the reconfigurable sequence using a grid search.

16. The non-transitory computer readable medium of claim 15 , wherein the retriever model is configured to process different fields of information in the information chunks differently.

17. The non-transitory computer readable medium of claim 15 , wherein the retriever model is configured to dynamically select a size of the one or more information chunks provided to the generative model.

18. The non-transitory computer readable medium of claim 15 , further containing instructions that when executed cause the at least one processor to provide the response to a user device associated with a user.

19. The non-transitory computer readable medium of claim 15 , wherein the instructions that when executed cause the at least one processor to determine the specified number of information chunks to be identified by each ranker in the reconfigurable sequence using the grid search comprise:

instructions that when executed cause the at least one processor to:

select values for the specified number of information chunks to be identified by each ranker in the reconfigurable sequence using a pre-defined grid;

use the retriever model and the generative model to process a validation set that includes annotated document chunks based on the selected values;

evaluate an accuracy of a system that includes the retriever model and the generative model based on the processing of the validation set; and

simultaneously tune the values for the specified number of information chunks to be identified by each ranker in the reconfigurable sequence based on results of the evaluating.

20. The non-transitory computer readable medium of claim 19 , wherein the instructions when executed cause the at least one processor to compare outputs generated by the retriever model and the generative model based on the validation set to annotations of the annotated document chunks, the annotations treated as ground truths.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 7, 2024
From: BRENNER, ELIOT P.; DASGUPTA, KOUSTUV; GUPTA, DINESH; HEGDE, MANJUNATH G.; PAJAK, AMY FRANCESCA; VENTURA DE MELO, GONCALO NUNO; BASHIR, ABDALLAH MOHAMED ABDO MOHAMED
To: GOLDMAN SACHS & CO. LLC
Reel/Frame 068215/0186 →
Priority Claims (1)
IN 202311078197 · Nov 17, 2023 · national
Continuity (1)
Continuation 18659799 · May 9, 2024
References Cited (26)
US 20240028909A1 · Wang · 2024 [cited by examiner]
US 20240289561A1 · Qadrud-Din · 2024 [cited by examiner]
Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, arXiv:2005.11401v4 [cs.CL], Apr. 2021, 19 pages. [cited by applicant]
Wikipedia, “Large Language Model”, Oct. 2023, 20 pages. [cited by applicant]
Wikipedia, “Self-Supervised Learning”, Oct. 2023, 6 pages. [cited by applicant]
Allenai, “A Repository of Language Instructions for NLP Tasks”, GitHub, Feb. 2023, 11 pages. [cited by applicant]
Bigscience-Workshop, “PromptSource”, GitHub, Jul. 2022, 7 pages. [cited by applicant]
LangChain, “Prompting Strategies”, Mar. 2024, 44 pages. [cited by applicant]
Weng, “Prompt Engineering”, Lil'Log, Mar. 2023, 80 pages. [cited by applicant]
Hwchase17, “LangChainHub”, GitHub, Apr. 2023, 10 pages. [cited by applicant]
Yuan et al., “Self-Rewarding Language Models”, arXiv:2401.10020v2 [cs.CL], Feb. 2024, 23 pages. [cited by applicant]
Khattab et al., “ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT”, arXiv:2004.12832v2 [cs.IR], Jun. 2020, 10 pages. [cited by applicant]
Rafailov et al., “Direct Preference Optimization: Your Language Model is Secretly a Reward Model”, arXiv:2305.18290v2 [cs.LG], Dec. 2023, 27 pages. [cited by applicant]
Raudaschl, “RAG-Fusion: The Next Frontier of Search Technology”, GitHub, Sep. 2023, 7 pages. [cited by applicant]
Pradeep et al., “RankVicuna: Zero-Shot Listwise Document Reranking with Open-Source Large Language Models”, arXiv:2309.15088v1 [cs.IR], Sep. 2023, 10 pages. [cited by applicant]
Cormack et al., “Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods”, SIGIR '09: Proceedings of the 32nd international ACM SIGIR Conference, Jul. 2009, 2 pages. [cited by applicant]
Wikipedia, “Bradley—Terry model”, Jun. 2023, 3 pages. [cited by applicant]
U.S. Securities and Exchange Commission, “Credit Agreement among Dunkin Finance Corp.”, Nov. 2010, 34 pages. [cited by applicant]
Bclavie, “Welcome to RAGatouille”, GitHub, Mar. 2024, 17 pages. [cited by applicant]
Lucidrains, “Self-Rewarding Language Model”, GitHub, Apr. 2024, 25 pages. [cited by applicant]
Ouyang et al., “Training language models to follow instructions with human feedback”, arXiv:2203.02155v1 [cs.CL], Mar. 2022, 68 pages. [cited by applicant]
Myscale, “RQABench: Retrieval QA Benchmark”, GitHub, Sep. 2024, 11 pages. [cited by applicant]
Ma et al., “Query Rewriting for Retrieval-Augmented Large Language Models”, arXiv:2305.14283v3 [cs.CL], Oct. 2023, 13 pages. [cited by applicant]
Yu et al., “Generate rather than Retrieve: Large Language Models are Strong Context Generators”, arXiv:2209.10063v3 [cs.CL], Jan. 2023, 27 pages. [cited by applicant]
Ye et al., “Prompt Engineering a Prompt Engineer”, arXiv:2311.05661v2 [cs.CL], Feb. 2024, 31 pages. [cited by applicant]
Brenner et al., “Retrieval-Augmented Generation (RAG) System Optimization”, U.S. Appl. No. 18/659,799, filed May 9, 2024, 54 pages. [cited by applicant]
Cited By (8)
US 12,517,941 US 12,524,417 US 12,536,373 US 12,547,634 US 12,561,314 US 12,566,755 US 12,639,275 US 12,650,993