IP Library › Granted Patent US 12,399,907
Granted Patent B2
US 12,399,907 · App. 18/898,506 · Granted Aug 26, 2025

Copilot implementation: recursive iteration of retrieval augmented generation (RAG)

Inventors: Elaine Kelsey (Corvallis, OR); Elliot Nicholas Robson (Seoul, KR); Sazzad Mahmud Nasir (Muncie, IN); Jeffrey Thomas Yarbro (Memphis, TN); Robert Oscar Robson (Corvallis, OR); Lauren Elizabeth Egerton (New York, NY); Spencer Thomas Ward (Kent, WA); Brendan Michael Kelly (Somerville, MA)
Assignee: THIA ST Co.
G06F16/3329G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,399,907
App. No.
18/898,506
Filed
Sep 26, 2024
Granted
Aug 26, 2025
Kind
B2
Art Unit
2166
USPC
707/758
Abstract

Apparatus and methods are disclosed for implementing a copilot as a network of microservices including specialized large language models (LLMs) or other trained machine learning (ML) tools. This architecture supports flexible, customizable, or dynamically determinable dataflow. Compared to much larger competing LLMs, comparable or superior performance is achieved, while significantly reducing computation time and hardware requirements, even to a single compute node with a single GPU. Examples incorporate a retrieval microservice, as least one data producer, and a core microservice. Based on client input, the retrieval microservice can perform multiple iterations of retrieval augmented generation (RAG). At each iteration, output (based on any preceding iterations' results or the client input) is transmitted to a data producer, and results received therefrom. Eventually, based on these results, an output is transmitted toward the core microservice for generation of a response to the client input. Variations and additional techniques are disclosed.

Claims (140)

1. A computer-implemented method for augmenting input for a core microservice of a copilot, comprising:

receiving second input, which reflects first input from a client;

recursively performing a plurality of iterations of retrieval augmented generation (RAG), each of the iterations comprising:

transmitting third output to a data producer; and

responsive to the transmitting, receiving fourth input from the data producer;

wherein, on an initial iteration of the iterations, the third output is based on the second input; and

wherein, on subsequent iterations of the iterations, the third output is based on the fourth input of an immediately preceding iteration of the iterations;

wherein the iterations are terminated upon meeting a stopping criterion comprising: two successive iterations of the iterations returning respective fourth inputs whose similarity is greater than or equal to a predetermined threshold; and

transmitting, toward the core microservice, fourth output comprising at least part of the fourth input received on a final iteration of the iterations;

wherein the fourth output enables the core microservice to generate a response to the first input from the client.

2. The computer-implemented method of claim 1 , further comprising:

receiving the first input from the client;

extracting one or more first tokens from the first input;

determining one or more second tokens associated with but distinct from the first tokens; and

generating a first output combining the first and second tokens; and

transmitting the first output to a retrieval microservice, where the first output is received as the second input.

3. The computer-implemented method of claim 1 , wherein the second input comprises a plurality of tokens, and the method further comprises:

executing the recursively performing acts individually for each of the plurality of tokens to obtain respective final fourth inputs on respective final iterations;

wherein the fourth output comprises, for each of the plurality of tokens, the fourth input of a final iteration of the iterations.

4. The computer-implemented method of claim 1 , wherein the fourth output comprises the fourth input of each of the plurality of iterations.

5. The computer-implemented method of claim 1 , wherein one or more of the third output, the fourth input, or the fourth output comprises an array of elements, each of the elements comprising a document, a vector representation of a document, or a token derived from the second input.

6. One or more computer-readable media storing instructions which, when executed by one or hardware processors, cause the one or more hardware processors to perform operations for augmenting input for a core microservice of a copilot, the operations comprising:

receiving second input, which reflects first input from a client application;

recursively performing a plurality of iterations of retrieval augmented generation (RAG), each of the iterations comprising:

transmitting third output to a data producer; and

responsive to the transmitting, receiving fourth input from the data producer;

wherein, on an initial iteration of the iterations, the third output is based on the second input; and

wherein, on subsequent iterations of the iterations, the third output is based on the fourth input of an immediately preceding iteration of the iterations;

scoring, for relevance to the first or second input, constituents of the fourth inputs received on respective iterations of the plurality of iterations;

transmitting, toward the core microservice, fourth output comprising the fourth input received on a final iteration of the iterations;

retaining, in the fourth output, high-scoring constituents among the constituents of the fourth inputs; and

discarding low-scoring constituents among the constituents of the fourth inputs;

wherein the fourth output enables the core microservice to generate a response to the first input from the client application.

7. A system comprising:

one or more hardware processors, with memory coupled thereto; and

one or more computer readable media storing instructions comprising a plurality of modules which, when executed by the one or more hardware processors, implement respective microservices, the microservices forming a weakly connected network of microservices configured as a copilot for one or more first client applications;

wherein each of the microservices is configured to:

receive input from (i) a respective first group comprising one or more others of the microservices or (ii) one or more second client applications; and

transmit output to (i) a second group comprising one or more of the microservices or (ii) one or more third client applications;

wherein a plurality of the microservices incorporate respective trained machine learning tools;

wherein the network of microservices comprises at least a retrieval microservice, one or more data producers, a qualification microservice, and a core microservice;

wherein the retrieval microservice is configured to:

receive second input, which reflects first input from a given one of the one or more first client applications;

recursively perform a plurality of iterations of retrieval augmented generation (RAG), each of the iterations comprising:

transmitting third output to a respective data producer of the one or more data producers; and

responsive to the transmitting, receiving fourth input from the respective data producer;

wherein, on an initial iteration of the iterations, the third output is based on the second input; and

wherein, on subsequent iterations of the iterations, the third output is based on the fourth input of an immediately preceding iteration of the iterations; and

transmit, toward the core microservice, fourth output comprising the fourth input received on a final iteration of the iterations;

wherein the fourth output enables the core microservice to generate a response to the first input from the given first client application; and

wherein the qualification microservice configured to:

receive a ninth input based on the fourth output;

compare the ninth input with a graphical model of a knowledge corpus incorporated in the copilot;

determine whether the copilot is competent to act on the ninth input;

upon determining that the copilot is competent:

determine and transmit ninth output based on the ninth input, toward the core microservice; and

upon determining that the copilot is not competent, transmit a notification indicating lack of competence.

8. The system of claim 7 , wherein the respective data producer, at a given iteration of the iterations, comprises:

a document repository storing a plurality of documents; and

an index storing vector representations of the stored documents.

9. The system of claim 7 wherein, for all of the iterations, the respective data producer is a common data producer.

10. The system of claim 7 wherein, for given first and second iterations of the iterations, the respective data producer for the first iteration is distinct from the respective data producer for the second iteration.

11. A computer-implemented method for augmenting input for a core microservice of a copilot, comprising:

receiving second input, which reflects first input from a client;

recursively performing a plurality of iterations of retrieval augmented generation (RAG), each of the iterations comprising:

transmitting third output to a data producer;

responsive to the transmitting, receiving fourth input from the data producer;

scoring, for relevance to the first or second input, constituents of the fourth input;

retaining high-scoring constituents among the constituents of the fourth input; and

discarding low-scoring constituents among the constituents of the fourth input;

wherein, on an initial iteration of the iterations, the third output is based on the second input; and

wherein, on subsequent iterations of the iterations, the third output is based on the fourth input of an immediately preceding iteration of the iterations;

transmitting, toward the core microservice, fourth output comprising the retained constituents of the fourth input received on a final iteration of the iterations;

wherein the fourth output enables the core microservice to generate a response to the first input from the client.

12. A computer-implemented method for augmenting input for a core microservice of a copilot, comprising:

receiving second input, which reflects first input from a client;

recursively performing a plurality of iterations of retrieval augmented generation (RAG), each of the iterations comprising:

transmitting third output to a data producer; and

responsive to the transmitting, receiving fourth input from the data producer;

wherein, on an initial iteration of the iterations, the third output is based on the second input; and

wherein, on subsequent iterations of the iterations, the third output is based on the fourth input of an immediately preceding iteration of the iterations;

wherein the iterations are terminated upon meeting a stopping criterion comprising: an increase in volume over successive iterations being less than or equal to a predetermined threshold, the volume on each of the successive iterations being defined by the second input and the fourth input on the respective iteration, in a graphical representation of a knowledge domain; and

transmitting, toward the core microservice, fourth output comprising at least part of the fourth input received on a final iteration of the iterations;

wherein the fourth output enables the core microservice to generate a response to the first input from the client.

13. One or more computer-readable media storing instructions which, when executed by one or hardware processors, cause the one or more hardware processors to perform first operations for augmenting input for a core microservice of a copilot and second operations of a data producer, the first operations comprising:

receiving second input, which reflects first input from a client application;

recursively performing a plurality of iterations of retrieval augmented generation (RAG), each of the iterations comprising:

transmitting third output to the data producer; and

responsive to the transmitting, receiving fourth input from the data producer;

wherein, on an initial iteration of the iterations, the third output is based on the second input; and

wherein, on subsequent iterations of the iterations, the third output is based on the fourth input of an immediately preceding iteration of the iterations; and

transmitting, toward the core microservice, fourth output comprising the fourth input received on a final iteration of the iterations;

wherein the fourth output enables the core microservice to generate a response to the first input from the client application; and

the second operations comprising:

receiving, on a given iteration of the plurality of iterations, a seventh input based on the third output;

retrieving, from one or more databases, database objects relevant to the seventh input; and

determining and transmitting a seventh output based on the retrieved database objects;

wherein, on the given iteration, the fourth input corresponding to the third output is based on the seventh output.

14. The one or more computer-readable media of claim 13 , wherein the iterations are terminated upon meeting a stopping criterion, wherein:

the stopping criterion comprises completion of a predetermined number of the iterations.

15. The one or more computer-readable media of claim 13 , wherein the iterations are terminated upon meeting a stopping criterion, wherein:

the stopping criterion comprises an amount of new data in the fourth input being less than or equal to a first predetermined threshold.

16. One or more computer-readable media storing instructions which, when executed by one or hardware processors, cause the one or more hardware processors to perform first operations for augmenting input for a core microservice of a copilot and second operations of a qualification microservice, the first operations comprising:

receiving second input, which reflects first input from a client application;

recursively performing a plurality of iterations of retrieval augmented generation (RAG), each of the iterations comprising:

transmitting third output to a data producer; and

responsive to the transmitting, receiving fourth input from the data producer;

wherein, on an initial iteration of the iterations, the third output is based on the second input; and

wherein, on subsequent iterations of the iterations, the third output is based on the fourth input of an immediately preceding iteration of the iterations; and

transmitting, toward the core microservice, fourth output comprising the fourth input received on a final iteration of the iterations;

the second operations comprising:

receiving a ninth input based on the fourth output;

comparing the ninth input with a graphical model of a knowledge corpus incorporated in the copilot;

determining whether the copilot is competent to act on the ninth input;

upon determining that the copilot is competent:

determining and transmitting ninth output based on the ninth input, toward the core microservice; and

upon determining that the copilot is not competent, transmitting a notification indicating lack of competence;

wherein the ninth output enables the core microservice to generate a response to the first input from the client application.

17. A system comprising:

one or more hardware processors, with memory coupled thereto; and

one or more computer readable media storing instructions comprising a plurality of modules which, when executed by the one or more hardware processors, implement respective microservices, the microservices forming a weakly connected network of microservices configured as a copilot for one or more first client applications;

wherein each of the microservices is configured to:

receive input from (i) a respective first group comprising one or more others of the microservices or (ii) one or more second client applications; and

transmit output to (i) a second group comprising one or more of the microservices or (ii) one or more third client applications;

wherein a plurality of the microservices incorporate respective trained machine learning tools;

wherein the network of microservices comprises at least a retrieval microservice, one or more data producers, and a core microservice;

wherein the retrieval microservice is configured to:

receive second input, which reflects first input from a given one of the one or more first client applications;

recursively perform a plurality of iterations of retrieval augmented generation (RAG), each of the iterations comprising:

transmitting third output to a respective data producer of the one or more data producers, wherein the respective data producer, at a given iteration of the iterations, is a database microservice; and

responsive to the transmitting, receiving fourth input from the respective data producer;

wherein, on an initial iteration of the iterations, the third output is based on the second input; and

wherein, on subsequent iterations of the iterations, the third output is based on the fourth input of an immediately preceding iteration of the iterations; and

transmit, toward the core microservice, fourth output comprising the fourth input received on a final iteration of the iterations;

wherein the fourth output enables the core microservice to generate a response to the first input from the given first client application; and

wherein the database microservice is configured to:

receive, on the given iteration, a seventh input based on the third output;

retrieve, from one or more databases, database objects relevant to the seventh input; and

determine and transmit a seventh output based on the retrieved database objects;

wherein the fourth input corresponding to the third output, on the given iteration, is based on the seventh output.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2024
From: EDUWORKS CORPORATION
To: THIA ST CO.
Reel/Frame 069643/0557 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2024
From: KELSEY, ELAINE; ROBSON, ELLIOT NICHOLAS; NASIR, SAZZAD MAHMUD; YARBRO, JEFFREY THOMAS; ROBSON, ROBERT OSCAR; EGERTON, LAUREN ELIZABETH; WARD, SPENCER THOMAS; KELLY, BRENDAN MICHAEL
To: EDUWORKS CORPORATION
Reel/Frame 069093/0060 →
Continuity (3)
Provisional Application 63561654 · Mar 5, 2024
Provisional Application 63620329 · Jan 12, 2024
Related Publication 20250231973A1 · Jul 17, 2025
References Cited (129)
US 8442940B1 · Faletti et al. · 2013 [cited by applicant]
US 10713664B1 · Alagappan et al. · 2020 [cited by applicant]
US 10762114B1 · Annunziata et al. · 2020 [cited by applicant]
US 10803127B2 · Alexander et al. · 2020 [cited by applicant]
US 10992780B1 · Rudrappa Goniwada · 2021 [cited by applicant]
US 11182748B1 · Neckermann et al. · 2021 [cited by applicant]
US 11379715B2 · Tang et al. · 2022 [cited by applicant]
US 11443102B1 · Wilson et al. · 2022 [cited by applicant]
US 11922121B2 · Wang · 2024 [cited by applicant]
US 12079570B1 · Mondlock et al. · 2024 [cited by applicant]
US 12093658B1 · Silver et al. · 2024 [cited by applicant]
US 12136043B1 · Davis et al. · 2024 [cited by applicant]
US 12210973B2 · Johnson · 2025 [cited by applicant]
US 20150199646A1 · Taylor et al. · 2015 [cited by applicant]
US 20160246824A1 · Furuhashi et al. · 2016 [cited by applicant]
US 20170262811A1 · Ovadya · 2017 [cited by applicant]
US 20170262949A1 · Jay · 2017 [cited by applicant]
US 20180089593A1 · Patel et al. · 2018 [cited by applicant]
US 20190325353A1 · Aftab et al. · 2019 [cited by applicant]
US 20200050946A1 · Lecue et al. · 2020 [cited by applicant]
US 20200057946A1 · Singaraju et al. · 2020 [cited by applicant]
US 20200069208A1 · Keane · 2020 [cited by applicant]
US 20200186433A1 · Cui · 2020 [cited by applicant]
US 20200241944A1 · Derdak et al. · 2020 [cited by applicant]
US 20200242484A1 · Lecue et al. · 2020 [cited by applicant]
US 20200257733A1 · Mei et al. · 2020 [cited by applicant]
US 20200357001A1 · Lopez Garcia et al. · 2020 [cited by applicant]
US 20210035047A1 · Mossoba et al. · 2021 [cited by applicant]
US 20210089779A1 · Chan et al. · 2021 [cited by applicant]
US 20210109995A1 · Mihindukulasooriya et al. · 2021 [cited by applicant]
US 20210174347A1 · Rose · 2021 [cited by applicant]
US 20210201128A1 · Xu et al. · 2021 [cited by applicant]
US 20210233030A1 · Preuss et al. · 2021 [cited by applicant]
US 20210286831A1 · Girardi · 2021 [cited by examiner]
US 20210295238A1 · Poon et al. · 2021 [cited by applicant]
US 20210312299A1 · Segal et al. · 2021 [cited by applicant]
US 20210312399A1 · Asokan et al. · 2021 [cited by applicant]
US 20210357705A1 · Sung et al. · 2021 [cited by applicant]
US 20210365782A1 · Huang et al. · 2021 [cited by applicant]
US 20220067665A1 · Westerheide et al. · 2022 [cited by applicant]
US 20220164683A1 · Hao et al. · 2022 [cited by applicant]
US 20220232085A1 · Panikkar et al. · 2022 [cited by applicant]
US 20220237892A1 · Anderton-Yang · 2022 [cited by applicant]
US 20220284312A1 · Brecque · 2022 [cited by applicant]
US 20220336060A1 · Koop et al. · 2022 [cited by applicant]
US 20230133373A1 · McGonnell · 2023 [cited by applicant]
US 20230214192A1 · Makhija et al. · 2023 [cited by applicant]
US 20230214238A1 · Yitzhaki et al. · 2023 [cited by applicant]
US 20230237348A1 · Polleri et al. · 2023 [cited by applicant]
US 20230289691A1 · Sabourin · 2023 [cited by applicant]
US 20240012842A1 · Kislal · 2024 [cited by examiner]
US 20240037949A1 · Johnston et al. · 2024 [cited by applicant]
US 20240087743A1 · Grimm et al. · 2024 [cited by applicant]
US 20240095679A1 · Nowak et al. · 2024 [cited by applicant]
US 20240111498A1 · Vaughn · 2024 [cited by applicant]
US 20240119383A1 · Bowers et al. · 2024 [cited by applicant]
US 20240176629A1 · Reddy · 2024 [cited by applicant]
US 20240202177A1 · Bandlamudi et al. · 2024 [cited by applicant]
US 20240303569A1 · Yu et al. · 2024 [cited by applicant]
US 20240356881A1 · Wheeler · 2024 [cited by applicant]
US 20240362503A1 · Prasad et al. · 2024 [cited by applicant]
US 20240370476A1 · Madisetti · 2024 [cited by applicant]
US 20240386015A1 · Crabtree · 2024 [cited by applicant]
US 20240403005A1 · Friddle · 2024 [cited by applicant]
US 20240412720A1 · Vasylyev · 2024 [cited by applicant]
US 20250005288A1 · Amatriain-Rubio et al. · 2025 [cited by applicant]
US 20250069308A1 · Cameron et al. · 2025 [cited by applicant]
Abideen, “Training at Scale: Chinchilla Scaling Laws for Compute-Optimal Training of LLMS,” available from https://medium.com/@zaiinn440/training-at-scale-chinchilla-scaling-laws-for-compute-optimal-training-of-llms-eca… [cited by applicant]
Agüera Y Arcas, “Do Large Language Models Understand Us?” [cited by applicant]
Alayrac et al., “Flamingo: a Visual Language Model for Few-Shot Learning,” ArXiv 2204.14198v2, pp. 1-54 (Nov. 15, 2022). [cited by applicant]
Ayub, “GPT-4o: Successor of GPT-4,” available from https://hamidayub.medium.com/gpt-40-successor-of-gpt-4-8207acf9104e, pp. 1-7 (May 14, 2024). [cited by applicant]
Binz et al., “Using cognitive psychology to understand GPT-3,” ArXiv 2206.14576v1, pp. 1-19 (Jun. 21, 2022). [cited by applicant]
Bzdok et al., “Statistics versus machine learning,” [cited by applicant]
Dunn et al., “SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine,” ArXiv 1704.05179v3, 5 pages (Jun. 11, 2017). [cited by applicant]
Goertzel, “Generative AI vs. AGI: The Cognitive Strengths and Weaknesses of Modern LLMs,” ArXiv 2309.10371v1, pp. 1-92 (Sep. 19, 2023). [cited by applicant]
Gu et al., “Mamba: Linear-Time Sequence Modeling with Selective State Spaces,” ArXiv 2312.00752v1, pp. 1-37 (Dec. 2023). [cited by applicant]
Haase et al., “Artificial muses: Generative Artificial Intelligence Chatbots Have Risen to Human-Level Creativity,” ArXiv 2303.12003v1, pp. 1-17, (Mar. 2023). [cited by applicant]
Henderson et al., “A Repository of Conversational Datasets,” ArXiv 1904.06472v2, 10 pages (May 29, 2019). [cited by applicant]
Hoffmann et al., “Training Compute-Optimal Large Language Models,” ArXiv 2203.15556v1, pp. 1-36 (Mar. 29, 2022). [cited by applicant]
Hugging Face “Sentence Transformers,” downloaded from https://huggingface.co/sentence-transformers on Dec. 28, 2023, pp. 1-8. [cited by applicant]
Jiang et al., “Mistral 7B,” ArXiv 2310.06825v1, pp. 1-9 (Oct. 10, 2023). [cited by applicant]
Kaplan et al., “Scaling Laws for Neural Language Models,” ArXiv 2001.08361v1, pp. 1-30 (Jan. 23, 2020). [cited by applicant]
Khashabi et al., “GooAQ: Open Question Answering with Diverse Answer Types,” ArXiv 2104.08727v2, 13 pages (Sep. 10, 2021). [cited by applicant]
Khattab et al., “ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT,” ArXiv 2004.12832v2, 10 pages (Jun. 4, 2020). [cited by applicant]
Koetsier, “GPT-4 Beats 90% of Lawyers Trying To Pass The Bar,” available from https://www.forbes.com/sites/johnkoetsier/2023/03/14/gpt-4-beats-90-of-lawyers-trying-to-pass-the-bar/, pp. 1-5 (Mar. 2023). [cited by applicant]
Kosinski, “Theory of Mind Might Have Spontaneously Emerged in Large Language Models,” ArXiv 2302.02083v5, 30 pages (Nov. 2023). [cited by applicant]
Koupaee et al., “WikiHow: A Large Scale Text Summarization Dataset,” ArXiv 1810.09305v1, 5 pages (Oct. 18, 2018). [cited by applicant]
Laurencon et al., “Introducing IDEFICS: An Open Reproduction of State-of-the-Art Visual Language Model,” available from https://huggingface.co/blog/idefics, pp. 1-8 (Aug. 2023). [cited by applicant]
Lewis et al., “PAQ: 65 Million Probably-Asked Questions and What You Can Do With Them,” ArXiv 2102.07033v1, 16 pages, (Feb. 13, 2021). [cited by applicant]
Liu et al., “How good are Large Language Models at Out-of-Distribution Detection?” ArXiv 2308.10261v2, 12 pages (Aug. 23, 2023). [cited by applicant]
Mahowald et al., “Dissociating Language and Thought in Large Language Models,” ArXiv 2301.06627v2, pp. 1-41, (Nov. 4, 2023). [cited by applicant]
Martínez, “Re-Evaluating GPT-4's Bar Exam Performance,” downloaded from https://ssrn.com/abstract=4441311 on Dec. 28, 2023, pp. 1-15. [cited by applicant]
Momennejad et al., “Evaluating Cognitive Maps and planning in Large Language Models with CogEval,” [cited by applicant]
Nature “Understanding ChatGPT is a bold new challenge for science,” [cited by applicant]
Naveed et al., “A Comprehensive Overview of Large Language Models,” ArXiv 2307.06435v7, pp. 1-46 (Dec. 27, 2023). [cited by applicant]
Neuronup, “Cognitive Functions,” downloaded from https://neuronup.us/areas-of-intervention/cognitive-functions/ on Dec. 28, 2023, pp. 1-19. [cited by applicant]
OpenAI, “GPT-4 Technical Report,” ArXiv 2303.08774v4, pp. 1-100 (Dec. 19, 2023). [cited by applicant]
Ouyang et al., “Training language models to follow instructions with human feedback,” ArXiv 2203.02155v1, pp. 1-68 (Mar. 4, 2022). [cited by applicant]
Piantadosi et al., “Meaning without reference in large language models,” ArXiv 2208.02957v1, 8 pages (Aug. 5, 2022). [cited by applicant]
Prystawski et al., “Why think step-by-step? Reasoning emerges from the locality of experience,” ArXiv 2304.03843v1, pp. 1-11 (Apr. 7, 2023). [cited by applicant]
Ruis et al., “The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolutions by LLMS,” ArXiv 2210.14986v2, pp. 1-79 (Dec. 3, 2023). [cited by applicant]
Serapio-García et al., “Personality Traits in Large Language Models,” ArXiv 2307.00184v3, pp. 1-53 (Sep. 21, 2023). [cited by applicant]
Stiennon et al., “Learning to summarize from human feedback,” [cited by applicant]
Sumers et al., “Cognitive Architectures for Language Agents,” ArXiv 2309.02427v2, pp. 1-30 (Sep. 27, 2023). [cited by applicant]
Sun et al., “Retentive Network: A Successor to Transformer for Large Language Models,” ArXiv 2307.08621v4, pp. 1-14 (Aug. 9, 2023). [cited by applicant]
Tay et al., “UL2: Unifying Language Learning Paradigms,” ArXiv 2205.05131v3, pp. 1-39 (Feb. 28, 2023). [cited by applicant]
Touvron et al., “Llama 2: Open Foundation and Fine-Tuned Chat Models,” ArXiv 2307.09288v2, pp. 1-77 (Jul. 19, 2023). [cited by applicant]
Ullman, “Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks,” ArXiv 2302.08399v5, pp. 1-11 (Mar. 14, 2023). [cited by applicant]
Vaswani et al., “Attention Is All You Need,” ArXiv 1706.03762v7, pp. 1-15 (Aug. 2, 2023). [cited by applicant]
Wang et al., “What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization,” ArXiv 2204.05832v1, pp. 1-26 (Apr. 12, 2022). [cited by applicant]
Wang et al., “Augmenting Language Models with Long-Term Memory,” ArXiv 2306.07174v1, pp. 1-14 (Jun. 12, 2023). [cited by applicant]
Wikipedia, “GPT-4o,” downloaded Sep. 4, 2024 from https://en.wikipedia.org/wiki/GPT-4o, pp. 1-5. [cited by applicant]
Yao et al., “React: Synergizing Reasoning and Acting in Language Models,” ArXiv 2210.03629v3, pp. 1-33 (Mar. 10, 2023). [cited by applicant]
Ziegler et al., “Fine-Tuning Language Models from Human Preferences,” ArXiv 1909.08593v2, 26 pages (Jan. 8, 2020). [cited by applicant]
Bhardwaj et al., “Pre-training LLMs using human-like development data corpus,” ArXiv 2311.04666v4, 7 pages (Jan. 10, 2024). [cited by applicant]
Xu et al., “S [cited by applicant]
Agarwal et al., “Knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training,” ArXiv:2010.12688 v2 (Mar. 2021). [cited by applicant]
Dorsch et al., “GraphGuard: Enhancing Data Quality in Knowledge Graph Pipelines,” SEMIIM, 14 pages, (Nov. 2023). [cited by applicant]
Liu et al., “Multi-stage pre-training over simplified multimodal pre-training models,” ArXiv:2017.14596 (Jul. 22, 2021). [cited by applicant]
Morisot, “Add a SideNet to your MainNet,” ArXiv:2007.13512v1, 15 pages (Jul. 2020). [cited by applicant]
Bodor et al., “From Development to Deployment: An Approach to MLOps Monitoring for Machine Learning Model,” 14th International Conference on Intelligent Systems: Theories and Applications, 7 pages (Nov. 2023). [cited by applicant]
Feng et al., “Knowledge Card: Filling LLMs' Knowledge Gaps With Plug-in Specialized Language Models,” ArXiv:2305.09955v2, pp. 1-24 (Oct. 2023). [cited by applicant]
Jeong, “A Study on the Implementation of Generative AI Services Using an Enterprise Data-Based LLM Application Architecture,” ArXiv 2309.01105v1, pp. 1-26 (Sep. 2023). [cited by applicant]
Kaddour et al., “Challenges and Applications of Large Language Models,” ArXiv 2307.10169v1, pp. 1-72 (Jul. 2023). [cited by applicant]
Liang et al. “Modular Retrieval for Generalization and Interpretation,” ArXiv:2303.13419v1, pp. 1-15 (Mar. 2023). [cited by applicant]
PCT/US2024/061299 Invitation to Pay Additional Fees, including Partial Search Report and Provisional Opinion, 15 pages (Apr. 9, 2025). [cited by applicant]
PCT/US2024/061934 Invitation to Pay Additional Fees, including Partial Search Report and Provisional Opinion, 20 pages (Apr. 4, 2025). [cited by applicant]
Roca et al., “Microservice chatbot architecture for chronic patient support,” Journal of Biomedical Informatics, pp. 1-19 (Oct. 2019). [cited by applicant]
Shao et al., “Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy,” ArXiv:2305.15294v1, pp. 1-12 (May 2023). [cited by applicant]