IP Library Granted Patent US 12,499,387
Granted Patent B2
US 12,499,387 · App. 19/088,106 · Granted Dec 16, 2025

Question answer (QA) pair generation using agentic workflow system and method that aligns artificial intelligence model with domain-specific principles

Inventors: Stefanos Poulis (Vienna, VA); Andrew J. Bauer (Vienna, VA); Diego A. Mesa (Vienna, VA); Robin J. Clark (Vienna, VA); Patrick C. Condo (Vienna, VA)
Assignee: Seekr Technologies Inc.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,387
App. No.
19/088,106
Granted
Dec 16, 2025
Kind
B2
Abstract

A question and answer (QA) pairs generation system and method uses an agentic workflow system and method to align generative artificial intelligence (a large language model (LLM) or a large multimodal model (LMM)) with the principles of a specific domain so that the generative artificial intelligence is better able to respond to a user query in the specific domain.

Claims (35)

1 . A method, comprising:

retrieving, by a computer system, a plurality of pieces of content that together embody one or more principles for a specific domain;

creating, on the computer system for each piece of retrieved content, a structured representation of one or more pieces of data in each piece of retrieved content;

performing, on the computer system on the structured representation for each piece of retrieved content, a recursive summarization to generate a summary for each piece of retrieved content that is stored back into the structured representation for each piece of retrieved content;

performing, by the computer system based on the structured representation for each piece of retrieved content that includes the summary, a relevance assessment of each node in the structured representation of the one or more pieces of data in each piece of retrieved content to identify a node relevant to a fine tuning direction for an artificial intelligence model (AIM) to generate an initial plurality of question answer pairs (QA pairs) aligned to the specific domain;

performing, by the computer system on each initial QA pair, an answer refinement for the relevant node to generate a refined QA pair and an acceptable plurality of QA pairs aligned to the one or more principles of the specific domain; and

adding, by the computer system, the refined QA pair to the summary of the relevant node.

2 . The method of claim 1 , wherein the structured representation of each piece of retrieved content is a document tree.

3 . The method of claim 2 , wherein each node of the document tree stores one of a piece of text and a structural element of the piece of retrieved content.

4 . The method of claim 1 further comprising training the artificial intelligence model (AIM) using the acceptable plurality of QA pairs to align the AIM with the specific domain.

5 . The method of claim 4 further comprising receiving a query and generating an aligned response by the aligned AIM to the received query.

6 . The method of claim 4 , wherein the AIM is one of a large language model and a large multimodal model.

7 . The method of claim 6 , wherein the specific domain is one of an industry standard, a civility score, an enterprise domain, a set of pieces of content from a computer and a blog post.

8 . The method of claim 1 further comprising generating one or more alignment processes using the acceptable QA pairs.

9 . The method of claim 8 further comprising outputting a response to a query from a trained artificial intelligence model (AIM) and adjusting, using the one or more alignment processes, the output response from the trained AIM to generate an aligned response to the query that is aligned to the specific domain.

10 . The method of claim 9 , wherein the trained AIM is one of a large language model and a large multimodal model.

11 . The method of claim 10 , wherein the specific domain is one of an industry standard, a civility score, an enterprise domain, a set of pieces of content from a computer and a blog post.

12 . A system, comprising:

a computer system having a hardware processor and a plurality of lines of instructions executed by the hardware processor so that the hardware processor is configured to:

retrieve a plurality of pieces of content that together embody one or more principles for a specific domain;

create, for each piece of retrieved content, a structured representation of one or more pieces of data in each piece of retrieved content;

perform, on the structured representation for each piece of retrieved content, a recursive summarization to generate a summary for each piece of retrieved content that is stored back into the structured representation for each piece of retrieved content;

perform, based on the structured representation for each piece of retrieved content that includes the summary, a relevance assessment of each node in the structured representation of the one or more pieces of data in each piece of retrieved content to identify a node relevant to a fine tuning direction for an artificial intelligence model (AIM) to generate an initial first plurality of question answer pairs (QA pairs) aligned to the specific domain;

perform, on each initial QA pair, an answer refinement for the relevant node to generate a refined QA pair and an acceptable plurality of QA pairs aligned to the specific domain; and

add the refined QA pair to the summary of the relevant node.

13 . The system of claim 12 , wherein the structured representation of each piece of retrieved content is a document tree.

14 . The system of claim 13 , wherein each node in the document tree stores one of a piece of text and a structural element of the piece of retrieved content.

15 . The system of claim 12 , wherein the hardware processor is further configured to train the artificial intelligence model (AIM) using the acceptable plurality of QA pairs to align the AIM with the specific domain.

16 . The system of claim 15 , wherein the hardware processor is further configured to receive a query and generate an aligned response by the aligned AIM to the received query.

17 . The system of claim 15 , wherein the AIM is one of a large language model and a large multimodal model.

18 . The system of claim 17 , wherein the specific domain is one of an industry standard, a civility score, an enterprise domain, a set of pieces of content from a computer and a blog post.

19 . The system of claim 12 , wherein the hardware processor is further configured to generate one or more alignment processes using the acceptable QA pairs.

20 . The system of claim 19 , wherein the hardware processor is further configured to output a response to a query from a trained artificial intelligence model (AIM) and adjust, using the one or more alignment processes, the output response from the trained AIM to generate an aligned response to the query that is aligned to the specific domain.

21 . The system of claim 20 , wherein the trained AIM is one of a large language model and a large multimodal model.

22 . The system of claim 21 , wherein the specific domain is one of an industry standard, a civility score, an enterprise domain, a set of pieces of content from a computer and a blog post.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2025
From: POULIS, STEFANOS; BAUER, ANDREW J.; MESA, DIEGO A.; CLARK, ROBIN J.; CONDO, PATRICK C.
To: SEEKR TECHNOLOGIES INC.
Reel/Frame 070601/0940 →
Continuity (5)
Continuation 19008444 · Jan 2, 2025
Continuation In Part 19004001 · Dec 27, 2024
Continuation 18646104 · Apr 25, 2024
Continuation In Part 18599955 · Mar 8, 2024
Related Publication 20250285027A1 · Sep 11, 2025
References Cited (104)
US 5696962A · Kupiec · 1997 [cited by applicant]
US 5909510A · Nakayama · 1999 [cited by applicant]
US 6026388A · Liddy et al. · 2000 [cited by applicant]
US 6119114A · Smadja · 2000 [cited by applicant]
US 6651057B1 · Jin et al. · 2003 [cited by applicant]
US 6847969B1 · Mathal et al. · 2005 [cited by applicant]
US 7062485B1 · Jin et al. · 2006 [cited by applicant]
US 7120925B2 · D'Souza et al. · 2006 [cited by applicant]
US 7197497B2 · Cossock · 2007 [cited by applicant]
US 7313622B2 · Lee et al. · 2007 [cited by applicant]
US 7475404B2 · Hamel · 2009 [cited by applicant]
US 7606810B1 · Jeavons · 2009 [cited by applicant]
US 7925973B2 · Allaire et al. · 2011 [cited by applicant]
US 7933893B2 · Walker et al. · 2011 [cited by applicant]
US 8195666B2 · Jeavons · 2012 [cited by applicant]
US 8219911B2 · Clarke-Martin et al. · 2012 [cited by applicant]
US 8478758B2 · Jeavons · 2013 [cited by applicant]
US 10733452B2 · Attorre · 2020 [cited by applicant]
US 11875240B1 · Bosnjakovic · 2024 [cited by applicant]
US 11893981B1 · Clark et al. · 2024 [cited by applicant]
US 11921731B2 · Baeza-Yates et al. · 2024 [cited by applicant]
US 12124932B1 · Poulis et al. · 2024 [cited by applicant]
US 12174903B1 · Poulis et al. · 2024 [cited by applicant]
US 12182678B1 · Poulis et al. · 2024 [cited by applicant]
US 12210535B1 · Poulis et al. · 2025 [cited by applicant]
US 12254872B2 · Clark et al. · 2025 [cited by applicant]
US 20050091340A1 · Facemire · 2005 [cited by applicant]
US 20050144158A1 · Capper · 2005 [cited by applicant]
US 20060117348A1 · D'Souza et al. · 2006 [cited by applicant]
US 20070038567A1 · Allaire et al. · 2007 [cited by applicant]
US 20070038931A1 · Allaire et al. · 2007 [cited by applicant]
US 20080010142A1 · O'Brien et al. · 2008 [cited by applicant]
US 20080104113A1 · Wong · 2008 [cited by applicant]
US 20080221983A1 · Ausiannik et al. · 2008 [cited by applicant]
US 20090197581A1 · Gupta et al. · 2009 [cited by applicant]
US 20090248668A1 · Zheng · 2009 [cited by applicant]
US 20100100545A1 · Jeavons · 2010 [cited by applicant]
US 20100313116A1 · Hyman · 2010 [cited by applicant]
US 20110166918A1 · Allaire et al. · 2011 [cited by applicant]
US 20110191163A1 · Allaire et al. · 2011 [cited by applicant]
US 20120078895A1 · Chu-Carroll · 2012 [cited by applicant]
US 20120143792A1 · Wang · 2012 [cited by applicant]
US 20130318063A1 · Ayzenshtat · 2013 [cited by applicant]
US 20150095014A1 · Marimuthu · 2015 [cited by applicant]
US 20160021037A1 · Hewitt · 2016 [cited by applicant]
US 20180101534A1 · Alexander, Jr. · 2018 [cited by applicant]
US 20190065744A1 · Gaustad · 2019 [cited by applicant]
US 20190082224A1 · Bradley · 2019 [cited by applicant]
US 20190147062A1 · Kim · 2019 [cited by applicant]
US 20190163327A1 · Otero · 2019 [cited by applicant]
US 20200125639A1 · Doyle · 2020 [cited by applicant]
US 20200126533A1 · Doyle · 2020 [cited by applicant]
US 20210004420A1 · Mittal · 2021 [cited by applicant]
US 20210019339A1 · Ghulati · 2021 [cited by applicant]
US 20210058352A1 · Fogu et al. · 2021 [cited by applicant]
US 20210097239A1 · Arora et al. · 2021 [cited by applicant]
US 20210240700A1 · Ling et al. · 2021 [cited by applicant]
US 20230316000A1 · Mukherjee · 2023 [cited by applicant]
US 20240111498A1 · Vaughn · 2024 [cited by applicant]
US 20240184991A1 · Mahabaleshwarkar · 2024 [cited by applicant]
US 20240202221A1 · Siebel · 2024 [cited by applicant]
US 20240281472A1 · LaRhette · 2024 [cited by applicant]
US 20240419713A1 · Siebel · 2024 [cited by applicant]
US 20250021568A1 · Poulis et al. · 2025 [cited by applicant]
US 20250094821A1 · Hettige · 2025 [cited by applicant]
US 20250131028A1 · Siebel · 2025 [cited by applicant]
US 20250200392A1 · Whalen · 2025 [cited by applicant]
Baulepur, “Aligning Language Models with Factuality and Truthfulness” Thesis submitted in partial fulfillment of Bachelor of Science in Computer Science, University of Illinois at Urbana-Champaign, 2023, 50 pages. [cited by applicant]
Azaria, et al., “The Internal State of an LLM Knows When its Lying”, School of Computer Science, Ariel University, Israel and Machine Learning Dept., Carnegie Mellon University, Pittsburgh, PA, Apr. 2023, 10 pages. [cited by applicant]
Lee, et al., “Linguistic Properties of Truthful Response,” University of Pennsylvania, PA, USA., Jun. 2023, 6 pages. [cited by applicant]
Poulis, “Algorithms for Interactive Machine Learning”, Dissertation submitted in partial fulfillment of degree of Doctor of Philosophy in Computer Science, University of California, San Diego, 2019, 148 pages. [cited by applicant]
Yang, et al., “RefGPT: Reference—Truthful & Customized Dialogues Generation by GPTs and for GPTs”, Shanghai Jiao Tong University, Hong Kong Polytechnical University, Beijing University of Posts and Telecommunications, M… [cited by applicant]
Pan, et al., “On the Risk of Misinformation Pollution with Large Language Models”, National University of Singapore, University of California, Santa Barbara, University of Waterloo, MBZUAI, Zhejiang University, May 2023… [cited by applicant]
McKenna, et al., “Sources of Hallucination by Large Language Models on Inference Tasks”, University of Edinburgh, Google Research, Macquarie University, May 2023, 17 pages. [cited by applicant]
Sun, “Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision”, 37th Conference on Neural Information Processing Systems, 2023. (Year: 2023). [cited by applicant]
Siriwardhana, “Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for Open Domain Question Answering”, Transactions of the Association for Computational Linguistics, vol. 11, pp. 1-17, 2023. … [cited by applicant]
Zhang, “Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching” Jun. 2024, 30 pgs, https://arxiv.org/abs/2406.06326. [cited by applicant]
Rozner, “Knowledge Editing in Language Models via Adapted Direct Preference Optimization” Sep. 2024, 13 pgs, https://arxiv.org/abs/2406.09920. [cited by applicant]
Ye, “Qilin Med: Multi-stage Knowledge Injection Advanced Medical Large Language Model”, 13 pgs, Apr. 2024. [cited by applicant]
Ovadia, “Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs”, Jan. 2024, 14 pgs, https://arxiv.org/abs/2312.05934. [cited by applicant]
Zhu, “FanOutQA: a Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models” Jun. 2024, 20 pgs, https://arxiv.org/abs/2402.14116. [cited by applicant]
Soudani, “Fine Tuning vs. Retrieval Augmented Generation for Less Popular Knowledge” Dec. 2024, 11 pgs, https://arxiv.org/abs/2403.01432. [cited by applicant]
Zhang, “RAFT: Adapting Language Model to Domain Specific RAG” 12 pgs, Jun. 2024, https://arxiv.org/abs/2403.10131. [cited by applicant]
Angelov, “Explainable artificial intelligence: an analytical review”, 13 pages, Jun. 10, 2021, WIREs Data Mining Knowl Discov. 2021;11:e1424, https://doi.org/10.1002/widm.1424. [cited by applicant]
Bai, et al., “Constitutional AI: Harmlessness from AI Feedback”, 34 pages, arXiv:2212.08073v1 [cs CL] Dec. 15, 2022. [cited by applicant]
Confalonieri, et al., “A Historical perspective of explainable Artificial Intelligence”, WIREs Data Mining and Knowledge Discovery published by Wiley Periodicals LLC., Sep. 2020, WIREs Data Mining Knowl Discov. 2021;11:… [cited by applicant]
Das, “Opportunities and Challenges in Explainable Artificial Intelligence (XAI): a Survey”, Department of Electrical and Computer Engineering, University of Texas at San Antonio, San Antonio, TX, 78249, arXiv:2006.11371… [cited by applicant]
Dawson, Algorithmic Adjudication and Constitutional AI—The Promise of a Better AI Decision Making Future?, 29 pages, 27 SMU Science & Technology L. Rev. 11 (2024). [cited by applicant]
Goush, et al., “A Closer Look at the Limitations of Instruction Tuning”, 31 pages, arXiv:2402.05119v5 [cs.CL] Jul. 14, 2024. [cited by applicant]
Dosilovic, et al., “Explainable Artificial Intelligence: a Survey”, 7 pages, University of Zagreb, Conference Paper ⋅ May 2018, DOI: 10.23919/MIPRO.2018.8400040. [cited by applicant]
Gunning, et al., “XAI—Explainable Artificial Intelligence”, 6 pages, City, University of London Institutional Repository, 2019, Science Robotics, 4(37), caay7120. Doi: 10.1126/scirobotics.aay7120. [cited by applicant]
Huang, et al., “Collective Constitutional AI: Aligning a Language Model with Public Input”, 23 pages, arXiv:2406.07814v1 [cs.Al] Jun. 12, 2024. [cited by applicant]
Mecklenburg, et al., “Injecting New Knowledge Into Large Language Models via Supervised Fine-Tuning”, 16 pages, arXiv:2404.00213v2 [cs.CL] Apr. 2, 2024. [cited by applicant]
Mitra, et al., “AgentInstruct: Toward Generative Teaching with Agentic Flows”, 32 pages, arXiv:2407.03502v1 [cs.AI] Jul. 3, 2024. [cited by applicant]
Shen, et al., “Large Language Model Alignment: a Survey”, 76 pages, College of Intelligence and Computing, Tianjin University, Tianjin China, arXiv:2309.15025v1 [cs.CL] Sep. 26, 2023. [cited by applicant]
Wang, et al., “Chain-of-Thought Reasoning without Prompting”, 23 pages, 2024 Google DeepMind, arXiv:2402.10200v2 [cs.CL] May 23, 2024. [cited by applicant]
Xu, et al., “A Survey on Knowledge Distillation of Large Language Models”, 43 pages, arXiv:2402.13116v4 [cs.CL] Oct. 21, 2024. [cited by applicant]
Zeng, et al., Scaling of Search and Learning: a Roadmap to Reproduce ol from Reinforcement Learning Perspective, 51 pages, arXiv:2412.14135v1 [cs.AI] Dec. 18, 2024. [cited by applicant]
Gekham, et al., “Does Fine-Tuning LLMS on New Knowledge Encourage Hallucinations?”, 20 pages, Technion—Israel Institute of Technology, Google Research, arXiv:2405.05904v3 [cs.CL] Oct. 1, 2024. [cited by applicant]
Durante, “Agent AI: Surveying the Horizons of Multimodel Interaction”, Jan. 25, 2024 (Year: 2024). [cited by applicant]
Zhang, et al., Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering, 11 pages. arXiv:2311.06503v1 [cs.CL] Nov. 11, 2023. [cited by applicant]
Cui, et al., “Ada-Instruct: Adapting Instruction Generators for Complex Reasoning” Shanghi University of Finance and Economics, arXiv:2310.04484v2 [cs.CL] Oct. 10, 2023. [cited by applicant]
Abiri, “Public Constitutional AI”, 2024. [cited by applicant]
Claude, “Collective Constitutional AI:Aligning a Language Model with Public Input”,2023. Webpage: file///Collective%20Constitutional%20 AI_%20. [cited by applicant]