IP Library › Granted Patent US 12,625,845
Granted Patent B2
US 12,625,845 · App. 18/667,690 · Granted May 12, 2026

Responding to a user query using machine learning

Inventors: Jon Saad-Falcon (Savannah, GA); Joseph D. Barrow (Alexandria, VA); Varun Manjunatha (San Diego, CA); Anusha Prakash (San Jose, CA); Ryan A. Rossi (San Jose, CA); Franck Dernoncourt (Spokane, WA); Alexa F Siu (San Jose, CA); Ani Nenkova Nenkova (Philadelphia, PA); Seunghyun Yoon (San Jose, CA)
Assignee: ADOBE INC.
G06F16/148
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,625,845
App. No.
18/667,690
Granted
May 12, 2026
Kind
B2
Abstract

A method, apparatus, non-transitory computer readable medium, and system for data processing include obtaining a query relating to a document and identifying metadata for the document based on the query, where the metadata describes a structure including a plurality of portions of the document. Some embodiments including generating, using a machine learning model, a retrieval command based on the query and the metadata, selectively retrieving at least one of the plurality of portions of the document based on the retrieval command, and generating, using the machine learning model, a response to the query based on the at least one of the plurality of portions of the document.

Claims (39)

1 . A method for data processing, comprising:

obtaining, from a computing device, a query relating to a document; and

in response to the obtaining of the query from the computing device:

identifying, by a metadata component executed on a data processing system, metadata for the document based on the query, wherein the metadata describes a structure of the document including a plurality of portions of the document;

generating, by a query component executed on the data processing system, a retrieval command prompt, wherein the retrieval command prompt includes the query, the metadata, a plurality of executable functions, and an instruction to generate, based on the query and the metadata, a retrieval command including at least one executable function of the plurality of executable functions and an argument for the at least one executable function;

generating, by a language generation machine learning model of the data processing system executing an attention mechanism, the retrieval command based on the instruction by computing a first set of attention weights corresponding to the instruction, wherein the retrieval command is based on the first set of attention weights and includes an executable function of the plurality of executable functions and an argument generated by the language generation machine learning model for the executable function;

selectively retrieving at least one portion of the plurality of portions of the document based on a context window size of the language generation machine learning model of the data processing system executing the attention mechanism by executing the executable function according to the argument for the executable function;

generating, by the language generation machine learning model of the data processing system executing the attention mechanism, a response to the query based on the retrieved portion of the document by computing a second set of attention weights corresponding to the retrieved portion of the document, wherein the second set of attention weights is different from the first set of attention weights, wherein the retrieval command is based on the second set of attention weights, and wherein the response comprises natural language text; and

displaying the response to a user via the computing device.

2 . The method of claim 1 , wherein: the query specifies the portion of the document.

3 . The method of claim 1 , wherein: the metadata comprises a hierarchical tree of structural elements included in the document.

4 . The method of claim 3 , wherein obtaining the metadata comprises: generating the hierarchical tree of structural elements based on text of the document.

5 . The method of claim 1 , wherein: the language generation machine learning model is trained to generate text in response to a natural language query.

6 . A non-transitory computer readable medium storing instructions that, when executed by a processor, cause the processor to:

obtain, from a computing device, a query relating to a document; and

in response to the obtaining of the query from the computing device:

identify, by a metadata component executed on a data processing system, metadata for the document based on the query, wherein the metadata describes a structure of the document including a plurality of portions of the document;

generate, by a query component executed on the data processing system, a retrieval command prompt, wherein the retrieval command prompt includes the query, the metadata, a plurality of executable functions, and an instruction to generate, based on the query and the metadata, a retrieval command including at least one executable function of the plurality of executable functions and an argument for the at least one executable function;

generate, by a language generation machine learning model of the data processing system executing an attention mechanism, the retrieval command based on the instruction by computing a first set of attention weights corresponding to the instruction, wherein the retrieval command is based on the first set of attention weights and includes an executable function of the plurality of executable functions and an argument generated by the language generation machine learning model for the executable function;

selectively retrieve at least one portion of the plurality of portions of the document based on a context window size of the language generation machine learning model of the data processing system executing the attention mechanism by executing the executable function according to the argument for the executable function;

generate, by the language generation machine learning model of the data processing system executing the attention mechanism, a response to the query based on the retrieved portion of the document by computing a second set of attention weights corresponding to the retrieved portion of the document and different from the first set of attention weights, wherein the retrieval command is based on the second set of attention weights and wherein the response comprises natural language text; and

display the response to a user via the computing device.

7 . The non-transitory computer readable medium of claim 6 , wherein: the query specifies the portion of the document.

8 . The non-transitory computer readable medium of claim 6 , wherein: the metadata comprises a hierarchical tree of structural elements included in the document.

9 . The non-transitory computer readable medium of claim 8 , wherein the instructions further cause the processor to: generate the hierarchical tree of structural elements based on text of the document.

10 . The non-transitory computer readable medium of claim 6 , wherein: the language generation machine learning model is trained to generate text in response to a natural language query.

11 . A data processing system, comprising:

a memory; and

a processing device coupled to the memory, the processing device configured to perform operations comprising:

obtaining, from a computing device, a query relating to a document; and

in response to the obtaining of the query from the computing device:

identifying, by a metadata component executed on a data processing system, metadata for the document based on the query, wherein the metadata describes a structure of the document including a plurality of portions of the document;

generating, by a query component executed on the data processing system, a retrieval command prompt, wherein the retrieval command prompt includes the query, the metadata, a plurality of executable functions, and an instruction to generate, based on the query and the metadata, a retrieval command including at least one executable function of the plurality of executable functions and an argument for the at least one executable function;

generating, by a language generation machine learning model of the data processing system executing an attention mechanism, the retrieval command based on the instruction by computing a first set of attention weights corresponding to the instruction, wherein the retrieval command is based on the first set of attention weights and includes an executable function of the plurality of executable functions and an argument generated by the language generation machine learning model for the executable function;

selectively retrieving at least one portion of the plurality of portions of the document based on a context window size of the language generation machine learning model of the data processing system executing the attention mechanism by executing the executable function according to the argument for the executable function;

generating, by the language generation machine learning model of the data processing system executing the attention mechanism, a response to the query based on the retrieved portion of the document by computing a second set of attention weights corresponding to the retrieved portion of the document, wherein the second set of attention weights is different from the first set of attention weights, wherein the retrieval command is based on the second set of attention weights, and wherein the response comprises natural language text; and

displaying the response to a user via the computing device.

12 . The system of claim 11 , further comprising: a metadata component configured to identify the metadata for the document based on the query.

13 . The system of claim 12 , wherein the metadata component is further configured to: generate a hierarchical tree of structural elements based on text of the document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2024
From: SAAD-FALCON, JON; BARROW, JOSEPH D.; MANJUNATHA, VARUN; PRAKASH, ANUSHA; ROSSI, RYAN A.; DERNONCOURT, FRANCK; SIU, ALEXA F; NENKOVA, ANI NENKOVA; YOON, SEUNGHYUN
To: ADOBE INC.
Reel/Frame 067450/0570 →
Continuity (1)
Related Publication 20250355833A1 · Nov 20, 2025
References Cited (33)
US 5751286A · Barber · 1998 [cited by examiner]
US 12299081B1 · Asi · 2025 [cited by examiner]
US 20040215664A1 · Hennings · 2004 [cited by examiner]
US 20090287698A1 · Marmaros · 2009 [cited by examiner]
US 20160210294A1 · Komarov · 2016 [cited by examiner]
US 20220342920A1 · Miller · 2022 [cited by examiner]
US 20220382975A1 · Gu · 2022 [cited by examiner]
US 20230306201A1 · Bayomi · 2023 [cited by examiner]
Asai, et al., “Task-aware Retrieval With Instructions”, arXiv preprint arXiv:2211.09260v2 [cs.CL] Dec. 19, 2022, 25 pages. [cited by applicant]
Dasigi, et al., “A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers”, arXiv preprint arXiv:2105.03011v1 [cs.CL] May 7, 2021, 12 pages. [cited by applicant]
Feng, et al., “Knowledge Refinement via Interaction Between Search Engines and Large Language Models”, arXiv preprint arXiv:2305.07402v2 [cs.CL] May 21, 2023, 14 pages. [cited by applicant]
Flesch, “A new readability yardstick”, Journal of Applied Psychology, 32(3):221-33, Jun. 1948, available at https://pubmed.ncbi.nlm.nih.gov/18867058/. [cited by applicant]
Gao, et al., “Precise Zero-Shot Dense Retrieval without Relevance Labels”, arXiv preprint arXiv:2212.10496v1 [cs.IR] Dec. 20, 2022, 11 pages. [cited by applicant]
Gulcehre, et al., “Reinforced Self Training (ReST) for Language Modeling”, arXiv preprint arXiv:2308.08998v2 [cs.CL] Aug. 21, 2023, 23 pages. [cited by applicant]
Kwiatkowski, et al., “Natural Questions: A Benchmark for Question Answering Research”, Transactions of the Association for Computational Linguistics, 7:453-466, 2019, 14 pages. [cited by applicant]
Landeghem, et al., “Document Understanding Dataset and Evaluation (DUDE)”, arXiv preprint arXiv:2305.08455v3 [cs.CV] Sep. 11, 2023, 22 pages. [cited by applicant]
Li, et al., “API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs”, arXiv preprint arXiv:2304.08244v2 [cs.CL] Oct. 25, 2023, 15 pages. [cited by applicant]
Liang, et al., “TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs,” arXiv preprint arXiv:2303.16434v1 [cs.AI] Mar. 29, 2023, 27 pages. [cited by applicant]
Lin, et al., “How to Train Your DRAGON: Diverse Augmentation Towards Generalizable Dense Retrieval”, arXiv preprint arXiv:2302.07452v1 [cs.IR] Feb. 15, 2023, 15 pages. [cited by applicant]
Mathew, et al., “DocVQA: A Dataset for VQA on Document Images”, arXiv preprint arXiv:2007.00398v3 [cs.CV] Jan. 5, 2021, 23 pages. [cited by applicant]
Mialon, et al., “Augmented Language Models: a Survey”, arXiv preprint arXiv:2302.07842v1 [cs.CL] Feb. 15, 2023, 33 pages. [cited by applicant]
Parisi, et al., “TALM: Tool Augmented Language Models”, arXiv preprint arXiv:2205.12255v1 [cs.CL] May 24, 2022, 6 pages. [cited by applicant]
Patil, et al., “Gorilla: Large Language Model Connected with Massive APIs”, arXiv preprint arXiv:2305.15334v1 [cs.CL] May 24, 2023, 18 pages. [cited by applicant]
Pereira, et al., “Visconde: Multi-document QA with GPT-3 and Neural Reranking”, arXiv preprint arXiv:2212.09656v1 [cs.CL] Dec. 19, 2022, 11 pages. [cited by applicant]
Press, et al., “Measuring and Narrowing the Compositionality Gap in Language Models”, arXiv preprint arXiv:2210.03350v3 [cs.CL] Oct. 17, 2023, 25 pages. [cited by applicant]
Rajpurkar, et al., “SQuAD: 100,000+ Questions for Machine Comprehension of Text”, arXiv preprint arXiv:1606.05250v3 [cs.CL] Oct. 11, 2016, 10 pages. [cited by applicant]
Schick, et al., “Toolformer: Language Models Can Teach Themselves to Use Tools”, arXiv preprint arXiv:2302.04761v1 [cs.CL] Feb. 9, 2023, 17 pages. [cited by applicant]
Wang, et al., “Glue: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding”, arXiv preprint arXiv:1804.07461v3 [cs.CL] Feb. 22, 2019, 20 pages. [cited by applicant]
Yao, et al., “ReAct: Synergizing Reasoning and Acting in Language Models”, arXiv preprint arXiv:2210.03629v3 [cs.CL] Mar. 10, 2023, 33 pages. [cited by applicant]
Yu, “Augmentation-Adapted Retriever Improves Generalization of Language Models as Generic Plug-In”, arXiv preprint arXiv:2305.17331v1 [cs.CL] May 27, 2023, 14 pages. [cited by applicant]
Zhao, et al., “Retrieving Multimodal Information for Augmented Generation: A Survey”, arXiv preprint arXiv:2303.10868v3 [cs.CL] Dec. 1, 2023, 21 pages. [cited by applicant]
Zheng, et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”, arXiv preprint arXiv:2306.05685v4 [cs.CL] Dec. 24, 2023, 29 pages. [cited by applicant]
Zhuang, et al., “Toolqa: A Dataset for LLM Question Answering with External Tools”, arXiv preprint arXiv:2306.13304v1 [cs.CL] Jun. 23, 2023, 25 pages. [cited by applicant]