IP Library Granted Patent US 12,417,352
Granted Patent B1
US 12,417,352 · App. 18/327,676 · Granted Sep 16, 2025

Systems and methods for using a large language model for large documents

Inventors: Vineeth Chinmaya Murthy (Bengaluru, IN); Rafal Powalski (Warsaw, PL); Atinderpal Singh (Bathinda, IN); Sławomir Jan Biel (Warsaw, PL); Hariharan Thirugnanam (Bangalore, IN); Bartosz Topolski (Warsaw, IN)
Assignee: Instabase, Inc.
G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,352
App. No.
18/327,676
Granted
Sep 16, 2025
Kind
B1
Abstract

Systems and methods for using a machine learning model for a set of one or more documents are disclosed. Exemplary implementations may: create a set of document segments from the set of one or more documents; create a set of semantic vectors; create a query vector that semantically represents a query from a user; determine a subset of the set of semantic vectors based on at least two different comparisons involving the query vector; create a combination of the individual document segments that are associated with the subset of the set of semantic vectors; provide a prompt to the machine learning model, using the created combination of the individual document segments as context; present replies from the machine learning model, and/or perform other steps.

Claims (37)

1. A system configured for using a machine learning model to extract information from a set of one or more documents, wherein the set of one or more documents spans at least 200 pages, the system comprising:

one or more hardware processors configured by machine-readable instructions to:

create a set of document segments from the set of one or more documents;

create, using the machine learning model, a set of semantic vectors, wherein individual semantic vectors are associated with individual document segments;

store the set of semantic vectors in a vector database;

effectuate a presentation of a user interface, the user interface being configured to obtain a query from a user, wherein the query pertains to extracting particular information from the set of one or more documents;

create, using the machine learning model, a query vector that semantically represents the query;

determine a subset of the set of semantic vectors, wherein the determination is based on both:

(i) a first type of comparison of the set of semantic vectors with the query vector, and

(ii) a second type of comparison of the set of semantic vectors with the query vector, and wherein the determination of the subset of the set of semantic vectors is further based on relative positions of the individual document segments in proximity to other individual document segments based on at least one of (i) and/or (ii);

create a combination of the individual document segments that are associated with the subset of the set of semantic vectors such that a quantity of information represented by the subset of the set of semantic vectors is within a capacity of the machine learning model to use as context;

provide a prompt to the machine learning model, using the created combination of the individual document segments as context, wherein the prompt is based on the query; and

present to the user, through the user interface, one or more replies obtained from the machine learning model in reply to the prompt, wherein the one or more replies are related to the particular information as extracted from the set of one or more documents.

2. The system of claim 1 , wherein the first type of comparison compares similarity between an individual semantic vector with the query vector, wherein the similarity represents natural language searching.

3. The system of claim 1 , wherein the second type of comparison compares an individual semantic vector with the query vector in a manner that represents keyword searching.

4. The system of claim 1 , wherein the determination of the subset of the set of semantic vectors is further based on: (iii) absolute positions of individual document segments within the set of one or more documents.

5. The system of claim 1 , wherein the machine learning model is limited to a predetermined number of tokens as the context for the prompt, and wherein the combination of the individual document segments is created such that the predetermined number of tokens is not exceeded.

6. The system of claim 1 , wherein the machine learning model is a large language model.

7. The system of claim 6 , wherein the large language model has been trained on at least a million documents, wherein the large language model includes a neural network using over a billion parameters and/or weights.

8. The system of claim 7 , wherein the large language model is based on or derived from Generative Pre-trained Transformer 3 (GPT3) or a successor of Generative Pre-trained Transformer 3 (GPT3).

9. A computer-implemented method using one or more hardware processors to extract information from a set of one or more documents through a machine learning model, wherein the set of one or more documents spans at least 200 pages, the method comprising:

creating a set of document segments from the set of one or more documents;

creating, using the machine learning model, a set of semantic vectors, wherein individual semantic vectors are associated with individual document segments;

storing the set of semantic vectors in a vector database;

effectuating a presentation of a user interface, wherein the user interface obtains a query from a user, wherein the query pertains to extracting particular information from the set of one or more documents;

creating, using the machine learning model, a query vector that semantically represents the query;

determining a subset of the set of semantic vectors, wherein the determination is based on both (i) a first type of comparison of the set of semantic vectors with the query vector, and (ii) a second type of comparison of the set of semantic vectors with the query vector, and wherein the determination of the subset of the set of semantic vectors is further based on relative positions of the individual document segments in proximity to other individual document segments based on at least one of (i) and/or (ii);

creating a combination of the individual document segments that are associated with the subset of the set of semantic vectors such that a quantity of information represented by the subset of the set of semantic vectors is within a capacity of the machine learning model to use as context;

providing a prompt to the machine learning model, using the created combination of the individual document segments as context, wherein the prompt is based on the query; and

presenting to the user, through the user interface, one or more replies obtained from the machine learning model in reply to the prompt, wherein the one or more replies are related to the particular information as extracted from the set of one or more documents.

10. The computer-implemented method of claim 9 , wherein the first type of comparison compares similarity between an individual semantic vector with the query vector, wherein the similarity represents natural language searching.

11. The computer-implemented method of claim 9 , wherein the second type of comparison compares an individual semantic vector with the query vector in a manner that represents keyword searching.

12. The computer-implemented method of claim 9 , wherein the determination of the subset of the set of semantic vectors is further based on: (iii) absolute positions of individual document segments within the set of one or more documents.

13. The computer-implemented method of claim 9 , wherein the machine learning model is limited to a predetermined number of tokens as the context for the prompt, and wherein the combination of the individual document segments is created such that the predetermined number of tokens is not exceeded.

14. The computer-implemented method of claim 9 , wherein the machine learning model is a large language model.

15. The computer-implemented method of claim 14 , wherein the large language model has been trained on at least a million documents, wherein the large language model includes a neural network using over a billion parameters and/or weights.

16. The computer-implemented method of claim 15 , wherein the large language model is based on or derived from Generative Pre-trained Transformer 3 (GPT3) or a successor of Generative Pre-trained Transformer 3 (GPT3).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2023
From: MURTHY, VINEETH CHINMAYA; POWALSKI, RAFAL; THIRUGNANAM, HARIHARAN; SINGH, ATINDERPAL; BIEL, SLAWOMIR JAN; TOPOLSKI, BARTOSZ
To: INSTABASE, INC.
Reel/Frame 064001/0370 →
References Cited (75)
US 5848184A · Taylor · 1998 [cited by applicant]
US 5898795A · Bessho · 1999 [cited by applicant]
US 7689431B1 · Carmel · 2010 [cited by applicant]
US 7720318B1 · Phinney · 2010 [cited by applicant]
US 7725423B1 · Pricer · 2010 [cited by applicant]
US 8254681B1 · Poncin · 2012 [cited by applicant]
US 9275030B1 · Fang · 2016 [cited by applicant]
US 9607058B1 · Gupta · 2017 [cited by applicant]
US 10642832B1 · Neumann · 2020 [cited by applicant]
US 10679089B2 · Annis · 2020 [cited by applicant]
US 11315353B1 · Cahn · 2022 [cited by applicant]
US 11947604B2 · Roitman · 2024 [cited by applicant]
US 11995394B1 · Morariu · 2024 [cited by applicant]
US 12182125B1 · Buniatyan · 2024 [cited by applicant]
US 20020064316A1 · Takaoka · 2002 [cited by applicant]
US 20040181749A1 · Chellapilla · 2004 [cited by applicant]
US 20040223648A1 · Hoene · 2004 [cited by applicant]
US 20050289182A1 · Pandian · 2005 [cited by applicant]
US 20080148144A1 · Tatsumi · 2008 [cited by applicant]
US 20080212901A1 · Castiglia · 2008 [cited by applicant]
US 20080291486A1 · Isles · 2008 [cited by applicant]
US 20090076935A1 · Knowles · 2009 [cited by applicant]
US 20090132590A1 · Huang · 2009 [cited by applicant]
US 20120072859A1 · Wang · 2012 [cited by applicant]
US 20120204103A1 · Stevens · 2012 [cited by applicant]
US 20140200880A1 · Neustel · 2014 [cited by applicant]
US 20140214732A1 · Carmeli · 2014 [cited by applicant]
US 20150012422A1 · Ceribelli · 2015 [cited by applicant]
US 20150169951A1 · Khintsitskiy · 2015 [cited by applicant]
US 20150169995A1 · Panferov · 2015 [cited by applicant]
US 20150278197A1 · Bogdanova · 2015 [cited by applicant]
US 20160014299A1 · Saka · 2016 [cited by applicant]
US 20160275526A1 · Becanovic · 2016 [cited by applicant]
US 20180189592A1 · Annis · 2018 [cited by applicant]
US 20180329890A1 · Ito · 2018 [cited by applicant]
US 20190138660A1 · White · 2019 [cited by applicant]
US 20190171634A1 · Nowakiewicz · 2019 [cited by applicant]
US 20190286900A1 · Pepe, Jr. · 2019 [cited by applicant]
US 20190340949A1 · Meisner · 2019 [cited by examiner]
US 20200004749A1 · Slezak · 2020 [cited by applicant]
US 20200089946A1 · Mallick · 2020 [cited by applicant]
US 20200104359A1 · Patel · 2020 [cited by applicant]
US 20200159848A1 · Yeo · 2020 [cited by applicant]
US 20200311349A1 · Balasubramanian · 2020 [cited by applicant]
US 20200320072A1 · Hormati · 2020 [cited by applicant]
US 20200364343A1 · Atighetchi · 2020 [cited by applicant]
US 20200379673A1 · Le Gallo-Bourdeau · 2020 [cited by applicant]
US 20210034621A1 · Patel · 2021 [cited by applicant]
US 20210258448A1 · Inoue · 2021 [cited by applicant]
US 20220164346A1 · Mitra · 2022 [cited by applicant]
US 20220398858A1 · Cahn · 2022 [cited by applicant]
US 20220414075A1 · Li · 2022 [cited by applicant]
US 20220414430A1 · Li · 2022 [cited by applicant]
US 20220414492A1 · Jezewski · 2022 [cited by applicant]
US 20230044564A1 · Jezewski · 2023 [cited by applicant]
US 20230315731A1 · Xu · 2023 [cited by applicant]
US 20230334889A1 · Cahn · 2023 [cited by applicant]
US 20230385261A1 · Siddiqui · 2023 [cited by applicant]
US 20240096125A1 · Yebes Torres · 2024 [cited by applicant]
US 20240202539A1 · Poirier · 2024 [cited by applicant]
US 20240221007A1 · Hormati · 2024 [cited by applicant]
US 20240311407A1 · Barron · 2024 [cited by applicant]
US 20240338361A1 · Hazel · 2024 [cited by applicant]
US 20250045314A1 · Madnani · 2025 [cited by applicant]
US 20250077527A1 · Vaughn · 2025 [cited by applicant]
US 20250086190A1 · Azarmi · 2025 [cited by applicant]
CN 117951274A · 2024 [cited by applicant]
CN 118332072A · 2024 [cited by applicant]
CN 118656482A · 2024 [cited by applicant]
CN 118939782A · 2024 [cited by applicant]
Chaudhuri et al., “Extraction of type style-based meta-information from imaged documents”, IJDAR (2001) 3: 138-149. (Year: 2001). [cited by applicant]
Doermann et al., “Image Based Typographic Analysis of Documents”, Proceedings of 2nd International Conference on Document Analysis and Recognition, pp. 769-773, 1993 IEEE. (Year: 1993). [cited by applicant]
Shafait (“Document image analysis with OCRopus,” IEEE 13th International Mulititopic Conference; Date of Conference: Dec. 14-15, 2009) (Year: 2009) 6 pages. [cited by applicant]
Singh et al. (A Proposed Approach for Character Recognition Using Document Analysis with OCR, Second InternationalConference on Intelligent Computing and Control Systems: Date of Conference: Jun. 14-15, 2018) (Year: 201… [cited by applicant]
Slavin et al., “Matching Digital Documents Based on OCR”, 2019 XXI International Conference Complex Systems: Control and Modeling Problems (CSCMP), pp. 177-181 , published on Sep. 1, 2019. (Year: 2019). [cited by applicant]
Cited By (2)
US 12,717,850 US 12,724,815