IP Library › Granted Patent US 12,505,145
Granted Patent B2
US 12,505,145 · App. 18/510,121 · Granted Dec 23, 2025

Method of classifying a very large corpus of documents

Inventors: Aidan Randle-Conde (Manchester, GB); Jimmie Weiss (Bad Rappenau, DE); Dave Ruel (Portland, OR); Julien Masanes (Paris, FR)
Assignee: HANZO LTD
G06F16/355G06F40/205G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,145
App. No.
18/510,121
Granted
Dec 23, 2025
Kind
B2
Abstract

A method for sorting candidate documents into several sets associated with a reference document, each document stored by a client device memory, wherein the method includes a device processor performing for each reference document, generating a first prompt for a first large language model requesting generation of at least one question determining the relevance of a candidate document to a reference document; (b) for each candidate document, generating at least one second prompt for a second large language model requesting the answer to at least one reference question; assigning each candidate document to the set associated with a reference document as a function of the value(s) that has been received for a second prompt containing a reference question associated to the reference document.

Claims (42)

1 . A method for sorting a plurality of candidate documents into several sets each associated with a reference document, each candidate and reference document being stored by a memory of a client device, wherein the method comprises performing, by a processor of the device, steps of:

(a) for each reference document

generating a first prompt for a first large language model requesting the generation of at least one question aiming to determine whether a candidate document is relevant to the content of the reference document, the first prompt containing a text of the reference document;

transmitting the generated first prompt to a first server implementing the first large language model;

receiving in reply to the first prompt at least one reference question corresponding to the reference document;

(b) for each candidate document,

generating at least one second prompt for a second large language model requesting the answer to at least one of the reference questions for the candidate document, the second prompt containing the reference question and a text of the candidate document;

transmitting each generated second prompt to a second server implementing the second large language model;

receiving in reply to each second prompt a value representative of the answer to the reference question;

(c) assigning each candidate document to the set associated with a reference document as a function of the value(s) that has been received for a second prompt containing a reference question associated to the reference document,

wherein the text of the candidate document contained in the second prompt is the entire candidate document if it fits within performant context-window size limits of the second large language model, or else a chunk of the candidate document fitting within performant context-window size limits of the second large language model.

2 . The method according to claim 1 , wherein generating a first prompt includes inserting the text of the reference document into a generic template of first prompt, and generating a second prompt includes inserting the reference question and the text of the candidate document into a generic template of second prompt.

3 . The method according to claim 1 , wherein the first prompt requests the generation of a predefined number of questions, the predefined number of reference questions corresponding to the reference document being received in reply to the first prompt.

4 . The method according to claim 3 , wherein step (a) comprises merging two reference questions corresponding to the same reference document into one if it is possible.

5 . The method according to claim 1 , wherein the first large language model has more parameters than the second large language model.

6 . The method according to claim 1 , wherein step (a) previously comprises parsing the reference document with the first large language model so as to identify a summary part of the reference document, the text of the reference document contained in the first prompt being the summary part of the reference document.

7 . The method according to claim 1 , wherein step (a) previously comprises parsing the reference document with the first large language model so as to identify parts of the reference documents, each reference question received for the reference document being mapped on to one or more of the parts of the reference document.

8 . The method according to claim 1 , wherein the second prompt contains several entire candidate documents at once if it fits within performant context-window size limits of the second large language model.

9 . The method according to claim 1 , wherein step (b) comprises, if the value replied to a second prompt containing an entire candidate document is representative of the answer yes, restarting step (b) with several second prompts containing only a part of the candidate document.

10 . The method according to claim 1 , wherein the value representative of the answer to the reference question is a boolean or a score representative of the probability that the answer to the reference question is yes.

11 . A client device for sorting a plurality of candidate documents into several sets each associated with a reference document, comprising a memory storing each candidate and reference document, wherein the method comprises a processor configured to implement:

(a) for each reference document

generating a first prompt for a first large language model requesting the generation of at least one question aiming to determine whether a candidate document is relevant to the content of the reference document, the first prompt containing a text of the reference document;

transmitting each generated first prompt to a first server implementing the first large language model;

receiving in reply to the first prompt at least one reference question corresponding to the reference document;

(b) for each candidate document,

generating at least one second prompt for a second large language model requesting the answer to at least one of the reference questions for the candidate document, the second prompt containing the reference question and a text of the candidate document;

transmitting each generated second prompt to a second server implementing the second large language model,

receiving in reply to each second prompt a value representative of the answer to the reference question;

(c) assigning each candidate document to the set associated with a reference document as a function of the values(s) that has been received for a second prompt containing a reference question associated to the reference document,

wherein the text of the candidate document contained in the second prompt is the entire candidate document if it fits within performant context-window size limits of the second large language model, or else a chunk of the candidate document fitting within performant context-window size limits of the second large language model.

12 . A non-transitory computer-readable medium comprising program code instructions stored thereon for implementing a method for sorting a plurality of candidate documents into several sets each associated with a reference document, each candidate and reference document being stored by a memory of a client device, wherein the method comprises performing, by a processor of the device, steps of:

(a) for each reference document

generated a first prompt for a first large language model requesting the generation of at least one question aiming to determine whether a candidate document is relevant to the content of the reference document, the first prompt containing a text of the reference document;

transmitting each generated first prompt to a first server implementing the first large language model;

receiving in reply to the first prompt at least one reference question corresponding to the reference document;

(b) for each candidate document,

generating at least one second prompt for a second large language model requesting the answer to at least one of the reference questions for the candidate document, the second prompt containing the reference question and a text of the candidate document;

transmitting each generated second prompt to a second server implementing the second large language model,

receiving in reply to each second prompt a value representative of the answer to the reference question;

(c) assigning each candidate document to the set associated with a reference document as a function of the values(s) that has been received for a second prompt containing a reference question associated to the reference document,

wherein the text of the candidate document contained in the second prompt is the entire candidate document if it fits within performant context-window size limits of the second large language model, or else a chunk of the candidate document fitting within performant context-window size limits of the second large language model.

Assignments (4)
SECURITY INTEREST Recorded Sep 17, 2024
From: HANZO ARCHIVES, INC.
To: ASHGROVE CAPITAL LLP
Reel/Frame 068607/0677 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE ASSIGNEE'S RESIDENCE PREVIOUSLY RECORDED ON REEL 66732 FRAME 591. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 3, 2024
From: RANDLE-CONDE, AIDAN; WEISS, JIMMIE; RUEL, DAVE; MASANES, JULIEN
To: HANZO LTD
Reel/Frame 067055/0521 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE CITY PREVIOUSLY RECORDED AT REEL: 66202 FRAME: 788. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 5, 2024
From: RANDLE-CONDE, AIDAN; WEISS, JIMMIE; RUEL, DAVE; MASANES, JULIEN
To: HANZO LTD
Reel/Frame 066732/0591 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2024
From: RANDLE-CONDE, AIDAN; WEISS, JIMMIE; RUEL, DAVE; MASANES, JULIEN
To: HANZO LTD
Reel/Frame 066202/0788 →
Continuity (1)
Related Publication 20250156468A1 · May 15, 2025
References Cited (45)
US 6415250B1 · van den Akker · 2002 [cited by examiner]
US 6640224B1 · Chakrabarti · 2003 [cited by examiner]
US 7747427B2 · Lee · 2010 [cited by examiner]
US 8214363B2 · Chaudhary · 2012 [cited by examiner]
US 8327265B1 · Vogel · 2012 [cited by examiner]
US 11436529B1 · Londeree · 2022 [cited by examiner]
US 11860914B1 · Qadrud-Din · 2024 [cited by examiner]
US 11861321B1 · O'Kelly · 2024 [cited by examiner]
US 11972223B1 · DeFoor · 2024 [cited by examiner]
US 12197483B1 · Shmukler · 2025 [cited by examiner]
US 20040088157A1 · Lach · 2004 [cited by examiner]
US 20040148155A1 · Vogel · 2004 [cited by examiner]
US 20050114327A1 · Kumamoto · 2005 [cited by examiner]
US 20060155662A1 · Murakami · 2006 [cited by examiner]
US 20070162272A1 · Koshinaka · 2007 [cited by examiner]
US 20090274376A1 · Selvaraj · 2009 [cited by examiner]
US 20100299139A1 · Ferrucci · 2010 [cited by examiner]
US 20120278266A1 · Naslund et al. · 2012 [cited by applicant]
US 20130138430A1 · Eden · 2013 [cited by examiner]
US 20140379761A1 · Adamson · 2014 [cited by examiner]
US 20150154305A1 · Lightner · 2015 [cited by examiner]
US 20180025075A1 · Beller · 2018 [cited by examiner]
US 20200142856A1 · Neelamana · 2020 [cited by examiner]
US 20200143257A1 · Neelamana · 2020 [cited by examiner]
US 20210165807A1 · Asaf · 2021 [cited by examiner]
US 20210248420A1 · Zhong · 2021 [cited by examiner]
US 20210390297A1 · Sakaguchi · 2021 [cited by examiner]
US 20210398025A1 · Yamamoto · 2021 [cited by examiner]
US 20220058496A1 · Rusk · 2022 [cited by examiner]
US 20220108126A1 · Schieber · 2022 [cited by examiner]
US 20220414369A1 · Makani · 2022 [cited by examiner]
US 20230015667A1 · Stadermann · 2023 [cited by examiner]
US 20230078263A1 · Seledkin · 2023 [cited by examiner]
US 20230196017A1 · Geiman · 2023 [cited by examiner]
US 20230315790A1 · Zeng · 2023 [cited by examiner]
US 20240095460A1 · Xu · 2024 [cited by examiner]
US 20240160900A1 · Smith · 2024 [cited by examiner]
US 20240249191A1 · Bowman · 2024 [cited by examiner]
US 20240289561A1 · Qadrud-Din · 2024 [cited by examiner]
US 20240356875A1 · Medalion · 2024 [cited by examiner]
US 20240386037A1 · Chockalingam · 2024 [cited by examiner]
US 20240394286A1 · Honke · 2024 [cited by examiner]
US 20250005058A1 · Khosla · 2025 [cited by examiner]
US 20250005276A1 · Bhat · 2025 [cited by examiner]
US 20250156468A1 · Randle-Conde · 2025 [cited by examiner]