IP Library Granted Patent US 12,353,409
Granted Patent B2
US 12,353,409 · App. 18/640,448 · Granted Jul 8, 2025

Methods and systems for improved document processing and information retrieval

Inventors: Steven John Rennie (Yorktown Heights, NY); Marie Wenzel Meteer (Arlington, MA); David Nahamoo (Great Neck, NY); Dominique O'donnell (Portland, OR); Vaibhava Goel (Chappaqua, NY); Etienne Marcheret (White Plains, NY); Chul Sung (Fort Lee, NJ); Igor Roditis Jablokov (Raleigh, NC); Soonthorn Ativanichayaphong (New York, NY); Ajinkya Jitendra Zadbuke (Cambridge, MA); Carmi Rothberg (New York, NY); Ellen Eide Kislal (Leawood, KS)
Assignee: Pryon Incorporated
G06F16/24522G06F16/248G06F16/3329G06F40/30G06N5/04G06N20/00G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,353,409
App. No.
18/640,448
Granted
Jul 8, 2025
Kind
B2
Abstract

Disclosed are methods, systems, devices, apparatus, media, and other implementations that include a method for document processing (particularly for training of a machine learning question answering platform, and for ingestion of documents). The method includes obtaining a question dataset (e.g., either from public or private repositories of questions) comprising one or more source questions for document processing by a machine learning question-and-answer system that provides answer data in response to question data submitted by a user, modifying a source question from the question dataset to generate one or more augmented questions with equivalent semantic meanings as that of the source question, and processing a document with the one or more augmented questions.

Claims (58)

1. A method for document processing, the method comprising:

obtaining a question dataset that comprises one or more source questions for document processing by a machine-learning question-and-answer system that provides answer data in response to question data submitted by a user,

analyzing the source question to determine a specificity level for the source question, wherein analyzing the source question comprises determining the source question is overly verbose based on a comparison of the determined specificity level to one or more specificity threshold values,

modifying the source question from the question dataset to generate one or more augmented questions that have equivalent semantic meanings as that of the source question,

processing a document with the one or more augmented questions,

wherein modifying the source question comprises: simplifying the source question, in response to a determination that the source question is overly verbose, to exclude one or more semantic elements of the source question to generate a terse question with an equivalent semantic meaning to the source question, and

adding the one or more augmented questions to an augmented question dataset.

2. The method of claim 1 , wherein processing the document with the one or more augmented questions comprises one or more of:

training the machine learning question-and-answer system using the document and the one or more augmented questions; or

ingesting the document with the one or more augmented questions, subsequent to completion of the training of the machine learning question-and-answer system, to generate an ingested document that is searchable by the machine learning question-and-answer system.

3. The method of claim 1 , wherein analyzing the source question comprises one or more of: determining a word count for the source question, determining intent associated with the source question, or classifying the source question using a machine-learning question specificity model.

4. The method of claim 1 , wherein simplifying the source question comprises excluding the one or more semantic elements of the source question according to one or more of: a term-weighing scheme to assign values for words appearing in the source question, or one or more natural language processing (NPL) rules applied to the source question.

5. The method of claim 4 , wherein the term-weighing scheme comprises a term frequency-inverse document frequency (TF-IDF) scheme, wherein simplifying the source question comprises:

computing weight values for the words appearing in the source question according to the (TF-IDF) scheme; and

removing one or more of the words appearing in the source question based on the computed weight values, and subject to the NPL rules.

6. The method of claim 1 , wherein modifying the source question comprises:

generating a question variant message comprising the source question and one or more required output characteristics for a resultant augmented question paraphrasing of the source question; and

providing the question variant message to a generative large language model system configured to generate a resultant augmented question based on the source question and the one or more required output characteristics.

7. The method of claim 1 , further comprising:

iteratively modifying the one or more augmented questions to generate additional sets of augmented questions with similar semantic meanings as that of a preceding set of augmented questions.

8. A method for document processing, the method comprising:

obtaining a question dataset that comprises one or more source questions for document processing by a machine-learning question-and-answer system that provides answer data in response to question data submitted by a user,

analyzing the source question to determine a specificity level for the source question, wherein analyzing the source question comprises determining the source question is overly verbose based on a comparison of the determined specificity level to one or more specificity threshold values,

modifying the source question from the question dataset to generate one or more augmented questions that have equivalent semantic meanings as that of the source question, and

processing a document with the one or more augmented questions,

wherein modifying the source question comprises: simplifying the source question, in response to a determination that the source question is overly verbose, to exclude one or more semantic elements of the source question to generate a terse question with an equivalent semantic meaning to the source question, and

wherein the source question includes unstructured keywords, and wherein modifying the source question comprises expanding the source question, in response to a determination that the source question is overly terse, to include structural components for the unstructured keywords.

9. The method of claim 8 , wherein expanding the source question comprises:

determining a statement of intent and/or other contextual information associated with the keywords of the source question; and

adding to the source question semantic structural components determined based on the statement of intent and/or other contextual information.

10. A method for document processing, the method comprising:

obtaining a question dataset comprising one or more source questions for document processing by a machine learning question-and-answer system that provides answer data in response to question data submitted by a user,

analyzing the source question to determine specificity level for the source question,

wherein analyzing the source question comprises determining that the source question is overly terse based on a comparison of the determined specificity level to one or more specificity threshold values,

modifying the source question from the question dataset to generate one or more augmented questions with equivalent semantic meanings as that of the source question;

processing a document with the one or more augmented questions,

wherein the source question includes unstructured keywords, and wherein modifying the source question comprises expanding the source question, in response to a determination that the source question is overly terse, to include structural components for the unstructured keywords, and

adding the one or more augmented questions to an augmented question dataset.

11. The method of claim 10 , wherein processing the document with the one or more augmented questions comprises one or more of:

training the machine learning question-and-answer system using the document and the one or more augmented questions; or

ingesting the document with the one or more augmented questions, subsequent to completion of the training of the machine learning question-and-answer system, to generate an ingested document that is searchable by the machine learning question-and-answer system.

12. The method of claim 10 , wherein analyzing the source question comprises one or more of: determining a word count for the source question, determining intent associated with the source question, or classifying the source question using a machine-learning question specificity model.

13. The method of claim 10 , wherein analyzing the source question comprises:

determining the source question is one of overly verbose or overly terse based on a comparison of the determined specificity level to one or more specificity threshold values.

14. The method of claim 10 , wherein modifying the source question comprises:

simplifying the source question, in response to a determination that the source question is overly verbose, to exclude one or more semantic elements of the source question to generate a terse question with an equivalent semantic meaning to the source question.

15. The method of claim 14 , wherein simplifying the source question comprises excluding the one or more semantic elements of the source question according to one or more of: a term-weighing scheme to assign values for words appearing in the source question, or one or more natural language processing (NPL) rules applied to the source question.

16. The method of claim 15 , wherein the term-weighing scheme comprises a term frequency-inverse document frequency (TF-IDF) scheme, wherein simplifying the source question comprises:

computing weight values for the words appearing in the source question according to the (TF-IDF) scheme; and

removing one or more of the words appearing in the source question based on the computed weight values, and subject to the NPL rules.

17. The method of claim 10 , wherein expanding the source question comprises:

determining a statement of intent and/or other contextual information associated with the keywords of the source question; and

adding to the source question semantic structural components determined based on the statement of intent and/or other contextual information.

18. The method of claim 10 , wherein modifying the source question comprises:

generating a question variant message comprising the source question and one or more required output characteristics for a resultant augmented question paraphrasing of the source question; and

providing the question variant message to a generative large language model system configured to generate a resultant augmented question based on the source question and the one or more required output characteristics.

19. The method of claim 10 , further comprising:

iteratively modifying the one or more augmented questions to generate additional sets of augmented questions with similar semantic meanings as that of a preceding set of augmented questions.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Oct 31, 2025
From: PRYON INCORPORATED
To: FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 073438/0899 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2024
From: RENNIE, STEVEN JOHN; METEER, MARIE WENZEL; NAHAMOO, DAVID; O'DONNELL, DOMINIQUE; GOEL, VAIBHAVA; MARCHERET, ETIENNE; SUNG, CHUL; JABLOKOV, IGOR RODITIS; ATIVANICHAYAPHONG, SOONTHORN; ZADBUKE, AJINKYA JITENDRA; ROTHBERG, CARMI; KISLAL, ELLEN EIDE
To: PRYON INCORPORATED
Reel/Frame 067321/0958 →
Continuity (4)
Continuation PCTUS2023027320 · Jul 11, 2023
Provisional Application 63423527 · Nov 8, 2022
Provisional Application 63388046 · Jul 11, 2022
Related Publication 20240265041A1 · Aug 8, 2024
References Cited (46)
US 7254545B2 · Falcon · 2007 [cited by examiner]
US 8024338B2 · Brei · 2011 [cited by examiner]
US 8478780B2 · Cooper · 2013 [cited by examiner]
US 11010284B1 · Santiago et al. · 2021 [cited by applicant]
US 11080336B2 · Van Dusen · 2021 [cited by applicant]
US 11561987B1 · Sager et al. · 2023 [cited by applicant]
US 20070150458A1 · Chung et al. · 2007 [cited by applicant]
US 20070288439A1 · Rappaport et al. · 2007 [cited by applicant]
US 20080270380A1 · Ohrn et al. · 2008 [cited by applicant]
US 20110196852A1 · Srikanth et al. · 2011 [cited by applicant]
US 20120078873A1 · Ferrucci et al. · 2012 [cited by applicant]
US 20140304250A1 · Sankar et al. · 2014 [cited by applicant]
US 20150178853A1 · Byron et al. · 2015 [cited by applicant]
US 20160364377A1 · Krishnamurthy · 2016 [cited by applicant]
US 20170024465A1 · Yeh et al. · 2017 [cited by applicant]
US 20170060990A1 · Brown et al. · 2017 [cited by applicant]
US 20170308531A1 · Ma et al. · 2017 [cited by applicant]
US 20170344640A1 · Goldstein et al. · 2017 [cited by applicant]
US 20180046764A1 · Katwala et al. · 2018 [cited by applicant]
US 20180143975A1 · Casal et al. · 2018 [cited by applicant]
US 20180189630A1 · Boguraev et al. · 2018 [cited by applicant]
US 20180232376A1 · Zhu et al. · 2018 [cited by applicant]
US 20180260738A1 · Gabbai et al. · 2018 [cited by applicant]
US 20190042988A1 · Brown et al. · 2019 [cited by applicant]
US 20200257679A1 · Sheinin et al. · 2020 [cited by applicant]
US 20210022688A1 · Lee et al. · 2021 [cited by applicant]
US 20210097405A1 · McNeil et al. · 2021 [cited by applicant]
US 20210157975A1 · Gelosi · 2021 [cited by applicant]
US 20210287141A1 · Molloy et al. · 2021 [cited by applicant]
US 20210319066A1 · Boxwell et al. · 2021 [cited by applicant]
US 20210365481A1 · Chen · 2021 [cited by applicant]
US 20210397609A1 · Yurtsev et al. · 2021 [cited by applicant]
US 20210406264A1 · Nahamoo · 2021 [cited by examiner]
US 20210406735A1 · Nahamoo et al. · 2021 [cited by applicant]
US 20240135291A1 · Steenstra et al. · 2024 [cited by applicant]
US 20240154941A1 · Nagpal et al. · 2024 [cited by applicant]
CN 111897937A · 2020 [cited by applicant]
CN 112199254A · 2021 [cited by applicant]
CN 114090760A · 2022 [cited by applicant]
KR 100347799B1 · 2002 [cited by applicant]
KR 1020190079805A · 2019 [cited by applicant]
KR 1020190118744A · 2019 [cited by applicant]
KR 1020220083450A · 2022 [cited by applicant]
WO 8810470A1 · 1988 [cited by applicant]
Tianwei Xing et al., DeepSQA: Understanding Sensor Data via Question Answering. In Proceedings of the International Conference on Internet-of-Things Design and Implementation. Association for Computing Machinery, 106-11… [cited by examiner]
International Search Report and Written Opinion, PCT Application No. PCT/US2023/027230, mailed Nov. 20, 2023 (10 pages). [cited by applicant]