IP Library Granted Patent US 12,189,670
Granted Patent B2
US 12,189,670 · App. 18/344,205 · Granted Jan 7, 2025

Systems and methods for systematic literature review

Inventors: Reza Jafar (Vancouver, CA); Anna Forsythe (Miami, FL); Kristian Thorlund (Vancouver, CA); Junhan Liu (Toronto, CA)
Assignee: Cytel Inc.
G06F16/3346G06F16/3334G06F16/35
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,189,670
App. No.
18/344,205
Granted
Jan 7, 2025
Kind
B2
Abstract

Publication pre-screening may include the use of a trained model. A trained language model may be fine-tuned on a question-and-answer task and may be configured to receive a question that includes inclusion and exclusion criteria for a publication. The question may be formulated to include context information such as a title and abstract of the publication. An output of the model may be used to determine a selection of the publication.

Claims (74)

1. A computer-implemented method for automated systematic literature review, comprising:

obtaining a set of inclusion criteria and a set of exclusion criteria for a set of categories, the set of categories includes a population category, an intervention category, a study design category, and an outcome category;

obtaining data for a first publication of a study from a first database;

for each category in the set of categories, formulating a question based on the set of inclusion criteria, the set of exclusion criteria, and the data of the first publication;

for each category in the set of categories, generating an input to a trained language model, wherein each input includes the question;

processing the set of inputs with the trained language model to generate a set of probability outputs, wherein the trained language model is fine-tuned on a question-and-answer task;

determining a selection score by evaluating the set of probability outputs using a scoring function;

marking the first publication for selection based on the selection score;

obtaining second data for a second publication of a study from a second database;

determining if the second publication is a duplicate of the first publication; and

in response to determining that the second publication is the duplicate of the first publication, rejecting one of the first publication or the second publication based on a hierarchy rating of the first database and the second database.

2. The method of claim 1 , wherein the question has a yes or no answer.

3. The method of claim 1 , wherein the data of the first publication includes a title of the first publication and an abstract of the first publication.

4. A computer-implemented method for automated systematic literature review, comprising:

obtaining a set of inclusion criteria and a set of exclusion criteria for a set of categories, the set of categories includes a population category, an intervention category, a study design category, and an outcome category;

obtaining data for a first publication of a study from a first database;

generating inclusion keywords, wherein the inclusion keywords are generated based on the set of inclusion criteria;

generating exclusion keywords, wherein the exclusion keywords are generated based on the set of exclusion criteria;

for each category in the set of categories, formulating a question based on the set of inclusion criteria, the set of exclusion criteria, the inclusion keywords, the exclusion keywords, and the data for the first publication;

for each category in the set of categories, generating an input to a trained language model, wherein each input includes the question;

processing the set of inputs with the trained language model to generate a set of probability outputs, wherein the trained language model is fine-tuned on a question-and-answer task;

determining a selection score by evaluating the set of probability outputs using a scoring function; and

marking the first publication for selection based on the selection score.

5. The method of claim 4 , further comprising:

determining a frequency of occurrence of the inclusion keywords and the exclusion keywords in the data of the first publication; and

ordering the inclusion keywords and the exclusion keywords based on the frequency of occurrence.

6. The method of claim 1 , wherein the scoring function is based on a hierarchy of categories in the set of categories.

7. A system for automated systematic literature review, comprising:

an input generator configured to:

obtain a set of inclusion criteria and a set of exclusion criteria for a set of categories, the set of categories includes a population category, an intervention category, a study design category, and an outcome category;

generate inclusion keywords, wherein the inclusion keywords are generated based on the set of inclusion criteria;

generate exclusion keywords, wherein the exclusion keywords are generated based on the set of exclusion criteria; and

obtain data for a first publication of a study from a first database;

a question formulation module configured to:

for each category in the set of categories, formulate a question based on the set of inclusion criteria, the set of exclusion criteria, and the data of the first publication; and

for each category in the set of categories, generate an input, wherein each input includes the question;

a trained language model fine-tuned on a question-and-answer task configured to:

process the input to generate a set of probability outputs; and

a presentation module configured to:

determine a selection score by evaluating the set of probability outputs using a scoring function; and

mark the first publication for selection based on the selection score.

8. The system of claim 7 , wherein the question has a yes or no answer.

9. The system of claim 7 , wherein the data of the first publication includes a title of the first publication and an abstract of the first publication.

10. The system of claim 7 , wherein the input generator is further configured to:

determine a frequency of occurrence of the inclusion keywords and the exclusion keywords in the data of the first publication; and

order the inclusion keywords and the exclusion keywords based on the frequency of occurrence.

11. The system of claim 7 , wherein the scoring function is based on a hierarchy of categories in the set of categories.

12. One or more non-transitory, computer-readable media comprising computer-executable instructions that, when executed, cause at least one processor to perform actions comprising:

obtaining a set of inclusion criteria and a set of exclusion criteria for a set of categories, the set of categories includes a population category, an intervention category, a study design category, and an outcome category;

obtaining data for a first publication of a study from a first database;

for each category in the set of categories, formulating a question based on the set of inclusion criteria, the set of exclusion criteria, and the data of the first publication;

for each category in the set of categories, generating an input to a trained language model, wherein each input includes the question;

processing the set of inputs with the trained language model to generate a set of probability outputs, wherein the trained language model is fine-tuned on a question-and-answer task;

determining a selection score by evaluating the set of probability outputs using a scoring function;

marking the first publication for selection based on the selection score;

obtaining second data for a second publication of a study from a second database;

determining if the second publication is a duplicate of the first publication; and

in response to determining that the second publication is the duplicate of the first publication, rejecting one of the first publication or the second publication based on a hierarchy rating of the first database and the second database.

13. The one or more non-transitory, computer-readable media of claim 12 , wherein the question has a yes or no answer.

14. The one or more non-transitory, computer-readable media of claim 12 , wherein the data of the first publication includes a title of the first publication and an abstract of the first publication.

15. One or more non-transitory, computer-readable media comprising computer-executable instructions that, when executed, cause at least one processor to perform actions comprising:

obtaining a set of inclusion criteria and a set of exclusion criteria for a set of categories, the set of categories includes a population category, an intervention category, a study design category, and an outcome category;

generating inclusion keywords, wherein the inclusion keywords are generated based on the set of inclusion criteria;

generating exclusion keywords, wherein the exclusion keywords are generated based on the set of exclusion criteria;

obtaining data for a first publication of a study from a first database;

for each category in the set of categories, formulating a question based on the set of inclusion criteria, the set of exclusion criteria, the inclusion keywords, the exclusion keywords, and the data for the first publication;

for each category in the set of categories, generating an input to a trained language model, wherein each input includes the question;

processing the set of inputs with the trained language model to generate a set of probability outputs, wherein the trained language model is fine-tuned on a question-and-answer task;

determining a selection score by evaluating the set of probability outputs using a scoring function; and

marking the first publication for selection based on the selection score.

16. The one or more non-transitory, computer-readable media of claim 15 , further comprising instructions that cause at least one processor to perform actions comprising:

determining a frequency of occurrence of the inclusion keywords and the exclusion keywords in the data of the first publication; and

ordering the inclusion keywords and the exclusion keywords based on the frequency of occurrence.

17. The one or more non-transitory, computer-readable media of claim 12 , wherein the scoring function is based on a hierarchy of categories in the set of categories.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 2, 2024
From: JAFAR, REZA; FORSYTHE, ANNA; THORLUND, KRISTIAN; LIU, JUNHAN
To: CYTEL INC.
Reel/Frame 065996/0185 →
Continuity (2)
Provisional Application 63367277 · Jun 29, 2022
Related Publication 20240004910A1 · Jan 4, 2024
References Cited (19)
US 8799236B1 · Azari et al. · 2014 [cited by applicant]
US 11080483B1 · Islam · 2021 [cited by examiner]
US 11328796B1 · Jain · 2022 [cited by examiner]
US 20110010177A1 · Nakano et al. · 2011 [cited by applicant]
US 20150358270A1 · Jacobs et al. · 2015 [cited by applicant]
US 20170235848A1 · Van Dusen · 2017 [cited by examiner]
US 20170344886A1 · Tong · 2017 [cited by examiner]
US 20180018355A1 · Toivanen et al. · 2018 [cited by applicant]
US 20200401661A1 · Kota · 2020 [cited by examiner]
US 20210034812A1 · Mezaoui · 2021 [cited by examiner]
US 20210319907A1 · Harley · 2021 [cited by examiner]
US 20220100958A1 · Kiazand · 2022 [cited by examiner]
US 20220327356A1 · Rossiello · 2022 [cited by examiner]
US 20220375539A1 · Alvarez · 2022 [cited by examiner]
US 20230377748A1 · Yang · 2023 [cited by examiner]
US 20230377749A1 · Berisha · 2023 [cited by examiner]
US 20240004910A1 · Jafar · 2024 [cited by examiner]
WO 2024006431A1 · 2024 [cited by applicant]
PCT/US2023/026569 , “International Application Serial No. PCT/US2023/026569, International Search Report and Written Opinion mailed Sep. 29, 2023”, Cytel Inc., 17 pages. [cited by applicant]