IP Library › Granted Patent US 12,265,567
Granted Patent B2
US 12,265,567 · App. 17/733,351 · Granted Apr 1, 2025

Artificial intelligence assisted originality evaluator

Inventors: Yinghao Ma (Potomac, MD); Chengmin Jiang (Potomac, MD); Sonja Krane (Blue Bell, PA); Jofia Jose Prakash (Aldie, VA); Utpal Tejookaya (Springfield, VA); Jeroen Van Prooijen (Fairfax, VA); Jonathan Hansford (Burke, VA); Jinglei Li (Falls Church, VA); Wallace Scott (Fairfax Station, VA)
Assignee: American Chemical Society
G06F16/3347G06F16/328
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,567
App. No.
17/733,351
Granted
Apr 1, 2025
Kind
B2
Abstract

A method is disclosed, involving converting each structured text document stored in a database into one or more vectors, using the vectors of the structured text documents stored in a database to create a similarity search index, then for each structured text document from the database, searching the search index using the one or more vectors of the structured text document in order to generate a list of N other structured text document from the database similar to the structured text document based on said search; an storing each list of N other structured text document from the database similar to the structured text document in a table.

Claims (83)

1. A method comprising:

converting each structured text document stored in a database into one or more vectors, each structured text document in the database having a title, an abstract, and an author;

using the vectors of the structured text documents stored in a database to create one or more similarity search index;

for each structured text document from the database, searching the search index using the one or more vectors of the structured text document;

for each structured text document from the database, generating a list of at least N other structured text documents from the database similar to the structured text document based on said search, wherein N equals 50; and

storing each list of at least N other structured text documents from the database similar to the structured text document in a table.

2. The method of claim 1 , further comprising:

receiving a new structured text document, the new structured text document having a title, an abstract, and an author;

converting the new structured text document into one or more vectors; searching the search index using the one or more vectors of the new structured text document;

generating a second list of at least N structured text documents from the database similar to the new structured text document using the said search index; and

storing the second list of at least N structured text documents from the database similar to the new structured text document in a second table.

3. The method of claim 1 , wherein converting each structured text document stored in the database into one or more vectors comprises converting the structured text document into one or more vectors by means of SPECTER embedding.

4. The method of claim 3 , further comprising:

receiving a new structured text document, the new structured text document having a title, an abstract, and an author;

converting the new structured text document into one or more new vectors by means of SPECTER embedding;

searching the search index using the one or more new vectors of the new structured text document;

generating a second list of at least N structured text documents from the database similar to the new structured text document using the said search index; and

storing the second list of at least N structured text documents from the database similar to the new structured text document in a second table.

5. The method of claim 1 , wherein:

each structured text document stored in the database is associated with a full text.

6. The method of claim 5 , further comprising:

receiving a new structured text document, the new structured text document having a title, an abstract, an author, and a full text;

converting the new structured text document into one or more new vectors;

searching the search index using the one or more new vectors of the new structured text document;

generating a second list of N structured text documents from the database similar to the new structured text document using the said search index; and

storing the second list of at least N structured text documents from the database similar to the new structured text document in a second table.

7. The method of claim 1 , wherein:

each structured text document stored in the database is associated with a metadata.

8. The method of claim 7 , further comprising:

receiving a new structured text document, the new structured text document having a title, an abstract, an author, and a metadata;

converting the new structured text document into one or more new vectors; searching the search index using the one or more new vectors;

generating a second list of at least N structured text documents from the database similar to the new structured text document based on said search; and

storing the second list of at least N structured text documents from the database similar to the new structured text document in a second table.

9. The method of claim 1 , wherein:

converting each structured text document stored in the database into one or more vectors comprises converting each structured text document using a natural language processing algorithm; and

generating the list of at least N other structured text documents from the database similar to the structured text document based on said search comprises determining a similarity based on one or more of an Euclidian distance, a Gaussian distance, or a cosine similarity.

10. The method of claim 1 , wherein the database is a collection of training data.

11. The method of claim 1 , wherein the search index is one of a flat index type, a locally sensitive hash (LSH) type, an inverted file index type (IVF), or a Hierarchical Navigable Small World (HNSW) graph.

12. The method of claim 1 , further comprising querying the database to compile a list of authors of each structured text document on the list of at least N structured text documents.

13. A system for identifying similar structured text documents to another structured text document, comprising:

at least one processor, and

at least one non-transitory computer readable media storing instructions configured to cause the processor to perform operations comprising:

converting each structured text document stored in a database into one or more vectors, each structured text document in the database having a title, an abstract, and an author;

using the vectors of the structured text documents stored in a database to create a similarity search index;

for each structured text document from the database, searching the search index using the one or more vectors of the structured text document

for each structured text document from the database, generating a list of at least N other structured text documents from the database similar to the structured text document based on said search, wherein N is 50; and

storing each list of at least N other structured text documents from the database similar to the structured text document in a table.

14. The system of claim 13 , the operations further comprising:

receiving a new structured text document, the new structured text document having a title, an abstract, and an author;

converting the new structured text document into one or more new vectors;

searching the search index using the one or more new vectors of the new structured text document;

generating a second list of at least N structured text documents from the database similar to the new structured text document using the said search index; and

storing the second list of at least N other structured text documents from the database similar to the new structured text document in a second table.

15. The system of claim 13 , the operations further comprising:

converting the structured text document into one or more vectors by means of SPECTER embedding.

16. The system of claim 15 , the operations further comprising:

receiving a new structured text document, the new structured text document having a title, an abstract, and an author;

converting the new structured text document into one or more new vectors by means of SPECTER embedding;

searching the search index using the one or more new vectors of the new structured text document;

generating a second list of at least N structured text documents from the database similar to the new structured text document using the said search index; and

storing the second list of at least N other structured text documents from the database similar to the new structured text document in a second table.

17. The system of claim 13 , wherein:

each structured text document stored in the database is associated with a full text.

18. The system of claim 17 , the operations further comprising:

receiving a new structured text document, the new structured text document having a title, an abstract, an author, and a full text;

converting the new structured text document into one or more new vectors;

searching the search index using the one or more new vectors of the new structured text document;

generating a second list of N structured text document from the database similar to the new structured text document using the said search index; and

storing the second list of N structured text documents from the database similar to the new structured text document in a second table.

19. The system of claim 13 , wherein:

each structured text document stored in the database is associated with a metadata.

20. The system of claim 19 , the operations further comprising:

receiving a new structured text document, the new structured text document having a title, an abstract, an author, and a metadata;

converting the new structured text document into one or more new vectors;

comparing the one or more new vectors of each structured text document to the vectors of the structured text documents from the database; and

generating a second list of at least N structured text document from the database similar to the new structured text document based on said comparison;

storing the second list of at least N structured text documents from the database similar to the new structured text document in a second table.

21. The system of claim 13 , wherein:

converting each structured text document stored in the database into one or more vectors comprises converting each structured text document using a natural language processing algorithm; and

generating the list of at least N other structured text documents from the database similar to the structured text document based on said search comprises determining a similarity based on one or more of an Euclidian distance, a Gaussian distance, or a cosine similarity.

22. The system of claim 13 , wherein the database is a collection of training data.

23. The system of claim 13 , wherein the search index is one of a flat index type, a locally sensitive hash (LSH) type, an inverted file index type (IVF), or a Hierarchical Navigable Small World (HNSW) graph.

24. The system of claim 13 , further comprising querying the database to compile a list of authors of each structured text document on the list of at least N structured text documents.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2022
From: MA, YINGHAO; JIANG, CHENGMIN; KRANE, SONJA; PRAKASH, JOFIA JOSE; TEJOOKAYA, UTPAL; VAN PROOIJEN, JEROEN; HANSFORD, JONATHAN; LI, JINGLEI; SCOTT, WALLACE
To: AMERICAN CHEMICAL SOCIETY
Reel/Frame 061397/0989 →
Continuity (2)
Provisional Application 63181560 · Apr 29, 2021
Related Publication 20220350828A1 · Nov 3, 2022
References Cited (17)
US 9483532B1 · Zhang et al. · 2016 [cited by applicant]
US 11531707B1 · Samdani · 2022 [cited by examiner]
US 20130198181A1 · Amer-Yahia et al. · 2013 [cited by applicant]
US 20170212882A1 · Rollins et al. · 2017 [cited by applicant]
US 20180300415A1 · Rehurek · 2018 [cited by applicant]
US 20190332644A1 · Freed et al. · 2019 [cited by applicant]
US 20200050638A1 · Hancock · 2020 [cited by applicant]
US 20200293608A1 · Nelson et al. · 2020 [cited by applicant]
US 20200380989A1 · Antunes et al. · 2020 [cited by applicant]
US 20210065041A1 · Gopalan et al. · 2021 [cited by applicant]
Gelman, et al. Toward a Robust Method for Understanding the Replicability of Research. SDU@ AAAI, published Mar. 21, 2021 (eight pages). [cited by applicant]
International Search Report for PCT Application No. PCT/US 22/26931, dated Aug. 16, 2022 (seven pages). [cited by applicant]
International Search Report for PCT Application No. PCT/US 22/26916, dated Aug. 16, 2022 (nine pages). [cited by applicant]
International Search Report for PCT Application No. PCT/US 22/26910, dated Aug. 16, 2022 (eight pages). [cited by applicant]
Office Action from Canadian Intellectual Property Office in Canadian patent application No. 3,172,934, mailed Jan. 24, 2024 (5 pages). [cited by applicant]
Office Action from Canadian Intellectual Property Office in Canadian patent application No. 3,172,992, mailed Jan. 26, 2024 (4 pages). [cited by applicant]
Office Action from Canadian Intellectual Property Office in Canadian patent application No. 3,172,963, mailed Jan. 25, 2024 (4 pages). [cited by applicant]