IP Library › Granted Patent US 11,630,870
Granted Patent B2
US 11,630,870 · App. 17/142,996 · Granted Apr 18, 2023

Academic search and analytics system and method therefor

Inventor: Tarek A. M. Abdunabi (Ottawa, CA)
G06F16/951G06F9/547G06F16/9538G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,630,870
App. No.
17/142,996
Granted
Apr 18, 2023
Kind
B2
Abstract

An apparatus and method for academic search and analytics insights have been provided. The apparatus includes an ingestion component, obtaining data from external heterogeneous sources, to produce ingested data; a processing component for processing the ingested data; a search and analytics component for executing search queries on the ingested data and generating analytics insights on returned search result; and a storage component for storing the ingested data, the storage component acting as a communication data bus for the ingestion component, the processing component and the search and analytics component. Corresponding server and network system are have been provided.

Claims (89)

1. An apparatus for academic search and analytics insights, comprising:

a non-transitory computer-readable storage medium, storing computer readable instructions for execution by a processor, comprising:

an ingestion component, obtaining data from external heterogeneous sources, to produce ingested data, wherein the ingestion component further comprises:

an application programming interface, API, module for automatically fetching data from external databases having API interfaces;

a crawlers module for crawling, in parallel, predefined websites for extracting predetermined attributes; and

an enrichment module for automatically obtaining additional or missing attributes of the ingested data from other sources;

wherein a single index has been used to coordinate an operation of the API module, the crawler module and the enrichment module, thereby guiding what data needs to be acquired and eliminating duplicate work;

a processing component for processing the ingested data, the processing component comprising:

a Natural Language Processing (NLP) module operable on the ingested data to generate a pre-processed ingested data, the NLP module being in communication with:

a machine learning (ML) modeling module, using the processor for grouping related data into topics and identifying respective keywords;

a communities and influencers detection module for identifying communities of researchers; and

a trends and discovery prediction module for predicting predetermined use cases based on historic data;

and

a search and analytics component for executing search queries on the pre-processed ingested data, and outputs of the machine learning modeling module, communities and influencers detection module, and trend and discovery prediction module, and generating the analytics insights on returned search results.

2. The apparatus of claim 1 , wherein the enrichment module is further configured to obtain latitude and longitude coordinates of a university, and authors affiliated with the university, by querying an external application programming interface of an online mapping tool.

3. The apparatus of claim 1 , wherein the ingestion component further comprises a format conversion module for converting the ingested data from the external heterogeneous sources into a common format, which is suitable for further processing by the NLP module.

4. The apparatus of claim 1 , wherein the ingestion component further comprises a cross-referencing module for linking related information from the ingested data into a single record per entity.

5. The apparatus of claim 1 , wherein the trends and discovery prediction module is configured to predict one or more of the following:

a sudden increase or decrease in a number of publications of a specific research topic;

a future number of publications per keywords or topic based on historic data; and

a future number of citations per paper.

6. The apparatus of claim 1 , wherein the communities of researchers comprise one or more of the following:

community of researchers who collaborate with each other; and

community of citations.

7. The apparatus of claim 1 , wherein the search and analytics component further comprises:

a search module for fetching the ingested data from a storage component based on a search query, to produce fetched data;

an analytics module for generating analytics insights based on the fetched data; and

a dashboard module for presenting the fetched data along with the analytics insights, while independently allowing accessibility to topics and keywords from the ML modeling module without a need of submitting a new search query.

8. The apparatus of claim 1 , further comprising a storage component for storing the ingested data, the storage component acting as a communication data bus for the ingestion component, the processing component and the search and analytics component.

9. A computer implemented method for academic search and analytics insights, comprising:

employing at least one hardware processor for:

obtaining data from external heterogeneous sources, to produce ingested data, wherein the obtaining further comprises:

automatically fetching data from external databases having application programming interfaces, APIs;

crawling, in parallel, predefined websites for extracting predetermined attributes;

enriching the fetching and the crawling, comprising automatically obtaining additional or missing attributes from other sources; and

using a single index for coordinating the fetching, the crawling and the enriching, thereby guiding what data needs to be acquired and eliminating duplicate work;

processing the ingested data, comprising:

Natural Language Processing (NLP) of the ingested data to generate a pre-processed ingested data, the NLP being in communication with:

machine learning modeling using the at least one hardware processor for grouping related data into topics and identifying respective keywords;

identifying communities and influencers among researchers; and

predicting predetermined use cases based on historic data;

and

executing search queries on the pre-processed ingested data and outputs of the machine learning modeling, the identifying, and the predicting, and generating the analytics insights on returned search results.

10. The method of claim 9 , wherein the enriching further comprises obtaining latitude and longitude coordinates of a university, and authors affiliated with the university, by querying an external application programming interface of an online mapping tool.

11. The method of claim 9 , wherein the obtaining further comprises converting the ingested data from the external heterogeneous sources into a common format suitable for further processing by the NLP module.

12. The method of claim 9 , wherein the obtaining further comprises a cross-referencing the ingested data, comprising linking related information from the ingested data into a single record per entity.

13. The method of claim 9 , wherein the predicting comprises one or more of the following:

predicting a sudden increase or decrease in a number of publications of a specific research topic;

predicting a future number of publications per keywords or topic based on historic data; and

predicting a future number of citations per paper.

14. The method of claim 9 , wherein the executing further comprises:

fetching the ingested data from a storage component, based on a search query, to produce fetched data;

generating analytics insights based on the fetched data; and

presenting the fetched data along with the analytics insights at a dashboard, while independently allowing accessibility to topics and keywords from ML modeling module, without a need of submitting a new search query.

15. The method of claim 9 , further comprising storing the ingested data in a storage component, the storage component acting as a communication data bus for the ingestion component, the processing component and the search and analytics component.

16. A server computer for academic search and analytics insights, comprising:

at least one hardware processor;

a non-transitory computer readable storage medium having computer executable instructions stored thereon for execution by the at least one hardware processor, causing the at least one hardware processor to:

ingest data from external heterogeneous sources, to produce ingested data, wherein the computer executable instructions to ingest further cause the at least one hardware processor to:

fetch data from external databases having application programming interfaces, APIs;

crawl, in parallel, predefined websites for extracting predetermined attributes to produce crawled data;

enrich the fetch data and the crawled data, comprising obtaining additional or missing attributes from other sources; and

coordinate the fetch, the crawl and the enrich processes using a single index, thereby guiding what data needs to be acquired and eliminating duplicate work;

process the ingested data, comprising:

Natural Language Processing (NLP) of the ingested data to generate a pre-processed ingested data, the NLP being in communication with:

machine learning modeling using the at least one hardware processor for grouping related data into topics and identifying respective keywords;

identifying communities and influencers among researchers; and

predicting predetermined use cases based on historic data; and

execute search queries on the pre-processed ingested data and outputs of the machine learning modeling, the identifying, and the predicting, and generate the analytics insights on search results.

17. The server computer of claim 16 , wherein the computer executable instructions further cause the at least one hardware processor to store the ingested data in a storage component, the storage component acting as a communication data bus for the ingestion component, the processing component and the search and analytics component.

18. The server computer of claim 16 , wherein the computer executable instructions further cause the at least one hardware processor to cross-reference the ingested data, which comprises linking related information from the ingested data into a single record per entity.

19. A communication network, comprising:

at least one server computer for academic search and analytics insights, comprising:

a processor;

a non-transitory computer readable storage medium having computer executable instructions stored thereon for execution by the processor, causing the processor to:

ingest data from external heterogeneous sources, to produce ingested data, wherein the computer executable instructions to ingest further cause the at least one hardware processor to:

fetch data from external databases having application programming interfaces, APIs;

crawl, in parallel, predefined websites for extracting predetermined attributes to produce crawled data;

enrich the fetch data and the crawled data, comprising obtaining additional or missing attributes from other sources; and

coordinate the fetch, the crawl and the enrich processes using a single index, thereby guiding what data needs to be acquired and eliminating duplicate work;

process the ingested data, comprising:

Natural Language Processing (NLP) of the ingested data to generate a pre-processed ingested data, the NLP being in communication with:

machine learning modeling using the processor for grouping related data into topics and identifying respective keywords;

identifying communities and influencers among researchers; and

predicting predetermined use cases based on historic data;

and

execute search queries on the pre-processed ingested data and outputs of the machine learning modeling, the identifying, and the predicting, and generate the analytics insights on search results;

the server computer performing the academic search and analytics insights in response to a request from a client device.

20. The communication network of claim 19 , wherein the computer executable instructions further cause the at least one hardware processor to store the ingested data in a storage component, the storage component acting as a communication data bus for the ingestion component, the processing component and the search and analytics component.

Continuity (2)
Provisional Application 62957565 · Jan 6, 2020
Related Publication 20210209177A1 · Jul 8, 2021
Cited By (2)
US 12,197,518 US 12,613,789