IP Library › Granted Patent US 11,874,798
Granted Patent B2
US 11,874,798 · App. 17/486,554 · Granted Jan 16, 2024

Smart dataset collection system

Inventor: Hans-Martin Ramsl (Mannheim, DE)
Assignee: SAP SE
G06F16/164G06F16/1873G06F16/345G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,874,798
App. No.
17/486,554
Filed
Sep 27, 2021
Granted
Jan 16, 2024
Kind
B2
Art Unit
2154
USPC
707/736
Abstract

Datasets are available from different dataset servers and often lack well-defined metadata. Thus, comparing datasets is difficult. Additionally, there might be different versions of the same dataset which makes the search even more difficult. Using systems and methods described herein, quality scores, dataset versioning, topic identification, and semantic relatedness metadata is stored about datasets stored on dataset servers. A user interface is provided to allow a user to search for datasets by specifying search criteria (e.g., a topic and a minimum quality score) and to be informed of responsive datasets. The user interface may further inform the user of the quality scores of the responsive datasets, the versions of the responsive datasets, or other metadata. From the search results, the user may select and download one or more of the responsive datasets.

Claims (59)

1. A method comprising:

receiving, by one or more processors and via a network, a plurality of datasets from a plurality of sources;

generating, by the one or more processors, based on each dataset of the plurality of datasets, a topic for the dataset;

determining, based on the topic for each dataset of the plurality of datasets, a freshness for the dataset;

determining for each dataset of the plurality of datasets, a quality score for the dataset based on the dataset, the freshness of the dataset, and the source of the dataset;

storing, in association with each dataset of the plurality of datasets, the quality score, the topic, and the freshness for the dataset;

receiving, via a user interface, a search request comprising a search topic and sort criteria; and

in response to the search request, based on the search topic, the topics for the plurality of datasets, and the quality scores for the plurality of datasets, causing a user interface to be presented that identifies a list of datasets corresponding to the search topic, the list ordered according to the sort criteria.

2. The method of claim 1 , wherein:

the generating of the quality score for each dataset of the plurality of datasets comprises determining a number of spelling errors in the dataset.

3. The method of claim 1 , further comprising:

generating a vector representation of a semantic meaning for each dataset of the plurality of datasets;

determining a degree of similarity between a first dataset and a second dataset based on the vector representations for the first dataset and the second dataset; and

based on the determined degree of similarity and a predetermined threshold, linking the first dataset with the second dataset.

4. The method of claim 1 , further comprising:

determining, based on each dataset of the plurality of datasets, a suitability rating of the dataset for each of a plurality of artificial intelligence (AI) applications; and

causing a user interface to be presented that indicates at least a subset of the determined suitability ratings.

5. The method of claim 1 , wherein the generating of the quality score for each dataset comprises determining whether the source for the dataset is trusted.

6. The method of claim 5 , wherein the determining whether the source for the dataset is trusted comprising checking the source against a list of trusted sources.

7. The method of claim 1 , wherein the determining, based on the topic for each dataset of the plurality of datasets, the freshness of the dataset comprises querying the topic against a knowledge base.

8. A system comprising:

a memory that stores instructions; and

one or more processors configured by the instructions to perform operations comprising:

receiving, via a network, a plurality of datasets from a plurality of sources;

generating, by the one or more processors, based on each dataset of the plurality of datasets, a topic for the dataset;

determining, based on the topic for each dataset of the plurality of datasets, a freshness for the dataset;

determining for each dataset of the plurality of datasets, a quality score for the dataset based on the dataset, the freshness of the dataset, and the source of the dataset;

storing, in association with each dataset of the plurality of datasets, the quality score, the topic, and the freshness for the dataset;

receiving, via a user interface, a search request comprising a search topic and sort criteria; and

in response to the search request, based on the search topic, the topics for the plurality of datasets, and the quality scores for the plurality of datasets, causing a user interface to be presented that identifies a list of datasets corresponding to the search topic, the list ordered according to the sort criteria.

9. The system of claim 8 , wherein:

the generating of the quality score for each dataset of the plurality of datasets comprises determining a number of spelling errors in the dataset.

10. The system of claim 8 , wherein the operations further comprise:

generating a vector representation of a semantic meaning for each dataset of the plurality of datasets;

determining a degree of similarity between a first dataset and a second dataset based on the vector representations for the first dataset and the second dataset; and

based on the determined degree of similarity and a predetermined threshold, linking the first dataset with the second dataset.

11. The system of claim 8 , wherein the operations further comprise:

determining, based on each dataset of the plurality of datasets, a suitability rating of the dataset for each of a plurality of artificial intelligence (AI) applications; and

causing a user interface to be presented that indicates at least a subset of the determined suitability ratings.

12. The system of claim 8 , wherein the generating of the quality score for each dataset comprises determining whether the source for the dataset is trusted.

13. The system of claim 12 , wherein the determining whether the source for the dataset is trusted comprises checking the source against a list of trusted sources.

14. The system of claim 8 , wherein the determining, based on the topic for each dataset of the plurality of datasets, the freshness of the dataset comprises querying the topic against a knowledge base.

15. A non-transitory computer-readable medium that stores instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

receiving, via a network, a plurality of datasets from a plurality of sources;

generating, by the one or more processors, based on each dataset of the plurality of datasets, a topic for the dataset;

determining, based on the topic for each dataset of the plurality of datasets, a freshness for the dataset;

determining for each dataset of the plurality of datasets, a quality score for the dataset based on the dataset, the freshness of the dataset, and the source of the dataset;

storing, in association with each dataset of the plurality of datasets, the quality score, the topic, and the freshness for the dataset;

receiving, via a user interface, a search request comprising a search topic and sort criteria; and

in response to the search request, based on the search topic, the topics for the plurality of datasets, and the quality scores for the plurality of datasets, causing a user interface to be presented that identifies a list of datasets corresponding to the search topic, the list ordered according to the sort criteria.

16. The non-transitory computer-readable medium of claim 15 , wherein:

the generating of the quality score for each dataset of the plurality of datasets comprises determining a number of spelling errors in the dataset.

17. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

generating a vector representation of a semantic meaning for each dataset of the plurality of datasets;

determining a degree of similarity between a first dataset and a second dataset based on the vector representations for the first dataset and the second dataset; and

based on the determined degree of similarity and a predetermined threshold, linking the first dataset with the second dataset.

18. The non-transitory computer-readable medium of claim 15 , wherein the generating of the quality score for each dataset comprises determining whether the source for the dataset is trusted.

19. The non-transitory computer-readable medium of claim 18 , wherein the determining whether the source for the dataset is trusted comprises checking the source against a list of trusted sources.

20. The non-transitory computer-readable medium of claim 15 , wherein the determining, based on the topic for each dataset of the plurality of datasets, the freshness of the dataset comprises querying the topic against a knowledge base.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2021
From: RAMSL, HANS-MARTIN
To: SAP SE
Reel/Frame 057751/0780 →
Continuity (1)
Related Publication 20230096118A1 · Mar 30, 2023