IP Library Granted Patent US 9,928,284
Granted Patent B2
US 9,928,284 · App. 14/588,007 · Granted Mar 27, 2018

File recognition system and method

Inventor: Funmi Fapohunda (San Francisco, CA)
Assignee: ZEPHYR HEALTH, INC.
G06F17/30563
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,928,284
App. No.
14/588,007
Granted
Mar 27, 2018
Kind
B2
Abstract

In various embodiments, a system and method for recognizing and integrating datasets disclosed. The system comprises a processor and a memory unit. A training dataset is developed that forms the basis to measure similarities between and among the datasets. Incoming datasets are examined for features and a measurement is made to determine similarity with datasets in the training dataset. An estimate can be made of the probability that an incoming dataset contains specified attributes. The incoming dataset can be assigned an attribute based on the probability estimate.

Claims (34)

1. A system for recognizing and integrating datasets, the system comprising:

at least one processor and operatively associated memory, wherein the memory stores instructions executable by the at least one processor to:

develop a training dataset comprising at least one first dataset and at least one second dataset;

store, for each dataset, at least one attribute tag corresponding to an attribute in the dataset;

determine a similarity measure between the at least one first dataset and the at least one second dataset based on computing whether the at least one attribute tag is present in both datasets; and

identify a numerator comprising a first sum of inverse document frequency values of attributes in a set-intersection of the at least one first dataset and the at least one second dataset, and a denominator comprising a second sum of inverse document frequency values of attributes in a set-union of the at least one first dataset and the at least one second dataset, and further wherein the similarity measure is a value obtained by dividing the numerator by the denominator.

2. The system of claim 1 , wherein the instructions are further executable by the at least one processor to identify the set-intersection of attribute tags in the at least one first dataset and the at least one second dataset.

3. The system of claim 2 , wherein the instructions are further executable by the at least one processor to compute an inverse document frequency value relating to attribute tags in the set-intersection of the at least one first dataset and the at least one second dataset.

4. The system of claim 1 , wherein the instructions are further executable by the at least one processor to identify the set-union of attribute tags in the at least one first dataset and the at least one second dataset.

5. The system of claim 4 wherein the instructions are further executable by the at least one processor to compute an inverse document frequency value relating to attribute tags in the set-union of the at least one first dataset and the at least one second dataset.

6. A system for recognizing and integrating datasets, the system comprising:

at least one processor and operatively associated memory, wherein the memory stores instructions executable by the at least one processor to:

store a training dataset comprising at least one first dataset and at least one second dataset, wherein a similarity of the at least one first dataset and the at least one second dataset is measured based on the presence or absence, in the at least one first dataset and at least one second dataset, of specified attribute tags;

estimate the similarity of an incoming dataset to a dataset in the training dataset;

identify a group of k nearest neighbors, in the training dataset, to the incoming dataset based on the similarity estimate, wherein k is a quantity of datasets in the training dataset;

identify at least one candidate attribute from the k nearest neighbors in the training dataset; and

determine at least one probability measure, wherein the at least one probability measure quantifies a probability that the at least one candidate attribute is present in the incoming dataset, wherein determining the at least one probability measure comprises identifying a numerator comprising a first sum of similarity measures, identifying a denominator comprising a second sum of similarity measures, and dividing the numerator by the denominator.

7. The system of claim 6 , wherein the instructions are further executable by the at least one processor to predict whether an attribute will be present in the incoming dataset based on the at least one probability measure.

8. The system of claim 7 , wherein the instructions are further executable by the at least one processor to rank candidate attributes in order of estimated probability of presence in the incoming dataset based on the at least one probability measure.

9. A system for recognizing and integrating datasets, the system comprising:

at least one processor and operatively associated memory, wherein the memory stores instructions executable by the at least one processor to:

store a training dataset comprising at least one first dataset and at least one second dataset, wherein a similarity of the at least one first dataset and at least one second dataset is measured based on the presence or absence, in the at least one first dataset and at least one second dataset, of at least one attribute tag related to at least one corresponding attribute;

determine a group of k nearest neighbors within the training dataset, the determination based on a similarity measure between an incoming dataset and at least one dataset in the training dataset, wherein k is a quantity datasets in the training dataset;

determine a probability measure, based on the presence or absence of at least one attribute in at least one dataset within the group of k nearest neighbors, the probability measure quantifying a probability that an attribute is present in the incoming dataset; and

identify a numerator comprising a first sum of similarity measures, and a denominator comprised of a second sum of similarity measures, and, by dividing the numerator by the denominator, obtain the probability measure.

10. The system of claim 9 , wherein the instructions are further executable by the at least one processor to square each similarity measure prior to summation.

11. The system of claim 10 , wherein the instructions are further executable by the at least one processor to enter in the numerator only those similarity measures where a specified attribute is shared by the incoming dataset on the one hand and a dataset in the group of k datasets on the other.

12. The system of claim 10 , wherein the instructions are further executable by the at least one processor to enter in the denominator all similarity measures between the incoming dataset on the one hand and each dataset of the group of k datasets on the other.

13. The system of claim 10 , wherein the probability measure is an input into a determination whether to assign an attribute tag to at least one attribute in the incoming dataset.

14. The system of claim 10 , wherein the instructions are further executable by the at least one processor to determine whether the probability measure exceeds a predetermined threshold.

15. The system of claim 10 , wherein the instructions are further executable by the at least one processor to return a value that can be resolved to a Boolean value.

16. The system of claim 10 , wherein the instructions are further executable by the at least one processor to make a recommendation which attribute tag to assign.

17. The system of claim 10 , wherein the instructions are further executable by the at least one processor to solicit user input which attribute tag to assign.

18. The system of claim 17 , wherein the similarity measure so derived persists in the system so as to form a potential input to a further similarity computation.

Assignments (8)
RELEASE OF SECURITY INTEREST Recorded Aug 15, 2024
From: BARINGS FINANCE LLC, AS ADMINISTRATIVE AGENT
To: ANJU ZEPHYR HEALTH, LLC
Reel/Frame 068301/0345 →
SECURITY INTEREST Recorded Feb 26, 2019
From: ANJU ZEPHYR HEALTH, LLC
To: BARINGS FINANCE LLC, AS ADMINISTRATIVE AGENT
Reel/Frame 048447/0069 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2018
From: ZEPHYR HEALTH INC.
To: ANJU ZEPHYR HEALTH, LLC
Reel/Frame 047420/0790 →
RELEASE OF SECURITY INTEREST Recorded Oct 11, 2018
From: SILICON VALLEY BANK
To: ZEPHYR HEALTH INC.
Reel/Frame 047139/0452 →
RELEASE OF SECURITY INTEREST Recorded Oct 3, 2018
From: TRIPLEPOINT CAPITAL LLC
To: ZEPHYR HEALTH INC.
Reel/Frame 047060/0532 →
SECURITY INTEREST Recorded Mar 26, 2018
From: ZEPHYR HEALTH INC.
To: TRIPLEPOINT CAPITAL LLC
Reel/Frame 045347/0872 →
SECURITY INTEREST Recorded Mar 26, 2018
From: ZEPHYR HEALTH INC.
To: SILICON VALLEY BANK
Reel/Frame 045351/0838 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2015
From: FAPOHUNDA, FUNMI
To: ZEPHYR HEALTH, INC.
Reel/Frame 035175/0496 →
Continuity (1)
Related Publication 20160188701A1 · Jun 30, 2016