IP Library Granted Patent US 10,937,548
Granted Patent B2
US 10,937,548 · App. 15/333,688 · Granted Mar 2, 2021

Computerized system for efficient augmentation of data sets

Inventors: Won Hwa Kim (Madison, WI); Seong Jae Hwang (Madison, WI); Nagesh Adluru (Madison, WI); Sterling Johnson (Fitchburg, WI); Vikas Singh (Madison, WI)
Assignee: Wisconsin Alumni Research Foundation
G16H50/30G16H10/20G16H10/60G16H15/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,937,548
App. No.
15/333,688
Granted
Mar 2, 2021
Kind
B2
Abstract

A method of improving data sets, for example, of patients, each being characterized by relatively low-cost medical data, identifies those patients where the acquisition of higher cost medical data would best inform an estimate of the higher cost medical data for the remaining patients. In this way scarce medical resources can be more efficiently applied in characterizing a potential patient pool, for example, for a clinical trial when resources are not available for extensive medical characterization of each trial participant.

Claims (173)

1. A computerized system for selectively augmenting a data set of data describing objects of a group of related objects each object characterized by a first type of data, the computerized system comprising at least one electronic computer having a memory for holding a stored program and executing to:

(a) use the first type of data of the objects to generate a graph of the objects providing multiple nodes each representing a different object;

(b) use a wavelet expansion operating on the graph to identify a limited set of proxy objects from among the group and representative of the group with respect to a second type of data of the objects but being less than the number of objects in the group;

(c) based on the identification of the proxy objects, create an augmented data set by collecting the second type of data different from the first type of data for the proxy objects and not for the objects other than the proxy objects; and

(d) using the augmented data set of the proxy objects and information of the graph to produce estimated data estimating the second type of data for objects, represented in the group of the augmented data set, other than the proxy objects, where the second type of data has not been collected;

wherein the wavelet expansion is in accordance with the equation:

p

s

(

n

)

=

1

Z

s

ψ

n

(

s

,

n

)

=

1

Z

s

l

=

0

N

-

1

h

(

s

λ

l

)

χ

l

(

n

)

2

(

1

)

where:

n is a node index;

ψ n (s,n) is a mother wavelet function having a scale s and translation values localized at each node index n;

h( ) is a filter for wavelets;

λ l and χ l are pairs of eigenvalues and corresponding eigenvectors of a graph Laplacian L operator; and

Z s is a normalizing factor

Z

s

=

n

=

1

N

ψ

n

(

s

,

n

)

computed over a selected wavelet.

2. The computerized system of claim 1 wherein the estimation employs minimization of an estimation error in a frequency domain of the graph subject to a band limitation and a subsequent inverse transformation from the frequency domain back into the estimated data of the graph.

3. The computerized system of claim 1 wherein the estimated data is used to characterize the objects according to predetermined criterion.

4. The computerized system of claim 3 wherein the objects are patients for a clinical trial and the first type of data has a first cost and the second type of data has a second cost higher than the first cost and wherein the estimated data can be used for downstream analyses of the clinical study.

5. The computerized system of claim 4 wherein the first type of data and second type of data represent medical measurements related to Alzheimer risk and the clinical trial relates to analyses of the second type of data the analyses being at least one of categorical analyses and statistical analyses.

6. The computerized system of claim 1 wherein the graph is in non-Euclidean spaces and has nodes representing each object and edges based on a similarity of data elements of the first data set of the objects.

7. A method of selectively augmenting a data set providing related objects each characterized by a first type of data, the method operating on at least one electronic computer having a memory for holding a stored program and executing to perform the steps comprising:

(a) using the first type of data of the objects to generate a graph of the objects;

(b) using a wavelet expansion to identify proxy objects of the graph to identify a limited set of proxy objects from among the group and representative of the group with respect to a second type of data of the objects but being less than the number of objects in the group;

(c) based on the identification of the proxy objects, creating an augmented data set by collecting the second type of data different from the first type of data for the proxy objects; and

(d) using the augmented data set of the proxy objects and the graph to produce estimated data estimating the second type of data for objects of the group of the augmented data set other than the proxy objects where the second type of data has not been collected;

wherein the wavelet expansion is in accordance with the equation:

p

s

(

n

)

=

1

Z

s

ψ

n

(

s

,

n

)

=

1

Z

s

l

=

0

N

-

1

h

(

s

λ

l

)

χ

l

(

n

)

2

(

1

)

where:

n is a node index;

ψ n (s,n) is a mother wavelet function having a scale s and translation values localized at each node index n;

h( ) is a filter for wavelets;

λ l and χ l are pairs of eigenvalues and corresponding eigenvectors of a graph Laplacian L operator, and

Z s is a normalizing factor

Z

s

=

n

=

1

N

ψ

n

(

s

,

n

)

computed over a selected wavelet.

8. The method of claim 7 wherein the estimation employs minimization of an estimation error in a frequency domain of the graph subject to a band limitation and a subsequent inverse transformation from the frequency domain back into the graph.

9. The method of claim 7 wherein the objects are patients for a clinical trial and the first type of data has a first cost and the second type of data has a second cost higher than the first cost and wherein the estimated data can be used for downstream analyses of the clinical study.

10. The method of claim 7 wherein the first type of data and second type of data represent medical measurements related to Alzheimer's risk and the clinical trial relates to analyses of the second type of data, the analysis providing at least one of a categorical analyses and statistical analyses identifying the effects from Alzheimer's disease.

11. The method of claim 9 wherein the graph is in non-Euclidean spaces and has nodes representing each object and edges based on a similarity of data elements of the first data set of the objects.

Assignments (2)
CONFIRMATORY LICENSE Recorded Jul 3, 2018
From: UNIVERSITY OF WISCONSIN MADISON
To: NATIONAL INSTITUTES OF HEALTH (NIH), U.S. DEPT. OF HEALTH AND HUMAN SERVICES (DHHS), U.S. GOVERNMENT
Reel/Frame 046473/0837 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2017
From: HWANG, SEONG JAE; ADLURU, NAGESH; KIM, WON HWA; JOHNSON, STERLING; SINGH, VIKAS
To: WISCONSIN ALUMNI RESEARCH FOUNDATION
Reel/Frame 041562/0175 →
Continuity (1)
Related Publication 20180113990A1 · Apr 26, 2018