IP Library Granted Patent US 7,788,193
Granted Patent B2
US 7,788,193 · App. 11/929,354 · Granted Aug 31, 2010

Kernels and methods for selecting kernels for use in learning machines

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,788,193
App. No.
11/929,354
Granted
Aug 31, 2010
Kind
B2
Abstract

Learning machines, such as support vector machines, are used to analyze datasets to recognize patterns within the dataset using kernels that are selected according to the nature of the data to be analyzed. Where the datasets possesses structural characteristics, locational kernels can be utilized to provide measures of similarity among data points within the dataset. The locational kernels are then combined to generate a decision function, or kernel, that can be used to analyze the dataset. Where an invariance transformation or noise is present, tangent vectors are defined to identify relationships between the invariance or noise and the data points. A covariance matrix is formed using the tangent vectors, then used in generation of the kernel.

Claims (29)

1. A computer-implemented method for analyzing data comprising a text document to identify patterns in words or characters within the document, the method comprising:

inputting the data into a computing environment comprising one or more pre-processing program modules and one or more support vector machine modules stored on a drive or a system memory of a computer or computer network by:

dividing the data into a training dataset and a test dataset;

defining a kernel for structured data for execution by the one or more support vector machine modules by representing the training dataset as a collection of sequences of words or characters and an index set within the document structure, wherein the indices within the index set correspond to locations of words or characters within the document;

applying a vicinity function to the collection of sequences of words or characters to define a plurality of sequences of words or characters centered at different words or characters;

measuring similarity of pairs of sequences of words or characters centered at the different indices to define a locational kernel having a value corresponding to each of the different pairs of sequences of words or characters;

creating additional locational kernels by performing an operation selected from addition, scalar multiplication, multiplication, pointwise limits, transformation and convolution on the locational kernels;

combining the locational kernels and the additional locational kernels for the different sequences of words or characters by performing an operation to produce a kernel on a set of sequences of words or characters;

testing the kernel on the test data set having a known set of sequences of words or characters to determine whether an optimal solution has been achieved;

if the optimal solution has been achieved, applying the kernel on a set of sequences of words or characters to a document having an unknown structure to identify patterns within, and thereby extract knowledge from, the document; and

generating a display of the identified patterns of words or characters within the document having an unknown structure.

2. The method of claim 1 , wherein the step of combining the locational kernels comprises performing an operation selected from the group consisting of summing over the indices and fixing the indices on a given locational kernel.

3. The method of claim 1 , wherein the locational kernels comprise Gaussian radial basis function kernels having fixed indices and the kernel on a set of sequences of words comprises a product over all index positions of the locational kernels.

4. The method of claim 1 , wherein the kernel on a set of sequences of words or characters further incorporates transformation invariance and the method is generated by estimating at least one tangent vector for associating each data point with a local invariance and selecting the function so that its optimization incorporates a covariance matrix of a plurality of the tangent vectors.

5. The method of claim 1 , wherein the dataset has the characteristic of noise and the function is generated by estimating at least one noise vector associated with each data point and selecting the function so that its optimization incorporates a covariance matrix of a plurality of noise vectors.

6. The method of claim 1 , wherein the words or characters can appear in any order within the sequence and a penalty is applied according to a distance of each word or character from the index.

7. A computer-implemented method for analyzing data comprising a text document to identify patterns in words or characters within the document, the method comprising:

inputting the data into a memory and a computer processor programmed for executing one or more support vector machines;

dividing the data into a training dataset and a test dataset;

defining a kernel for structured data for execution by the one or more support vector machines by representing the training dataset as a collection of word or character strings and an index set within the document structure, wherein the indices within the index set correspond to locations of word or character strings within the document;

applying a vicinity function to the collection of word or character strings to define a plurality of word or character strings centered at different words or characters;

measuring similarity of pairs of word or character strings centered at the different indices to define a locational kernel having a value corresponding to each of the different pairs of word or character strings;

creating additional locational kernels by performing an operation selected from addition, scalar multiplication, multiplication, pointwise limits, transformation and convolution on the locational kernels;

combining the locational kernels and the additional locational kernels for the different word or character strings by performing an operation to produce a kernel on a set of word or character strings;

testing the kernel on the test data set having a known set of word or character strings to determine whether an optimal solution has been achieved;

if the optimal solution has been achieved, applying the kernel on a set of word or character strings to a document having an unknown structure to identify patterns within, and thereby extract knowledge from, the document; and

generating a display of the identified patterns of words or characters within the document having an unknown structure.

8. The method of claim 7 , wherein the step of combining the locational kernels comprises performing an operation selected from the group consisting of summing over the indices and fixing the indices on a given locational kernel.

9. The method of claim 7 , wherein the words or characters can appear in any order within the sequence and a penalty is applied according to a distance of each word or character from the index.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2008
From: MEMORIAL HEALTH TRUST, INC.; STERN, JULIAN N.; ROBERTS, JAMES; PADEREWSKI, JULES B.; FARLEY, PETER J.; ANDERSON, CURTIS; MATTHEWS, JOHN E.; SIMPSON, K. RUSSELL; O'HAYER, TIMOTHY P.; BERGERON, GLYNN; CARLS, GARRY L.; MCKENZIE, JOE
To: HEALTH DISCOVERY CORPORATION
Reel/Frame 020361/0542 →
CONSENT ORDER CONFIRMING FORECLOSURE SALE ON JUNE 1, 2004 Recorded Jan 11, 2008
From: BIOWULF TECHNOLOGIES, LLC
To: MEMORIAL HEALTH TRUST, INC.; STERN, JULIAN N.; ROBERTS, JAMES; PADEREWSKI, JULES B.; FARLEY, PETER J.; ANDERSON, CURTIS; MATTHEWS, JOHN E.; SIMPSON, K. RUSSELL; O'HAYER, TIMOTHY P.; BERGERON, GLYNN; CARLS, GARRY L.; MCKENZIE, JOE
Reel/Frame 020352/0896 →
NUNC PRO TUNC ASSIGNMENT Recorded Jan 1, 2008
From: BARTLETT, PETER; ELISSEEFF, ANDRE; SCHOELKOPF, BERNHARD; CHAPELLE, OLIVIER
To: BIOWULF TECHNOLOGIES, LLC
Reel/Frame 020305/0244 →