IP Library Granted Patent US 9,158,839
Granted Patent B2
US 9,158,839 · App. 13/592,811 · Granted Oct 13, 2015

Systems and methods for training and classifying data

Inventor: Alexander Karl Hudek (Toronto, CA)
Assignee: DiligenceEngine, Inc.
G06F17/30707G06N5/02G06N99/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,158,839
App. No.
13/592,811
Granted
Oct 13, 2015
Kind
B2
Abstract

A mechanism for training and classifying data is disclosed. The method includes receiving a data set having at least a first annotation and at least a second annotation. The first annotation and the second annotation represent characteristics within the data set. The method also includes determining a first identifier from the first annotation and a second identifier from the second annotation and associating the first identifier to the second identifier to generate a joined identifier. The method also includes computing feature weights and transition weights for the annotated data set based on the at least a first identifier, at least a second identifier, and at least a joined identifier and transitions between each of the first, the second and the joined identifiers. The method further includes receiving a second un-annotated data set and classifying the second data set based on the computed feature weights and the transition weights.

Claims (40)

1. A method comprising:

receiving, by a processing device, a data set, wherein the data set is annotated with at least a first annotation and at least a second annotation, wherein at least the first annotation and the second annotation represent characteristics within the data set;

determining, by the processing device, a first identifier from the first annotation and a second identifier from the second annotation;

associating, by the processing device, the first identifier to the second identifier to generate at least one joined identifier, wherein the at least one joined identifier is placed between the first annotation and the second annotation;

computing, by the processing device, a feature weight and a transition weight for the annotated data set based on the first annotation, the second annotation, and the at least one joined identifier and a transition between each of the first annotation, the second annotation and the at least one joined identifier;

receiving, by the processing device, a second data set, wherein the second data set is un-annotated; and

classifying, by the processing device, the second data set based on the computed feature weight and the transition weight from the first data set.

2. The method of claim 1 further comprising providing the classified second data set to a user.

3. The method of claim 1 wherein the data set and the second data set comprise a plurality of sentences having a plurality of words.

4. The method of claim 3 wherein each of the first annotation and the second annotation comprise at least one word.

5. The method of claim 4 wherein the feature weight characterizes the words that occur in each of the first identifier, second identifier and the joined identifier.

6. The method of claim 3 wherein the classifying comprising assigning at least one class label to each of the sentences in the second data set.

7. The method of claim 1 wherein the first annotation is adjoining to the second annotation in the data set.

8. The method of claim 1 wherein the data set is one of a text, image, video and audio and the second data set is one of the text, image, video and audio.

9. The method of claim 1 wherein the transition weight characterizes the transition that occurs between each of the first identifier, the second identifier and the at least one joined identifier.

10. A system comprising:

a memory;

a processing device coupled to the memory, the processing device configured to:

receive a data set, wherein the data set is annotated with at least a first annotation and at least a second annotation, wherein at least the first annotation and the second annotation represent within the data set;

determine a first identifier from the first annotation and a second identifier from the second annotation;

associate the first identifier to the second identifier to generate at least one joined identifier, wherein the at least one joined identifier is placed between the first annotation and the second annotation;

compute a feature weight and a transition weight for the annotated data set based on the first annotation, the second annotation, and the at least one joined identifier and a transition between each of the first annotation, the second annotation and the at least one joined identifier;

receive a second data set, wherein the second data set is un-annotated; and

classify the second data set based on the computed feature weight and the transition weight from the first data set.

11. The system of claim 10 wherein the processing device is further configured to provide the classified second data set to a user.

12. The system of claim 10 wherein the data set and the second data set comprise a plurality of sentences having a plurality of words.

13. The system of claim 12 wherein each of the first annotation and the second annotation comprise at least one word.

14. The system of claim 10 wherein the first annotation is adjoining to the second annotation in the data set.

15. The system of claim 12 wherein the classify comprise assign at least one class label to each of the sentences in the second data set.

16. A non-transitory machine-readable storage medium including data that, when accessed by a machine, cause the machine to perform operations comprising:

receiving, by a processing device, a data set, wherein the data set is annotated with at least a first annotation and at least a second annotation, wherein at least the first annotation and the second annotation represent characteristics within the data set;

determining, by the processing device, a first identifier from the first annotation and a second identifier from the second annotation;

associating, by the processing device, the first identifier to the second identifier to generate at least one joined identifier, wherein the at least one joined identifier is placed between the first annotation and the second annotation;

computing, by the processing device, a feature weight and a transition weight for the annotated data set based on the first annotation, the second annotation, and the at least one joined identifier and a transition between each of the first annotation, the second annotation and the at least one joined identifier;

receiving, by the processing device, a second data set, wherein the second data set is un-annotated; and

classifying, by the processing device, the second data set based on the computed feature weight and the transition weight from the first data set.

17. The non-transitory machine-readable storage medium of claim 16 , the operations further comprising providing the classified second data set to a user.

18. The non-transitory machine-readable storage medium of claim 16 wherein the data set and the second data set comprise a plurality of sentences having a plurality of words.

19. The non-transitory machine-readable storage medium of claim 18 wherein each of the first annotation and the second annotation comprise at least one word.

20. The non-transitory machine-readable storage medium of claim 18 wherein the classifying comprising assigning at least one class label to each of the sentences in the second data set.

Assignments (7)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE ASSIGNEE ADDING THE SECOND ASSIGNEE PREVIOUSLY RECORDED AT REEL: 058859 FRAME: 0104. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 18, 2022
From: KIRA INC.
To: KIRA INC.; ZUVA INC.
Reel/Frame 061964/0502 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNMENT OF ALL OF ASSIGNOR'S INTEREST PREVIOUSLY RECORDED AT REEL: 057509 FRAME: 0057. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jan 26, 2022
From: KIRA INC.
To: ZUVA INC.
Reel/Frame 058859/0104 →
SECURITY INTEREST Recorded Oct 29, 2021
From: KIRA INC.
To: OWL ROCK CAPITAL CORPORATION, AS COLLATERAL AGENT
Reel/Frame 057964/0784 →
SECURITY INTEREST Recorded Sep 16, 2021
From: ZUVA INC.
To: KIRA INC.
Reel/Frame 057509/0067 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2021
From: KIRA INC.
To: ZUVA INC.
Reel/Frame 057509/0057 →
CHANGE OF NAME Recorded Mar 28, 2019
From: DILIGENCEENGINE, INC
To: KIRA INC
Reel/Frame 049401/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2012
From: HUDEK, ALEXANDER KARL
To: DILIGENCEENGINE,INC.
Reel/Frame 028837/0552 →
Continuity (1)
Related Publication 20140058983A1 · Feb 27, 2014