IP Library Granted Patent US 11,269,943
Granted Patent B2
US 11,269,943 · App. 16/593,309 · Granted Mar 8, 2022

Semantic matching system and method

Inventors: Stefan Winzenried (Kilchberg, CH); Adrian Hossu (St. Gallen, CH)
Assignee: JANZZ LTD
G06F16/367G06F16/355
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,269,943
App. No.
16/593,309
Granted
Mar 8, 2022
Kind
B2
Abstract

A computer-based system and method for determining similarity between at least two heterogenous unstructured data records and for optimizing processing performance. A plurality of occupational data records is generated and, for each of the occupational data records, a respective vector is created to represent the occupational data record. Each of the vectors is sliced into a plurality of chunks. Thereafter, semantic matching of the chunks occurs in parallel, to compare at least one occupational data record to at least one other occupational data record simultaneously and substantially in real time. Thereafter, values representing similarities between at least two of the occupational data records are output.

Claims (28)

1. A computer-based method for determining similarity between at least two heterogenous unstructured data records and for optimizing processing performance, the method comprising:

generating, by at least one processor that is configured by executing code stored on non-transitory processor readable media, a plurality of occupational data records;

creating, by the at least one processor, for each of the occupational data records, a respective vector in an n-dimensional non-orthogonal unit vector space by calculating dot products between unit vectors corresponding to concepts from an ontology to represent the occupational data record;

slicing, by the at least one processor, each of the vectors into a plurality of chunks;

performing, by the at least one processor, semantic matching for each of the chunks in parallel to compare at least one occupational data record to at least one other occupational data record; and

outputting, by the at least one processor, values representing similarities between at least two of the occupational data records.

2. The method of claim 1 , wherein each of the respective vectors has magnitude and direction.

3. The method of claim 1 , further comprising applying correlation coefficients derived from information provided by an ontology.

4. The method of claim 1 , further comprising weighting vectorially represented concepts.

5. The method of claim 1 , further comprising storing information associated with dot products that are above zero or at least equal to a predefined threshold.

6. The method of claim 1 , wherein the matching step includes performing asymmetric comparisons.

7. The method of claim 6 , wherein the asymmetric comparisons are based on cosine similarity.

8. The method of claim 1 , wherein the output is sorted based on degree of similarity.

9. A computer-based system for determining similarity between at least two heterogenous unstructured data records and for optimizing processing performance, the system comprising:

at least one processor configured to access non-transitory processor readable media, the at least one processor further configured, when executing instructions stored on the non-transitory processor readable media, to:

generate a plurality of occupational data records;

create, for each of the occupational data records, a respective vector in an n-dimensional non-orthogonal unit vector space by calculating dot products between unit vectors corresponding to concepts from an ontology to represent the occupational data record;

slice each of the vectors into a plurality of chunks;

perform semantic matching for each of the chunks in parallel to compare at least one occupational data record to at least one other occupational data record; and

output values representing similarities between at least two of the occupational data records.

10. The system of claim 9 , wherein each of the respective vectors has magnitude and direction.

11. The system of claim 10 , wherein the at least one processor is further configured to:

apply correlation coefficients derived from information provided by an ontology.

12. The system of claim 10 , wherein the at least one processor is further configured to:

weight vectorially represented concepts.

13. The system of claim 9 , wherein the at least one processor is further configured to:

store information associated with dot products that are above zero or at least equal to a predefined threshold.

14. The system of claim 9 , wherein the matching step includes performing asymmetric comparisons.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2019
From: WINZENRIED, STEFAN; HOSSU, ADRIAN
To: JANZZ LTD
Reel/Frame 050756/0589 →
Continuity (2)
Continuation In Part 16045902 · Jul 26, 2018
Related Publication 20200104315A1 · Apr 2, 2020