IP Library Granted Patent US 11,625,423
Granted Patent B2
US 11,625,423 · App. 17/157,874 · Granted Apr 11, 2023

Linear late-fusion semantic structural retrieval

Inventors: Sean Moran (Putney, GB); Fanny Silavong (London, GB); Rob Otter (Witham, GB); Brett Sanford (Tunbridge Wells, GB); Antonios Georgiadis (London, GB); Sae Young Moon (Edmonton, CA)
Assignee: JPMORGAN CHASE BANK, N.A.
G06F16/334G06F16/322
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,625,423
App. No.
17/157,874
Granted
Apr 11, 2023
Kind
B2
Abstract

Systems and methods for generating a fusion score between electronic documents. The method includes receiving a first electronic document by a document management system. The method further includes extracting a first set of features from the first electronic document including at least one feature type indicating the hierarchical structure of the first electronic document. The method also includes receiving a second electronic document by the document management server. The method further includes extracting a second set of features from the second electronic document including at least one feature type indicating the hierarchical structure of the second electronic document. The method further includes generating a fusion score based on a comparison of the first set of features and the second set of features.

Claims (26)

1. A method comprising:

receiving a first electronic document by a document management server comprising a computer processor;

extracting a first set of features from the first electronic document, the first set of features comprising a hierarchical structure of the first electronic document and a content of the first electronic document;

receiving a second electronic document by the document management server;

extracting, by the document management server, a second set of features from the second electronic document, the second set of features comprising a hierarchical structure of the second electronic document and a content of the second electronic document;

generating, by the document management server, a similarity score for each feature by comparing each feature in the first set of features to a corresponding feature in the second set of features;

generating a weighted feature, by a machine learning model, for each feature of the first electronic document and the second electronic document, the weighted feature based on feedback from a previous similarity score; and

generating, by the document management server, a fusion score comprising a weighted combination of the similarity scores, wherein the weighted combination is based on the weighted feature, and the fusion score represents a similarity between the first electronic document and the second electronic document.

2. The method of claim 1 , wherein the step of extracting the first set of features comprises extracting a tree path structure from the first electronic document, wherein the tree path structure comprises a plurality of nodes.

3. The method of claim 2 , wherein the tree path structure further comprises semantic data for each node of the plurality of nodes, wherein the semantic data identifies a position of each node within the tree path structure, and wherein the semantic data includes structural and contextual information corresponding to each node.

4. The method of claim 1 , wherein generating the similarity score for each feature further comprises computing a positive pointwise mutual information between each feature and the second feature.

5. The method of claim 1 , wherein the similarity score for each feature is further generated using a term frequency-inverse document frequency between each feature and the second feature.

6. A system for generating a fusion score for electronic documents, the system comprising:

a document management system comprising at least one memory storing executable instructions and a processor configured to execute instructions, which cause the processor to perform operations comprising:

receive a first electronic document;

extract a first set of features from the first electronic document, the first set of features comprising a hierarchical structure of the first electronic document and a content of the first electronic document;

receive a second electronic document;

extract a second set of features from the second electronic document, the second set of features comprising a hierarchical structure of the second electronic document and a content of the second electronic document;

generate a similarity score for each feature by comparing each feature in the first set of features to a corresponding feature in the second set of features;

generating a weighted feature, by a machine learning model, for each feature of the first electronic document and the second electronic document, the weighted feature based on feedback from a previous similarity score; and

generating a fusion score comprising a weighted combination of the similarity scores, wherein the weighted combination is based on the weighted feature, and the fusion score represents a similarity between the first electronic document and the second electronic document.

7. The system of claim 6 , wherein the operation of extracting the first set of features comprises extracting a tree path structure from the first electronic document, wherein the tree path structure comprises a plurality of nodes.

8. The system of claim 7 , wherein the tree path structure further comprises semantic data for each node of the plurality of nodes, wherein the semantic data identifies a position of each node within the tree path structure.

9. The system of claim 6 , wherein the operation of generating the similarity score for each feature comprises computing a positive pointwise mutual information between each feature and the second feature.

10. The system of claim 6 , wherein the similarity score for each feature is further generated using a term frequency-inverse document frequency between each feature and the second feature.

11. The system of claim 6 further comprising a client device communicatively coupled to the document management system and operable to communicate the first electronic document or the second electronic document to the document management system.

Continuity (1)
Related Publication 20220237182A1 · Jul 28, 2022