IP Library › Granted Patent US 11,170,179
Granted Patent B2
US 11,170,179 · App. 16/021,112 · Granted Nov 9, 2021

Systems and methods for natural language processing of structured documents

Inventors: Jason D. Mills (Montclair, NJ); Zheng Wu (New York, NY); Jennifer Rabowsky (New York, NY); Kenny Song (Weehawken, NJ); Khaled Bugrara (Framingham, MA)
Assignee: JPMORGAN CHASE BANK, N.A.
G06F40/40G06F40/205G06F40/30G06K9/00469G06F40/106
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,170,179
App. No.
16/021,112
Granted
Nov 9, 2021
Kind
B2
Abstract

Systems and methods for natural language processing of structured documents. In another embodiment, in an information processing apparatus comprising at least one computer processor, a method for processing a structured document may include: (1) receiving a document; (2) parsing the document into a plurality of components using a statistical parser; (3) extracting a plurality of entities from each component; (4) identifying a potential relationship between two of the plurality of entities; (5) generating a numeric representation for the potential relationship; (6) confirming the potential relationship with a logical regression model; and (7) generating and storing a unified structured file for the document.

Claims (53)

1. A method for processing a structured document, comprising:

in an information processing apparatus comprising at least one computer processor:

receiving a structured legal document;

parsing the structured legal document into a plurality of document sections using a statistical parser comprising a neural network;

extracting a plurality of entities from each document section;

identifying a potential relationship between two of the plurality of entities;

generating a numeric representation for the potential relationship;

confirming the potential relationship with a logical regression model;

generating and storing a unified structured file for the structured legal document;

receiving feedback on an accuracy of at least one of the plurality of document sections identified using the statistical parser; and

updating the statistical parser based on the feedback.

2. The method of claim 1 , wherein the plurality of document sections comprise an article, a section, a subsection, or a subsubsection.

3. The method of claim 1 , wherein the statistical parser parses the structured legal document based on a first vector of word embeddings and a second vector of orthographic properties of words in the legal document.

4. The method of claim 1 , wherein the step of parsing the structured legal document into a plurality of document sections comprises identifying a relationship among the plurality of document sections.

5. The method of claim 1 , further comprising:

generating a score for each sentence or paragraph of the structured legal document.

6. The method of claim 5 , wherein the score is generated using a latent semantic indexing model.

7. The method of claim 5 , wherein the score is generated using a continuous bag-of-words model.

8. The method of claim 1 , wherein the potential relationship is based on an ontology.

9. The method of claim 1 , wherein the potential relationship is based on a hierarchical correspondence rule.

10. The method of claim 1 , wherein the numeric representation for the potential relationship is based on functional features, tail features, and head features of the potential relationship.

11. The method of claim 1 , further comprising:

identifying a plurality of defined terms in the structured legal document.

12. The method of claim 1 , wherein the logical regression model confirms each potential relationship as being true or false.

13. The method of claim 1 , further comprising:

generating a graphical representation of the structured legal document.

14. A method for processing a structured document, comprising:

in an information processing apparatus comprising at least one computer processor:

receiving a structured legal document;

parsing the structured legal document into a plurality of document sections using a statistical parser comprising a neural network;

extracting a plurality of entities from each document section using a Conditional Random Field model;

identifying a potential relationship between two of the plurality of entities;

generating a numeric representation for the potential relationship;

confirming the potential relationship with a logical regression model;

generating and storing a unified structured file for the structured legal document;

receiving feedback on an accuracy of at least one of the entities extracted using the Conditional Random Field model; and

updating the Conditional Random Field model based on the feedback.

15. The method of claim 14 , wherein the plurality of document sections comprise an article, a section, a subsection, or a subsubsection.

16. The method of claim 14 , wherein the statistical parser parses the structured legal document based on a first vector of word embeddings and a second vector of orthographic properties of words in the legal document.

17. A method for processing a structured document, comprising:

in an information processing apparatus comprising at least one computer processor:

receiving a structured legal document;

parsing the structured legal document into a plurality of document sections using a statistical parser comprising a neural network;

extracting a plurality of entities from each document;

identifying a potential relationship between two of the plurality of entities;

generating a numeric representation for the potential relationship;

confirming the potential relationship with a logical regression model;

generating and storing a unified structured file for the structured legal document;

receiving feedback on an accuracy of at least one of the potential relationships confirmed using the logical regression model; and

updating the logical regression model based on the feedback.

18. The method of claim 17 , wherein the plurality of document sections comprise an article, a section, a subsection, or a subsubsection.

19. The method of claim 17 , wherein the statistical parser parses the structured legal document based on a first vector of word embeddings and a second vector of orthographic properties of words in the legal document.

20. The method of claim 17 , wherein further comprising generating a score for each sentence or paragraph of the structured legal document, wherein the score is generated using a latent semantic indexing model or a continuous bag-of-words model.

Continuity (2)
Provisional Application 62527487 · Jun 30, 2017
Related Publication 20190005029A1 · Jan 3, 2019
Cited By (1)
US 12,597,284