IP Library Granted Patent US 12682235
Granted Patent B2
US 12682235 · App. 18/187,474 · Granted Jul 14, 2026

Method and system for identifying hierarchical relationships between data elements of document

Inventors: Bhaskar Kalita (Marlborough, MA); Karthik Kumar Veldandi (Mumbai, IN); Jeevan Prakash (Mumbai, IN); Alok Kumar Garg (Mumbai, IN); Sagar Kewalramani (Toronto, CA)
Assignee: Quantiphi Inc.
G06N3/08G06F16/93G06F40/103G06F40/279
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682235
App. No.
18/187,474
Granted
Jul 14, 2026
Kind
B2
Abstract

Disclosed is a method for identifying multi-level hierarchical relationships between data elements of a document, the method comprising receiving a plurality of sample documents each having a plurality of data elements arranged in a multi-level hierarchical data structure; classifying each of the plurality of data elements into a key entity field or a key field value based on a hierarchical relationship therebetween; identifying key entity fields, from among the classified key entity field of the plurality of data elements, having the hierarchical relationship therebetween; pairing the key entity field, with a corresponding key field value or an identified key entity field, to form a training dataset; and employing the training dataset on a neural network framework, having at least one of a textual modality or a visual modality, to identify the multi-level hierarchical relationships between the data elements of the document.

Claims (42)

1 . A computer-implemented method for identifying multi-level hierarchical relationships between data elements of a document, the method comprising:

receiving, from a data repository, a plurality of sample documents each having a plurality of data elements arranged in a dynamic nested tree representing a multi-level hierarchical data structure;

classifying, by at least one processor, each of the plurality of data elements into a key entity field or a key field value based on a hierarchical relationship between each of the plurality of data elements;

identifying, by the at least one processor, key entity fields, from among the plurality of data elements classified as key entity fields, having the hierarchical relationship therebetween;

pairing, by the at least one processor, key entity fields, with a corresponding key field value or an identified key entity field, to form a training dataset;

defining a parent and child relationship between the paired key entity field and one of the corresponding key field value or the another identified key entity field based on at least one of (i) a visual hierarchical indentation and (ii) a visual precedence, and by offsetting the paired key entity field with one of the corresponding key field value or the another identified key entity field to allow the paired key entity field, the corresponding key field value and the another identified key entity field to attain a fixed position in the sample documents; and

employing the training dataset comprising textual features and visual features including indentation of a given key entity field with respect to a previous key entity field, position of the key field values, font size of the key entity fields and the key field values, font type of the key entity fields and the key field values, and colour of the key entity fields and the key field values, for supervised learning of a neural network framework to identify the multi-level hierarchical relationships between the data elements of the document.

2 . The method according to claim 1 , wherein the multi-level hierarchical data structure is a dynamic nested tree, wherein the dynamic nested tree comprises one or more levels of hierarchy for one or more key entity fields, and wherein each of the one or more key entity fields hierarchically concludes with a corresponding key field value.

3 . The method according to claim 1 , wherein the key entity field is extracted from the document by:

defining a set of primary data elements, and

comparing the plurality of data elements with each of the set of primary data elements to extract the key entity field.

4 . The method according to claim 1 , wherein the hierarchical relationship is a parent and child relationship.

5 . The method according to claim 4 , wherein the parent and child relationship is defined, between the paired key entity field with one of the corresponding key field value or the identified key entity field, based on a visual precedence or a visual hierarchical indentation, respectively.

6 . The method according to claim 3 , wherein the parent and child relationship is defined by offsetting the paired key entity field with one of the corresponding key field value or the identified key entity field, wherein offsetting the paired key entity field with one of the corresponding key field value or the identified key entity field allows the paired key entity field, the corresponding key field value and the identified key entity field to attain a fixed position in the sample documents.

7 . The method according to claim 3 , wherein the hierarchical relationships between the paired key entity field with one of the corresponding key field value or the identified key entity field is represented using at least one of:

a set of key field values,

a table containing the set of key field values, and

a group of key entity fields along with corresponding set of key field values.

8 . The method according to claim 1 , wherein the hierarchical relationships between the paired key entity field with one of the corresponding key field value or the identified key entity field is represented by colour coding the set of key field values, and the group of key entity fields.

9 . The method according to claim 1 , wherein the key entity field comprises one of alphabetical characters or alphanumeric characters, and wherein the key field value comprises alphabetical characters, numerical characters or alphanumeric characters.

10 . The method according to claim 1 , wherein identification of the hierarchical relationships in the document enables in comparing the document with another document, and wherein the document is one of a policy, a contract, an agreement, a certificate or a scheme the plurality of data elements comprising a plurality of key entity fields and key field values associated plurality of key entity fields.

11 . A system for identifying multi-level hierarchical relationships between data elements of a document, the system comprising at least one processor configured to:

receive, from a data repository, a plurality of sample documents each having a plurality of data elements arranged in a dynamic nested tree representing a multi-level hierarchical data structure;

classify each of the plurality of data elements into a key entity field or a key field value based on a hierarchical relationship between each of the plurality of data elements;

identify key entity fields, from among the classified key entity field of the plurality of data elements, having the hierarchical relationship therebetween;

pair key entity fields, with a corresponding key field value or an identified key entity field, based on the hierarchical relationship therebetween to form a training dataset;

define, by the at least one processor, a parent and child relationship between the paired key entity field and one of the corresponding key field value or the another identified key entity field based on at least one of (i) a visual hierarchical indentation and (ii) a visual precedence, and by offsetting the paired key entity field with one of the corresponding key field value or the another identified key entity field to allow the paired key entity field, the corresponding key field value and the another identified key entity field to attain a fixed position in the sample documents; and

employ the training dataset comprising textual features and visual features including indentation of a given key entity field with respect to a previous key entity field, position of the key field values, font size of the key entity fields and the key field values, font type of the key entity fields and the key field values, and colour of the key entity fields and the key field values, for supervised learning of a neural network framework to identify the multi-level hierarchical relationships between the data elements of the document.

12 . The system according to claim 11 , wherein the multi-level hierarchical data structure is a dynamic nested tree, wherein the dynamic nested tree structure comprises one or more levels of hierarchy for one or more key entity fields, and wherein each of the one or more key entity fields hierarchically concludes with a corresponding key field value.

13 . The system according to claim 11 , wherein the at least one processor extracts the key entity field from the document by:

receiving a set of primary data elements, and

comparing the plurality of data elements with each of the set of primary data elements to extract the key entity field.

14 . The system according to claim 11 , wherein the multi-level hierarchical relationship is a parent and child relationship.

15 . The system according to claim 14 , wherein the at least one processor is configured to define the parent and child relationship, between the paired key entity field with one of the corresponding key field value or the identified key entity field, based on a visual hierarchical indentation or a visual precedence, respectively.

16 . The system according to claim 14 , wherein the at least one processor is configured to defined the parent child relationship by offsetting the paired key entity field with one of the corresponding key field value or the identified key entity field, wherein offsetting the paired key entity field with one of the corresponding key field value or the identified key entity field allows the paired key entity field, the corresponding key field value and the identified key entity field to attain a fixed position in the sample documents.

17 . The system according to claim 11 , wherein the at least one processor is configured to represent the hierarchical relationships between the paired key entity field with one of the corresponding key field value or the identified key entity field by:

a set of key field values,

a table containing the set of key field values, and

a group of key entity fields.

18 . A system according to claim 11 , wherein the at least one processor is configured to represent the hierarchical relationships between the paired key entity field with one of the corresponding key field value or the identified key entity field by colour coding the set of key field values, and the group of key entity fields.

19 . The system according to claim 11 , wherein the key entity field comprises one of alphabetical characters, alphanumeric characters, and wherein the key field value comprises alphabetical characters, numerical characters or alphanumeric characters.

20 . The system according to claim 11 , wherein identification of the hierarchical relationships in the document enables in comparing the document with another document, and wherein the document is one of a policy, a contract, an agreement, a certificate or a scheme the plurality of data elements comprising a plurality of key entity fields and key field values associated plurality of key entity fields.