AI-AUGMENTED AUDITING PLATFORM INCLUDING TECHNIQUES FOR AUTOMATED ADJUDICATION OF COMMERCIAL SUBSTANCE, RELATED PARTIES, AND COLLECTABILITY
Systems and methods for adjudicating AI-augmented automated analysis of documents in order to quickly and efficiently make various adjudications based on the documents are provided, including adjudications as to whether the documents represent underlying data that meets one or more predefined or dynamically-determined criteria. Criteria for adjudication may include commercial-substance criteria, related-party-transaction criteria, and/or collectability criteria. A system may receive a plurality of documents and generate a plurality of feature vectors by applying natural language processing techniques. The system may apply one or more classification models to the plurality of feature vectors to generate output data classifying each of the feature vectors. The system may identify, for each feature vector, a subset of closest matching prior feature vectors. Based on the classification and based on the identified subset, the system may adjudicate each feature vector with respect to commercial substance, including an adjudication classification and an adjudication confidence score.
1 . A system for classifying documents, the system comprising one or more processors configured to cause the system to:
receive data representing a document;
apply one or more natural language processing techniques to the received data to generate a feature vector representing the document;
identify, based on the feature vector, a second feature vector from a case library based on a similarity to the feature vector;
apply a plurality of models to the feature vector to compute respective changes for a plurality of characteristics represented by the document; and
determine, based on the identified second feature vector and based on the computed respective changes for the plurality of characteristics, an adjudication for the document, wherein the adjudication comprises an adjudication classification and an adjudication confidence score.
2 . The system of claim 1 , wherein:
the one or more processors are configured to identify, based on the feature vector, a cluster of feature vectors from the case library that has a highest level of similarity to the feature vector amongst feature vector clusters in the case library; and
wherein the determination of the adjudication is further based on the identified cluster of feature vectors.
3 . The system of claim 1 , wherein the plurality of characteristics comprises one or more of the following: a risk characteristic, a timing characteristic, and an amount characteristic.
4 . The system of claim 1 , wherein applying the plurality of models to the feature vector comprises computing a plurality of characteristic and comparing the plurality of computed characteristics to corresponding baseline characteristics obtained from an ERP data source to compute the respective changes.
5 . The system of any one of claim 1 , wherein computing the respective changes comprises generating a plurality of respective change values and a plurality of respective change confidence levels.
6 . The system of claim 1 , wherein applying the one or more natural language processing techniques to the received data to generate a feature vector comprises:
applying a plurality of sets of models in parallel to one another, wherein each of the sets of models is configured to process the received data to generate respective output data; and
storing the output data from each of the models in the feature vector.
7 . The system of claim 6 , wherein a first set of models of the plurality of sets of models comprises a first sentence classification module and a classification module configured to generate output data relating to a first type of content of the document.
8 . The system of claim 6 , wherein a second set of models of the plurality of sets of models comprises structural classification module, a linguistic modality classification module, and a classification module configured to generate output data relating to a second type of content of the document.
9 . The system of claim 6 , wherein a third set of models of the plurality of sets of models comprises a second sentence classification module and a classification module configured to generate output data relating to a third type of content of the document.
10 . The system of claim 1 , wherein determining the adjudication classification comprises determining whether the document meets commercial substance criteria.
11 . The system of claim 1 , wherein determining the adjudication classification and the adjudication confidence score comprises applying an adjudication reconciliation data processing operation based on data associated with the identified second feature vector and based on the computed respective changes for the plurality of characteristics.
12 . A non-transitory computer-readable storage medium storing instructions for classifying documents, the instructions configured to be executed by one or more processors to cause the system to:
receive data representing a document;
apply one or more natural language processing techniques to the received data to generate a feature vector representing the document;
identify, based on the feature vector, a second feature vector from a case library based on a similarity to the feature vector;
apply a plurality of models to the feature vector to compute respective changes for a plurality of characteristics represented by the document; and
determine, based on the identified second feature vector and based on the computed respective changes for the plurality of characteristics, an adjudication for the document, wherein the adjudication comprises an adjudication classification and an adjudication confidence score.
13 . A method for classifying documents, wherein the method is executed by a system comprising one or more processors, the method comprising:
receiving data representing a document;
applying one or more natural language processing techniques to the received data to generate a feature vector representing the document;
identifying, based on the feature vector, a second feature vector from a case library based on a similarity to the feature vector;
applying a plurality of models to the feature vector to compute respective changes for a plurality of characteristics represented by the document; and
determining, based on the identified second feature vector and based on the computed respective changes for the plurality of characteristics, an adjudication for the document, wherein the adjudication comprises an adjudication classification and an adjudication confidence score.