IP Library Granted Patent US 11,842,286
Granted Patent B2
US 11,842,286 · App. 17/988,684 · Granted Dec 12, 2023

Machine learning platform for structuring data in organizations

Inventors: Chaithanya Manda (Jersey City, NJ); Anupam Kumar (Jersey City, NJ); Solmaz Torabi (Austin, TX); Raman Kumar (New Delhi, IN); Anish Goswami (Pune, IN); Sidhant Agarwal (Ranchi, IN); Md Sharique (Bandel, IN); Diksha Malhotra (Mohali, IN); Garimella Venkata BhanuTeja (Vijayawada, IN); Arvind Singh (Dehradun, IN); Pavan Praneeth (Hyderabad, IN)
Assignee: ExlService Holdings, Inc.
G06N5/022G06F16/93G06F40/284G06F40/295G06V30/10G06V30/416G06N3/08G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,842,286
App. No.
17/988,684
Granted
Dec 12, 2023
Kind
B2
Abstract

An application instance that includes one or more machine learning models receives, from a subscriber computing system, a document comprising unstructured data. Based on the unstructured data, the application instance generates an optimized model input that includes a plurality of parsed document sections. For each parsed document section, the application instance generates an output set by performing, by a machine learning model, at least one key information extraction operation. The machine learning model transmits the output in structured form to a target application operated or hosted at least in part by a subscriber entity associated with the subscriber computing system.

Claims (53)

1. One or more non-transitory computer readable storage media excluding transitory signals, the storage media storing instructions, which when executed by at least one processor, cause the at least one processor to perform operations comprising:

receiving, at an application instance communicatively coupled to a subscriber computing system of a plurality of subscriber computing systems, a document comprising a first set of unstructured data;

performing pre-processing operations on at least a portion of the document comprising the first set of unstructured data, the pre-processing operations comprising generating an optimized model input comprising a plurality of parsed document sections;

for each parsed document section in the plurality of parsed document sections, generating a first output set by performing, by a trained machine learning model, at least one key information extraction operation, the key information extraction operation comprising one or more of extractive summarization, named entity extraction, or classification,

wherein the at least one trained machine learning model is trained using two or more of:

a second set of unstructured data previously received from the subscriber computing system,

a reference ontology,

a second output set corresponding to the second set of unstructured data, or

a confidence value for an item in the second output set; and

transmitting the first output set to a target application operated or hosted at least in part by the subscriber computing system or a subscriber entity associated with the subscriber computing system.

2. The media of claim 1 , wherein the subscriber computing system and the target application are provided by the subscriber entity, and wherein the application instance is provided by a provider entity different from the subscriber entity.

3. The media of claim 1 , wherein the application instance is on a virtual network associated with the subscriber entity.

4. The media of claim 1 , wherein the operations further comprise splitting the document comprising unstructured data into at least a first section and a second section, and wherein separate pre-processing operations are performed on the first section and the second section.

5. The media of claim 1 , wherein the document comprising unstructured data comprises at least one linked sub-document in a collection of sub-document objects relationally linked to the document, the pre-processing operations comprising:

traversing the collection of sub-document objects; and

for each sub-document object in the collection of sub-document objects, retrievably storing the sub-document object; and

generating a separate optimized model input.

6. The media of claim 5 , wherein the document comprising unstructured data comprises embedded markup-language code, and wherein traversing the collection of sub-document objects comprises retrieving an object indicated by the markup-language code.

7. The media of claim 5 , wherein the document comprising unstructured data is an email message, and wherein traversing the collection of sub-document objects comprises determining and accessing at least one attachment associated with the email message.

8. The media of claim 1 , wherein the document comprising unstructured data includes a combination of text and images, and wherein the pre-processing operations further comprise scanning the document to identify respective text or images.

9. The media of claim 1 ,

wherein the pre-processing operations comprise at least one of performing optical character recognition on at least a portion of the document comprising unstructured data or extracting a form field from the document comprising unstructured data,

and wherein the portion of the document is identified by determining relative coordinates of the portion of the document by a pre-processing convolutional neural network structured to generate the optimized model input.

10. The media of claim 1 , wherein generating optimized model input comprises applying a domain-specific ontology to the document.

11. The media of claim 10 , wherein the domain-specific ontology relates to at least one of a health condition or a medication.

12. The media of claim 1 , wherein generating optimized model input comprises determining a type of output needed based on at least one of: a previously stored setting, a subscriber-defined runtime parameter or a feature of the target application.

13. The media of claim 1 , wherein the first output set comprises a plurality of key-value pairs.

14. The media of claim 1 , wherein the first output set comprises a table.

15. The media of claim 1 , wherein the optimized model input comprises a plurality of sentence tokens, and wherein the first output set comprises a summary sentence generated based on the sentence tokens.

16. The media of claim 1 , wherein the optimized model input comprises a key-value pair.

17. A computer-implemented method, the method comprising:

receiving, at an application instance communicatively coupled to a subscriber computing system of a plurality of subscriber computing systems, a document comprising a first set of unstructured data;

performing pre-processing operations on at least a portion of the document comprising the first set of unstructured data, the pre-processing operations comprising generating an optimized model input comprising a plurality of parsed document sections;

for each parsed document section in the plurality of parsed document sections, generating a first output set by performing, by a trained machine learning model, at least one key information extraction operation, the key information extraction operation comprising one or more of extractive summarization, named entity extraction, or classification,

wherein the at least one trained machine learning model is trained using two or more of:

a second set of unstructured data previously received from the subscriber computing system,

a reference ontology,

a second output set corresponding to the second set of unstructured data, or

a confidence value for an item in the second output set; and

transmitting the first output set to a target application operated or hosted at least in part by the subscriber computing system or a subscriber entity associated with the subscriber computing system.

18. The method of claim 17 , wherein the application instance is on a virtual network associated with the subscriber entity.

19. A provider computing system comprising at least one processor, at least one memory, and one or more non-transitory computer readable medium storing instructions, which when executed by the at least one processor, perform operations comprising:

receive, at an application instance of the provider computing system, the application instance communicatively coupled to a subscriber computing system of a plurality of subscriber computing systems, a document comprising a first set of unstructured data;

perform pre-processing operations on at least a portion of the document comprising the first set of unstructured data, the pre-processing operations comprising generating an optimized model input comprising a plurality of parsed document sections;

for each parsed document section in the plurality of parsed document sections, generate a first output set, comprising operations to perform, by a trained machine learning model, at least one key information extraction operation, the key information extraction operation comprising one or more of extractive summarization, named entity extraction, or classification,

wherein the at least one trained machine learning model is trained using two or more of:

a second set of unstructured data previously received from the subscriber computing system,

a reference ontology,

a second output set corresponding to the second set of unstructured data, or

a confidence value for an item in the second output set; and

transmit the first output set to a target application operated or hosted at least in part by the subscriber computing system or a subscriber entity associated with the subscriber computing system.

20. The system of claim 19 , wherein the application instance is on a virtual network associated with the subscriber entity.

21. The system of claim 19 , wherein the reference ontology comprises at least one of medication unit information or medication dosage information.

Assignments (1)
SECURITY INTEREST Recorded Aug 18, 2026
From: EXLSERVICE HOLDINGS, INC.; OVERLAND SOLUTIONS, LLC; EXLSERVICE TECHNOLOGY SOLUTIONS, LLC
To: PNC BANK, NATIONAL ASSOCIATION
Reel/Frame 075693/0088 →
Continuity (2)
Provisional Application 63280062 · Nov 16, 2021
Related Publication 20230153641A1 · May 18, 2023
Cited By (3)
US 12,260,342 US 12,632,473 US 12,645,988