IP Library Granted Patent US 12,536,822
Granted Patent B2
US 12,536,822 · App. 18/343,816 · Granted Jan 27, 2026

Systems and methods for extracting key performance indicators from textual data

Inventors: Shreyansh S Nanawati (Kalyan, IN); Deeksha Thareja (New Delhi, IN); Mrityunjai Singh (Noida, IN); Aviral Sharma (Jaipur, IN); Jatin Lamba (Kheri, IN)
Assignee: Optum, Inc.
G06V30/148G06Q10/06393G06V30/413
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,822
App. No.
18/343,816
Granted
Jan 27, 2026
Kind
B2
Abstract

A method includes: determining a first bounding box in textual data; categorizing the first bounding box into a paragraph category among a list of categories including a header category, a section header category, the paragraph category, and a noise category; extracting a candidate noun from the text in the categorized first bounding box, as a candidate paragraph noun; extracting a candidate value from the text in the categorized first bounding box; generating a relationship between the candidate paragraph noun and the extracted candidate value; associating the candidate paragraph noun with a candidate header noun; and generating a key performance indicator using the candidate header noun, the candidate paragraph noun, the candidate value, and the generated relationship.

Claims (50)

1 . A method, performed by one or more processors of a computing system, for extracting a key performance indicator from textual data, the method comprising:

determining a first bounding box in textual data, wherein the first bounding box includes text grouped by position in the textual data;

categorizing the first bounding box into a paragraph category among a list of categories including a header category, a section header category, the paragraph category, and a noise category;

extracting a candidate noun from the text in the categorized first bounding box, as a candidate paragraph noun;

extracting a candidate value from the text in the categorized first bounding box;

generating a relationship between the candidate paragraph noun and the extracted candidate value;

associating the candidate paragraph noun with a candidate header noun, wherein the candidate header noun is an extracted candidate noun from text in a second bounding box categorized into the header category or the section header category; and

generating a key performance indicator using the candidate header noun, the candidate paragraph noun, the candidate value, and the generated relationship.

2 . The method of claim 1 , wherein determining the first bounding box includes using a segmentation model.

3 . The method of claim 2 , wherein the segmentation model uses image features, positional features, and text features to merge close and/or overlapping bounding boxes.

4 . The method of claim 1 , wherein extracting the candidate noun from the text in the categorized first bounding box includes using part-of-speech tags to extract the candidate noun from a sentence in the textual data.

5 . The method of claim 1 , wherein extracting the candidate value from the text in the categorized first bounding box includes using a pre-trained Named Entity Recognition model to extract the candidate value from a sentence in the textual data.

6 . The method of claim 1 , wherein generating the relationship includes using a classification model.

7 . The method of claim 6 , wherein the classification model is based on (i) word embedding of the candidate noun, (ii) word embedding of the candidate value, and (iii) averaged word embeddings of tokens present between the candidate noun and the candidate value.

8 . The method of claim 7 , wherein the classification model is a binary classifier which predicts whether a relationship exists between the candidate noun and the candidate value.

9 . The method of claim 1 , wherein extracting the candidate noun from the text in the categorized first bounding box includes extracting noun phrases from a sentence in the textual data.

10 . The method of claim 1 , wherein generating the relationship includes filtering candidate nouns that result in redundant key performance indicators.

11 . The method of claim 1 , wherein generating the key performance indicator includes standardizing the generated key performance indicator.

12 . The method of claim 1 , wherein the textual data is generated by performing optical character recognition on a document.

13 . The method of claim 12 , wherein the one or more processors of the computing system execute instructions for:

a text extractor to generate the textual data,

a segmentation generator to determine the first bounding box,

a candidate noun extractor to extract the candidate noun,

a candidate value extractor to extract the candidate value,

a relationship generator to generate the relationship,

an association generator to associate the candidate paragraph noun with the candidate header noun, and

a standardization generator to generate the key performance indicator.

14 . A system for extracting a key performance indicator from textual data, the system comprising:

one or more processors; and

at least one memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including:

determining a first bounding box in textual data, wherein the first bounding box includes text grouped by position in the textual data;

categorizing the first bounding box into a paragraph category among a list of categories including a header category, a section header category, the paragraph category, and a noise category;

extracting a candidate noun from the text in the categorized first bounding box, as a candidate paragraph noun;

extracting a candidate value from the text in the categorized first bounding box;

generating a relationship between the candidate paragraph noun and the extracted candidate value;

associating the candidate paragraph noun with a candidate header noun, wherein the candidate header noun is an extracted candidate noun from text in a second bounding box categorized into the header category or the section header category; and

generating a key performance indicator using the candidate header noun, the candidate paragraph noun, the candidate value, and the generated relationship.

15 . The system of claim 14 , wherein determining the first bounding box includes using a segmentation model that uses image features, positional features, and text features to optimize and merge close and overlapping bounding boxes.

16 . The system of claim 14 , wherein extracting the candidate noun from the text in the categorized first bounding box includes using part-of-speech tags to extract the candidate noun from a sentence in the textual data.

17 . The system of claim 14 , wherein extracting the candidate value from the text in the categorized first bounding box includes using a pre-trained Named Entity Recognition model to extract the candidate value from a sentence in the textual data.

18 . The system of claim 14 , wherein generating the relationship includes using a classification model that is a binary classifier which predicts whether a relationship exists between the candidate noun and the candidate value, based on (i) word embedding of the candidate noun, (ii) word embedding of the candidate value, and (iii) averaged word embeddings of tokens present between the candidate noun and the candidate value.

19 . The system of claim 14 , wherein extracting the candidate noun from the text in the categorized first bounding box includes extracting noun phrases from a sentence in the textual data.

20 . A non-transitory computer readable medium for extracting a key performance indicator from textual data, the non-transitory computer readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

determining a first bounding box in textual data, wherein the first bounding box includes text grouped by position in the textual data;

categorizing the first bounding box into a paragraph category among a list of categories including a header category, a section header category, the paragraph category, and a noise category;

extracting a candidate noun from the text in the categorized first bounding box, as a candidate paragraph noun;

extracting a candidate value from the text in the categorized first bounding box;

generating a relationship between the candidate paragraph noun and the extracted candidate value;

associating the candidate paragraph noun with a candidate header noun, wherein the candidate header noun is an extracted candidate noun from text in a second bounding box categorized into the header category or the section header category; and

generating a key performance indicator using the candidate header noun, the candidate paragraph noun, the candidate value, and the generated relationship.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2023
From: NANAWATI, SHREYANSH S; THAREJA, DEEKSHA; SINGH, MRITYUNJAI; SHARMA, AVIRAL; LAMBA, JATIN
To: OPTUM, INC.
Reel/Frame 064108/0510 →
Continuity (1)
Related Publication 20250005944A1 · Jan 2, 2025
References Cited (8)
US 10565502B2 · Scholtes · 2020 [cited by applicant]
US 10599767B1 · Mattera et al. · 2020 [cited by applicant]
US 11080295B2 · Chang et al. · 2021 [cited by applicant]
US 20180075128A1 · Srinivasan · 2018 [cited by examiner]
US 20220181029A1 · Colley et al. · 2022 [cited by applicant]
US 20240311578A1 · Laudij · 2024 [cited by examiner]
Kamaruddin, et al. (Automatic Extraction of Performance Indicators from Financial Statements). (Year: 2009). [cited by examiner]
Hillebrand et al., “KPI-Bert: A Joint Named Entity Recognition and Relation Extraction Model for Financial Reports,” IEEE, published on Aug. 3, 2022, retrieved from https://arxiv.org/pdf/2208.02140.pdf. [cited by applicant]