IP Library Granted Patent US 12675732
Granted Patent B2
US 12675732 · App. 17/660,940 · Granted Jul 7, 2026

Machine learning techniques for context-based document classification

Inventors: Rahul Singh (Bengaluru, IN); Rakesh P A (Bengaluru, IN); Vijaychandar Natesan (Bangalore, IN); Ramesh R. Ganesan (Bangalore, IN); Maureen E. Calkins (Hamburg, NY); Caroline N. Eccleston (Boise, ID)
Assignee: UnitedHealth Group Incorporated
G06N20/00G06F16/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12675732
App. No.
17/660,940
Granted
Jul 7, 2026
Kind
B2
Abstract

Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and/or the like for performing context-based document classification prediction using a hierarchical attention-based keyword classifier machine learning framework. Certain embodiments of the present invention utilize systems, methods, and computer program products that perform context-based document classification prediction using at least one of techniques using contextual keyword classifications, techniques using attention-based keyword classifier machine learning framework, techniques using a greedy matching indicator, and/or the like.

Claims (64)

1 . A computer-implemented method comprising:

identifying, by one or more processors, a keyword sequence associated with a plurality of keywords, wherein the plurality of keywords is a subset of a keyword repository comprising a group of candidate keywords;

performing, by the one or more processors, a plurality of contextual keyword classification routine iterations, wherein: (i) a contextual keyword classification routine iteration of the plurality of contextual keyword classification routine iterations is associated with a corresponding sequentially-selected keyword in the keyword sequence, and (ii) the contextual keyword classification routine iteration is configured to:

generate, using an attention-based keyword classifier machine learning framework and based at least in part on the corresponding sequentially-selected keyword, a contextual keyword classification,

determine a greedy matching indicator for the corresponding sequentially-selected keyword based at least in part on whether the contextual keyword classification for the corresponding sequentially-selected keyword is an affirmative contextual keyword classification by:

(1) identifying a document page associated with the corresponding sequentially-selected keyword,

(2) generating a page template for the document page, and

(3) in response to determining that the page template corresponds to at least one of a plurality of exclusionary page templates, determining a negative greedy matching indicator, and

in response to determining that the greedy matching indicator is an affirmative greedy matching indicator: (a) determine that a document data object is associated with an affirmative document classification, and (b) terminate the plurality of contextual keyword classification routine iterations; and

initiating, by the one or more processors, performance of one or more prediction-based actions based at least in part on the affirmative document classification.

2 . The computer-implemented method of claim 1 , wherein determining the plurality of exclusionary page templates comprises:

identifying a plurality of excluded training document pages;

generating a plurality of training page signatures corresponding to the plurality of excluded training document pages;

generating, based at least in part on the plurality of training page signatures, an excluded training document page cluster comprising a clustered document page subset of the plurality of excluded training document pages; and

generating an exclusionary page template for the excluded training document page cluster based at least in part on, training page signature, of the plurality of training page signatures, for the clustered document page subset.

3 . The computer-implemented method of claim 2 , wherein generating a particular exclusionary page template that is associated with a particular excluded training document page cluster comprises:

generating the particular exclusionary page template based at least in part on a training page signature distribution for the clustered document page subset of the particular exclusionary page template.

4 . The computer-implemented method of claim 3 , wherein generating the particular exclusionary page template is performed based at least in part on a page name distribution for the clustered document page subset of the particular exclusionary page template and a keyword distribution for the particular exclusionary page template.

5 . The computer-implemented method of claim 4 , wherein the page template for the document page is determined based at least in part on a page name for the document page and a keyword set for the document page.

6 . The computer-implemented method of claim 1 , wherein: (a) the attention-based keyword classifier machine learning framework comprises an attention-based encoder machine learning model and a keyword classifier machine learning model, (b) the attention-based encoder machine learning model is configured to generate an attention-based keyword encoding for the corresponding sequentially-selected keyword based at least in part on a defined number of context keywords of the keyword repository that occur within an attention window for the corresponding sequentially-selected keyword, and (c) the keyword classifier machine learning model is configured to generate the contextual keyword classification for the corresponding sequentially-selected keyword based at least in part on the attention-based keyword encoding.

7 . The computer-implemented method of claim 6 , wherein the attention-based encoder machine learning model is configured to perform operations corresponding to a bidirectional self-attention mechanism.

8 . A system comprising:

one or more processors; and

at least one memory storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to:

identify a keyword sequence associated with a plurality of keywords, wherein the plurality of keywords is a subset of a keyword repository comprising a group of candidate keywords;

perform a plurality of contextual keyword classification routine iterations, wherein: (i) a contextual keyword classification routine iteration of the plurality of contextual keyword classification routine iterations is associated with a corresponding sequentially-selected keyword in the keyword sequence, and (ii) the contextual keyword classification routine iteration is configured to:

generate, using an attention-based keyword classifier machine learning framework and based at least in part on the corresponding sequentially-selected keyword, a contextual keyword classification,

determine a greedy matching indicator for the corresponding sequentially-selected keyword based at least in part on whether the contextual keyword classification for the corresponding sequentially-selected keyword is an affirmative contextual keyword classification by:

(1) identifying a document page associated with the corresponding sequentially-selected keyword,

(2) generating a page template for the document page, and

(3) in response to determining that the page template corresponds to at least one of a plurality of exclusionary page templates, determining a negative greedy matching indicator, and

in response to determining that the greedy matching indicator is an affirmative greedy matching indicator: (a) determine that a document data object is associated with an affirmative document classification, and (b) terminate the plurality of contextual keyword classification routine iterations; and

initiate performance of one or more prediction-based actions based at least in part on the affirmative document classification.

9 . The system of claim 8 , wherein determining the plurality of exclusionary page templates comprises:

identifying a plurality of excluded training document pages;

generating a plurality of training page signatures corresponding to the plurality of excluded training document pages;

generating, based at least in part on the plurality of training page signatures, an excluded training document page cluster comprising a clustered document page subset of the plurality of excluded training document pages; and

generating an exclusionary page template for the excluded training document page cluster based at least in part on a training page signature, of the plurality of training page signatures, for the clustered document page subset.

10 . The system of claim 1 , wherein generating a particular exclusionary page template that is associated with a particular excluded training document page cluster comprises:

generating the particular exclusionary page template based at least in part on a training page signature distribution for the clustered document page subset of the particular exclusionary page template.

11 . The system of claim 10 , wherein generating the particular exclusionary page template is performed based at least in part on a page name distribution for the clustered document page subset of the particular exclusionary page template and a keyword distribution for the particular exclusionary page template.

12 . The system of claim 11 , wherein the page template for the document page is determined based at least in part on a page name for the document page and a keyword set for the document page.

13 . The system of claim 8 , wherein: (a) the attention-based keyword classifier machine learning framework comprises an attention-based encoder machine learning model and a keyword classifier machine learning model, (b) the attention-based encoder machine learning model is configured to generate an attention-based keyword encoding for the corresponding sequentially-selected keyword based at least in part on a defined number of context keywords of the keyword repository that occur within an attention window for the corresponding sequentially-selected keyword, and (c) the keyword classifier machine learning model is configured to generate the contextual keyword classification for the corresponding sequentially-selected keyword based at least in part on the attention-based keyword encoding.

14 . The system of claim 13 , wherein the attention-based encoder machine learning model is configured to perform operations corresponding to a bidirectional self-attention mechanism.

15 . One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to:

identify a keyword sequence associated with a plurality of keywords, wherein the plurality of keywords is a subset of a keyword repository comprising a group of candidate keywords;

perform a plurality of contextual keyword classification routine iterations, wherein: (i) a contextual keyword classification routine iteration of the plurality of contextual keyword classification routine iterations is associated with a corresponding sequentially-selected keyword in the keyword sequence, and (ii) the contextual keyword classification routine iteration is configured to:

generate, using an attention-based keyword classifier machine learning framework and based at least in part on the corresponding sequentially-selected keyword, a contextual keyword classification,

determine a greedy matching indicator for the corresponding sequentially-selected keyword based at least in part on whether the contextual keyword classification for the corresponding sequentially-selected keyword is an affirmative contextual keyword classification by:

(1) identifying a document page associated with the corresponding sequentially-selected keyword,

(2) generating a page template for the document page, and

(3) in response to determining that the page template corresponds to at least one of a plurality of exclusionary page templates, determining a negative greedy matching indicator, and

in response to determining that the greedy matching indicator is an affirmative greedy matching indicator: (a) determine that a document data object is associated with an affirmative document classification, and (b) terminate the plurality of contextual keyword classification routine iterations; and

initiate performance of one or more prediction-based actions based at least in part on the affirmative document classification.

16 . The one or more non-transitory computer-readable storage media of claim 15 , wherein determining the plurality of exclusionary page templates comprises:

identifying a plurality of excluded training document pages;

generating a plurality of training page signatures corresponding to the plurality of excluded training document pages;

generating, based at least in part on the plurality of training page signatures, an excluded training document page cluster comprising a clustered document page subset of the plurality of excluded training document pages; and

generating an exclusionary page template for the excluded training document page cluster based at least in part on a training page signature, of the plurality of training page signatures, for the clustered document page subset.

17 . The one or more non-transitory computer-readable storage media of claim 16 , wherein generating a particular exclusionary page template that is associated with a particular excluded training document page cluster comprises:

generating the particular exclusionary page template based at least in part on a training page signature distribution for the clustered document page subset of the particular exclusionary page template.

18 . The one or more non-transitory computer-readable storage media of claim 17 , wherein generating the particular exclusionary page template is performed based at least in part on a page name distribution for the clustered document page subset of the particular exclusionary page template and a keyword distribution for the particular exclusionary page template.

19 . The one or more non-transitory computer-readable storage media of claim 18 , wherein the page template for the document page is determined based at least in part on a page name for the document page and a keyword set for the document page.

20 . The one or more non-transitory computer-readable storage media of claim 15 , wherein: (a) the attention-based keyword classifier machine learning framework comprises an attention-based encoder machine learning model and a keyword classifier machine learning model, (b) the attention-based encoder machine learning model is configured to generate an attention-based keyword encoding for the corresponding sequentially-selected keyword based at least in part on a defined number of context keywords of the keyword repository that occur within an attention window for the corresponding sequentially-selected keyword, and (c) the keyword classifier machine learning model is configured to generate the contextual keyword classification for the corresponding sequentially-selected keyword based at least in part on the attention-based keyword encoding.