IP Library › Granted Patent US 12,299,543
Granted Patent B2
US 12,299,543 · App. 17/018,430 · Granted May 13, 2025

Anomalous text detection and entity identification using exploration-exploitation and pre-trained language models

Inventors: Vineet Shukla (Bangalore, IN); V Kishore Ayyadevara (Hyderabad, IN); Rohan Khilnani (Hyderabad, IN); Ravi Kumar Raju Gottumukkala (Bangalore, IN); Ankit Varshney (Delhi, IN); Rajat Gupta (Ghaziabad, IN)
Assignee: Optum Technology, Inc.
G06N20/00G06F16/3346G06F16/353G06F40/20G06N7/01G06V30/413
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,543
App. No.
17/018,430
Filed
Sep 11, 2020
Granted
May 13, 2025
Kind
B2
Art Unit
2174
USPC
706/12
Abstract

There is a need for more effective and efficient anomalous text detection. This need can be addressed by, for example, solutions for anomalous text detection that include the steps of performing a group of exploration-exploitation keyword extraction iterations based at least in part on one or more training corpus data entries until a per-iteration keyword list for an ultimate exploration-exploitation keyword extraction iteration satisfies a keyword list threshold condition; and subsequent to performing the exploration-exploitation keyword extraction iterations: processing one or more input corpus data entries using the language-model-based binary classification model to generate one or more inferred anomaly probabilities, processing the one or more input corpus data entries using the keyword model to generate explanatory metadata for the one or more inferred anomaly probabilities, and performing one or more prediction-based actions based at least in part on the one or more inferred anomaly probabilities and the explanatory metadata.

Claims (58)

1. A computer-implemented method comprising:

receiving, by one or more processors and via a plurality of exploration-exploitation keyword extraction iterations, a keyword list (i) based at least in part on a plurality of training corpus data entries and (ii) that satisfies a keyword list threshold condition, wherein an exploration-exploitation keyword extraction iteration of the plurality of exploration-exploitation keyword extraction iterations comprises:

(i) modifying the plurality of training corpus data entries by generating one or more anomalous data entries based at least in part on a prior per-iteration keyword list of a prior exploration-exploitation keyword extraction iteration that temporally precedes the exploration-exploitation keyword extraction iteration,

(ii) training a language-model-based binary classification model based at least in part on the one or more anomalous data entries, wherein the language-model-based binary classification model comprises:

(a) a bidirectional encoder layer that is configured to process the one or more anomalous data entries to generate an encoded representation,

(b) a dropout layer that is configured to process the encoded representation to generate a dropout representation,

(c) a fully-connected layer that is configured to process the dropout representation to generate a fully connected output; and

(d) a softmax layer that is configured to process the fully connected output to generate an anomaly probability for the one or more anomalous data entries,

(iii) inputting the one or more anomalous data entries to the language-model-based binary classification model to generate one or more per-entry anomaly probabilities, and

(iv) generating a per-iteration keyword list for the exploration-exploitation keyword extraction iteration based at least in part on the one or more per-entry anomaly probabilities, wherein the keyword list comprises the per-iteration keyword list for an ultimate exploration-exploitation keyword extraction iteration of the plurality of exploration-exploitation keyword extraction iterations; and

updating, by the one or more processors, a keyword model associated with the language-model-based binary classification model based at least in part on the keyword list, wherein the keyword model is configured to provide explanatory metadata for one or more inferred anomaly probabilities output by the language-model-based binary classification model.

2. The computer-implemented method of claim 1 , wherein the prior exploration-exploitation keyword extraction iteration is an initial exploration-exploitation keyword extraction iteration of the plurality of exploration-exploitation keyword extraction iterations and the one or more anomalous data entries of the initial exploration- exploitation keyword extraction iteration are generated by removing randomly-selected removable words from one or more training corpus data entries of the plurality of training corpus data entries.

3. The computer-implemented method of claim 1 , wherein updating the keyword model comprises:

inputting a keyword of the keyword of the keyword list to the bidirectional encoder layer to generate a per-keyword encoded representation of the keyword; and

generating, using a label-generating machine learning model, a labeled keyword cluster based at least in part on the per-keyword encoded representation.

4. The computer-implemented method of claim 3 , wherein generating the labeled keyword cluster comprises:

generating, using a principal component analysis layer, a reduced representation of the per-keyword encoded representation;

generating, using a t-distributed stochastic neighbor embedding layer, an unlabeled keyword cluster for the per-keyword encoded representation based at least in part on the reduced representation; and

generating, using a K-means layer, the labeled keyword cluster for the per-keyword encoded representation based on the unlabeled keyword cluster.

5. The computer-implemented method of claim 1 , wherein the keyword list threshold condition is satisfied when a deviation measure between the per-iteration keyword list and a prior per-iteration keyword list of the prior exploration-exploitation keyword extraction iteration satisfies a threshold deviation measure.

6. The computer-implemented method of claim 1 , wherein the keyword list threshold condition is satisfied when a keyword count for the per-iteration keyword list satisfies a threshold keyword count measure.

7. The computer-implemented method of claim 1 , further comprising:

inputting one or more input corpus data entries to the language-model-based binary classification model to generate the one or more inferred anomaly probabilities for the one or more input corpus data entries;

generating, using the keyword model, the explanatory metadata for the one or more inferred anomaly probabilities; and

initiating the performance of one or more prediction-based actions based at least in part on the one or more inferred anomaly probabilities and the explanatory metadata.

8. A system comprising memory and one or more processors communicatively coupled to the memory, the one or more processors configured to:

receive, via a plurality of exploration-exploitation keyword extraction iterations, a keyword list (i) based at least in part on a plurality of training corpus data entries and (ii) that satisfies a keyword list threshold condition, wherein an exploration-exploitation keyword extraction iteration of the plurality of exploration-exploitation keyword extraction iterations comprises:

(i) modifying the plurality of training corpus data entries by generating one or more anomalous data entries based at least in part on a prior per-iteration keyword list of a prior exploration-exploitation keyword extraction iteration that temporally precedes the exploration-exploitation keyword extraction iteration,

(ii) training a language-model-based binary classification model based at least in part on the one or more anomalous data entries, wherein the language-model-based binary classification model comprises:

(a) a bidirectional encoder layer that is configured to process the one or more anomalous data entries to generate an encoded representation,

(b) a dropout layer that is configured to process the encoded representation to generate a dropout representation of the input data object,

(c) a fully-connected layer that is configured to process the dropout representation to generate a fully connected output; and

(d) a softmax layer that is configured to process the fully connected output to generate an anomaly probability for the one or more anomalous data entries,

(iii) inputting the one or more anomalous data entries to the language-model-based binary classification model to generate one or more per-entry anomaly probabilities, and

(iv) generating a per-iteration keyword list for the exploration-exploitation keyword extraction iteration based at least in part on the one or more per-entry anomaly probabilities, wherein the keyword list comprises the per-iteration keyword list for an ultimate exploration-exploitation keyword extraction iteration of the plurality of exploration-exploitation keyword extraction iterations; and

update a keyword model associated with the language-model-based binary classification model based at least in part on the keyword list, wherein the keyword model is configured to provide explanatory metadata for one or more inferred anomaly probabilities output by the language-model-based binary classification model.

9. The system of claim 8 , wherein the prior exploration- exploitation keyword extraction iteration is an initial exploration-exploitation keyword extraction iteration of the plurality of exploration-exploitation keyword extraction iterations and the one or more anomalous data entries of the initial exploration-exploitation keyword extraction iteration are generated by removing randomly-selected removable words from one or more training corpus data entries of the plurality of training corpus data entries.

10. The system of claim 8 , wherein updating the keyword model comprises:

inputting a keyword of the keyword list to the bidirectional encoder layer to generate a per-keyword encoded representation of the keyword; and

generating, using a label-generating machine learning model, a labeled keyword cluster based at least in part on the per-keyword encoded representation.

11. The system of claim 10 , wherein generating the labeled keyword cluster comprises:

generating, using a principal component analysis layer, a reduced representation of the per-keyword encoded representation;

generating, using a t-distributed stochastic neighbor embedding layer, an unlabeled keyword cluster for the per-keyword encoded representation based at least in part on the reduced representation; and

generating, using a K-means layer, the labeled keyword cluster for the per-keyword encoded representation based on the unlabeled keyword cluster.

12. The system of claim 8 , wherein the keyword list threshold condition is satisfied when a keyword count for the per-iteration keyword list satisfies a threshold keyword count measure.

13. The system of claim 8 , wherein the keyword list threshold condition is satisfied when a deviation measure between the per-iteration keyword list and a prior per-iteration keyword list of the prior exploration-exploitation keyword extraction iteration satisfies a threshold deviation measure.

14. One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to:

receive, via a plurality of exploration-exploitation keyword extraction iterations, a keyword list (i) based at least in part on a plurality of training corpus data entries and (ii) that satisfies a keyword list threshold condition, wherein an exploration-exploitation keyword extraction iteration of the plurality of exploration-exploitation keyword extraction iterations comprises:

(i) modifying the plurality of training corpus data entries by generating one or more anomalous data entries based at least in part on a prior per-iteration keyword list of a prior exploration-exploitation keyword extraction iteration that temporally precedes the exploration-exploitation keyword extraction iteration,

(ii) training a language-model-based binary classification model based at least in part on the one or more anomalous data entries, wherein the language-model-based binary classification model comprises:

(a) a bidirectional encoder layer that is configured to process the one or more anomalous data entries to generate an encoded representation,

(b) a dropout layer that is configured to process the encoded representation to generate a dropout representation of the input data object,

(c) a fully-connected layer that is configured to process the dropout representation to generate a fully connected output; and

(d) a softmax layer that is configured to process the fully connected output to generate an anomaly probability for the one or more anomalous data entries,

(iii) inputting the one or more anomalous data entries to the language-model-based binary classification model to generate one or more per-entry anomaly probabilities, and

(iv) generating a per-iteration keyword list for the exploration-exploitation keyword extraction iteration based at least in part on the one or more per-entry anomaly probabilities, wherein the keyword list comprises the per-iteration keyword list for an ultimate exploration-exploitation keyword extraction iteration of the plurality of exploration-exploitation keyword extraction iterations; and

update a keyword model associated with the language-model-based binary classification model based at least in part on the keyword list, wherein the keyword model is configured to provide explanatory metadata for one or more inferred anomaly probabilities output by the language-model-based binary classification model.

15. The one or more non-transitory computer-readable storage media of claim 14 , wherein the prior exploration-exploitation keyword extraction iteration is an initial exploration-exploitation keyword extraction iteration of the plurality of exploration-exploitation keyword extraction iterations and the one or more anomalous data entries of the initial exploration-exploitation keyword extraction iteration are generated by removing randomly-selected removable words from one or more training corpus data entries of the plurality of training corpus data entries.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2020
From: SHUKLA, VINEET; AYYADEVARA, V KISHORE; KHILNANI, ROHAN; RAJU GOTTUMUKKALA, RAVI KUMAR; VARSHNEY, ANKIT; GUPTA, RAJAT
To: OPTUM TECHNOLOGY, INC.
Reel/Frame 053748/0561 →
Continuity (1)
Related Publication 20220083898A1 · Mar 17, 2022
References Cited (27)
US 7480640B1 · Elad et al. · 2009 [cited by applicant]
US 9306966B2 · Eskin et al. · 2016 [cited by applicant]
US 20100208984A1 · Bilenko · 2010 [cited by examiner]
US 20120253927A1 · Qin · 2012 [cited by examiner]
US 20170076217A1 · Krumm · 2017 [cited by examiner]
US 20170235784A1 · Seon et al. · 2017 [cited by applicant]
US 20180052849A1 · Jagmohan · 2018 [cited by examiner]
US 20180239992A1 · Chalfin · 2018 [cited by examiner]
US 20180314978A1 · Kajino · 2018 [cited by examiner]
US 20190036952A1 · Sim · 2019 [cited by examiner]
US 20190188212A1 · Miller · 2019 [cited by examiner]
US 20190287012A1 · Celikyilmaz et al. · 2019 [cited by applicant]
US 20200034436A1 · Chen · 2020 [cited by examiner]
US 20200104648A1 · Yadav · 2020 [cited by applicant]
US 20210019300A1 · Marathe · 2021 [cited by examiner]
US 20210287071A1 · Ben Fadhel · 2021 [cited by examiner]
US 20220058172A1 · John · 2022 [cited by examiner]
CN 109800219A · 2019 [cited by applicant]
IN 2225MU2015A · 2015 [cited by applicant]
“Contextual Anomaly Detection in Text Data”, by Amogh Mahapatra, taken from https://www.mdpi.com/1999-4893/5/4/469, published 2012, 30 pages. (Year: 2012). [cited by examiner]
Chalapathy, Raghavendra et al. “Deep Learning For Anomaly Detection: A Survey,” arXiv preprint arXiv:1901.03407, Jan. 10, 2019, pp. 1-50. [Retrieved from the Internet] <URL: https://arxiv.org/pdf/1901.03407.pdf>. [cited by applicant]
Denny, Matthew J. et al. “Text Preprocessing For Unsupervised Learning: Why It Matters, When It Misleads, and What To Do About It,” SSRN, Sep. 27, 2017, pp. 1-44. DOI: 10.2319.ssm.2849145. [cited by applicant]
Görnitz, Nico et al. “Toward Supervised Anomaly Detection,” Journal of Artificial Intelligence Research, Feb. 20, 2013, vol. 46, pp. 235-262. [cited by applicant]
Nag, Avishek. “Unsupervised Outlier Detection In Text Corpus Using Deep Learning,” Data Driven Investor, May 16, 2019, (17 pages). [Article, Online]. [Retrieved from the Internet Dec. 10, 2020] <URL: https://medium.com/… [cited by applicant]
Nourbakhsh, Armineh et al. “A Framework For Anomaly Detection Using Language Modeling, And Its Applications To Finance,” arXiv preprint arXiv:1908.09156. Aug. 24, 2019, (5 pages). [Retrieved from the Internet Dec. 10, 2… [cited by applicant]
Ruff, Lukas et al. “Self-Attentive, Multi-Context One-Class Classification for Unsupervised Anomaly Detection On Text,” In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 20… [cited by applicant]
Zhong, Chen et al. “Deep Actor-Critic Reinforcement Learning For Anomaly Detection,” arXiv:1908.10755v1 [cs.LG] Aug. 28, 2019, pp. 1-6. [Retrieved from the Internet Dec. 10, 2020] <URL: https://arxiv.org/pdf/1908.10755.… [cited by applicant]