IP Library › Granted Patent US 12,197,462
Granted Patent B2
US 12,197,462 · App. 18/153,624 · Granted Jan 14, 2025

Focusing unstructured data and generating focused data determinations from an unstructured data set

Inventors: Colum Foley (County Dublin, IE); Paul Ferguson (Dublin, IE)
Assignee: Optum Services (Ireland) Limited
G06F16/258G06N20/00G06Q40/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,197,462
App. No.
18/153,624
Granted
Jan 14, 2025
Kind
B2
Abstract

Embodiments provide for improvements in generating focused data from an unstructured data set. The focused data generated from the unstructured data set may provide data insight(s) into analysis of the unstructured data set, and/or provide for improved capabilities for a user to efficiently navigate through relevant data portions of such data utilizing a user interface, even when such relevant data portions are not immediately distinguishable without further processing of the unstructured data set. Some embodiments receive an unstructured data set, extract an identified relevant subset utilizing at least one high-level extractor model, extract low-level relevant data from the identified relevant subset utilizing at least one low-level extractor model, generate fraud probability data by applying at least the low-level relevant data and the identified relevant subset to a fraud processing model, and output at least the fraud probability data, identified relevant subset, and/or low-level relevant data, and/or derivations therefrom.

Claims (53)

1. A computer-implemented method comprising:

extracting, by one or more processors, an initial keyword set from an identified relevant subset of an unstructured data set, wherein:

(i) the identified relevant subset is generated based at least in part on at least one high-level extractor model, and

(ii) the initial keyword set is extracted based at least in part on a keyword extraction model that generates a keyword relevance score for an initial keyword of the initial keyword set;

identifying, by the one or more processors, an irrelevant keyword based at least in part on a keyword relevance threshold and the keyword relevance score for the initial keyword of the initial keyword set;

generating, by the one or more processors, an updated keyword set by removing the irrelevant keyword from the initial keyword set;

removing, by the one or more processors and from the updated keyword set, an unknown keyword;

generating, by the one or more processors, a filtered keyword set by applying a dictionary filter model to the updated keyword set; and

outputting, by the one or more processors, at least one keyword from the filtered keyword set.

2. The computer-implemented method of claim 1 , wherein outputting the at least one keyword from the filtered keyword set comprises:

causing rendering of the at least one keyword to a user interface.

3. The computer-implemented method of claim 1 , wherein the dictionary filter model is based at least in part on a central truth source.

4. The computer-implemented method of claim 1 , wherein the initial keyword set is extracted utilizing a local interpretable model-agnostic (LIME) model, a Shapley additive explanations (SHAP) model, or an attention-based machine-learning model.

5. The computer-implemented method of claim 1 , wherein the keyword extraction model generates trusted description data associated with the initial keyword of the initial keyword set, and wherein the computer-implemented method further comprises:

calculating the keyword relevance score for the initial keyword based at least in part on a similarity of a data portion associated with the initial keyword of the initial keyword set with the trusted description data.

6. The computer-implemented method of claim 1 , wherein the keyword relevance score does not exceed the keyword relevance threshold.

7. The computer-implemented method of claim 1 , wherein generating the updated keyword set by removing the irrelevant keyword from the initial keyword set comprises:

determining that the irrelevant keyword is not present in a known dictionary.

8. A computing system comprising a processor and memory including program code, the memory and the program code configured to, when executed by the processor, cause the computing system to:

extract an initial keyword set from an identified relevant subset of an unstructured data set, wherein

(i) the identified relevant subset is generated based at least in part on at least one high-level extractor model, and

(ii) the initial keyword set is extracted based at least in part on a keyword extraction model that generates a keyword relevance score for an initial keyword of the initial keyword set;

identify an irrelevant keyword based at least in part on a keyword relevance threshold and the keyword relevance score for the initial keyword of the initial keyword set;

generate an updated keyword set by removing the irrelevant keyword from the initial keyword set;

remove, from the updated keyword set, an unknown keyword;

generate a filtered keyword set by applying a dictionary filter model to the updated keyword set; and

output at least one keyword from the filtered keyword set.

9. The computing system of claim 8 , wherein outputting the at least one keyword from the filtered keyword set comprises:

cause rendering of the at least one keyword to a user interface.

10. The computing system of claim 8 , wherein the dictionary filter model is based at least in part on a central truth source.

11. The computing system of claim 8 , wherein the initial keyword set is extracted utilizing a local interpretable model-agnostic (LIME) model, a Shapley additive explanations (SHAP) model, or an attention-based machine-learning model.

12. The computing system of claim 8 , wherein the keyword extraction model generates trusted description data associated with the initial keyword of the initial keyword set, and wherein the computing system is further caused to:

calculate the keyword relevance score for the initial keyword based at least in part on a similarity of a data portion associated with the initial keyword of the initial keyword set with the trusted description data.

13. The computing system of claim 8 , wherein the keyword relevance score does not exceed the keyword relevance threshold.

14. The computing system of claim 8 , wherein to generate the updated keyword set by removing the irrelevant keyword from the initial keyword set the computing system is further caused to:

determine that the irrelevant keyword is not present in a known dictionary.

15. A computer program product comprising a non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium including instructions that, when executed by a computing system, cause the computing system to:

extract an initial keyword set from an identified relevant subset of an unstructured data set, wherein:

(i) the identified relevant subset is generated based at least in part on at least one high-level extractor model, and

(ii) the initial keyword set is extracted based at least in part on a keyword extraction model that generates a keyword relevance score for an initial keyword of the initial keyword set;

identify an irrelevant keyword based at least in part on a keyword relevance threshold and the keyword relevance score for the initial keyword of the initial keyword set;

generate an updated keyword set by removing the irrelevant keyword from the initial keyword set;

remove, from the updated keyword set, an unknown keyword;

generate a filtered keyword set by applying a dictionary filter model to the updated keyword set; and

output at least one keyword from the filtered keyword set.

16. The computer program product of claim 15 , wherein to output the at least one keyword from the filtered keyword set the computing system is caused to:

cause rendering of the at least one keyword to a user interface.

17. The computer program product of claim 15 , wherein the dictionary filter model is based at least in part on a central truth source.

18. The computer program product of claim 15 , wherein the initial keyword set is extracted utilizing a local interpretable model-agnostic (LIME) model, a Shapley additive explanations (SHAP) model, or an attention-based machine-learning model.

19. The computer program product of claim 15 , wherein the keyword extraction model generates trusted description data associated with the initial keyword of the initial keyword set, and wherein the computing system is further caused to:

calculate the keyword relevance score for the initial keyword based at least in part on a similarity of a data portion associated with the initial keyword of the initial keyword set with the trusted description data.

20. The computer program product of claim 15 , wherein to generate the updated keyword set by remove the irrelevant keyword from the initial keyword set the computing system is further caused to:

determine that the at least one irrelevant keyword is not present in a known dictionary.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2023
From: FERGUSON, PAUL; FOLEY, COLUM
To: OPTUM SERVICES (IRELAND) LIMITED
Reel/Frame 062360/0446 →
Continuity (3)
Continuation 18153439 · Jan 12, 2023
Provisional Application 63374252 · Sep 1, 2022
Related Publication 20240078610A1 · Mar 7, 2024
References Cited (46)
US 10102340B2 · Tanner, Jr. et al. · 2018 [cited by applicant]
US 10319468B2 · Ginsburg · 2019 [cited by applicant]
US 10552576B2 · Campbell · 2020 [cited by applicant]
US 11049594B2 · Chintamaneni et al. · 2021 [cited by applicant]
US 11361381B1 · Lehmuth et al. · 2022 [cited by applicant]
US 20110258054A1 · Pandey · 2011 [cited by examiner]
US 20130006655A1 · Van Arkel et al. · 2013 [cited by applicant]
US 20130054259A1 · Wojtusiak et al. · 2013 [cited by applicant]
US 20140081652A1 · Klindworth · 2014 [cited by applicant]
US 20160239617A1 · Farooq et al. · 2016 [cited by applicant]
US 20160253461A1 · Sohr et al. · 2016 [cited by applicant]
US 20170017760A1 · Freese et al. · 2017 [cited by applicant]
US 20170322930A1 · Drew · 2017 [cited by examiner]
US 20190034589A1 · Chen et al. · 2019 [cited by applicant]
US 20200013124A1 · Obee et al. · 2020 [cited by applicant]
US 20200104731A1 · Oliner · 2020 [cited by examiner]
US 20200311601A1 · Robinson et al. · 2020 [cited by applicant]
US 20200381090A1 · Apostolova et al. · 2020 [cited by applicant]
US 20210056113A1 · Macan tSaoir · 2021 [cited by examiner]
US 20210109915A1 · Godden · 2021 [cited by examiner]
US 20210304749A1 · Singh · 2021 [cited by examiner]
US 20210313022A1 · Chaballout · 2021 [cited by applicant]
US 20220012611A1 · Moradi et al. · 2022 [cited by applicant]
US 20220309592A1 · Zahora et al. · 2022 [cited by applicant]
US 20220392048A1 · Henry et al. · 2022 [cited by applicant]
US 20230048097A1 · Clausen et al. · 2023 [cited by applicant]
US 20230101817A1 · Sinha · 2023 [cited by examiner]
US 20230131694A1 · Saber et al. · 2023 [cited by applicant]
US 20230195443A1 · Eberlein et al. · 2023 [cited by applicant]
US 20230224493A1 · Foley et al. · 2023 [cited by applicant]
US 20230385705A1 · Takehara et al. · 2023 [cited by applicant]
US 20240078245A1 · Foley et al. · 2024 [cited by applicant]
US 20240078609A1 · Foley et al. · 2024 [cited by applicant]
CN 108334935A · 2018 [cited by applicant]
CN 109636061A · 2019 [cited by applicant]
CN 116070630A · 2023 [cited by applicant]
WO 2022057057A1 · 2022 [cited by applicant]
NonFinal Office Action for U.S. Appl. No. 18/153,602, dated Dec. 20, 2023, (12 pages), United States Patent and Trademark Office, US. [cited by applicant]
Jain, Ankit. “Claim Analysis and Fraud Detection Using Business Intelligence|Blog,” Nalashaa, Feb. 2, 2017, (3 pages), [Retrieved from the Internet Oct. 7, 2022] <URL: https://www.nalashaa.com/claim-analysis-fraud-detec… [cited by applicant]
Kim, Byung-Hak et al. “Deep Claim: Payer Response Prediction From Claims Data With Deep Learning,” arXiv preprint arXiv:2007.06229v1 [cs.LG], Jul. 13, 2020, (9 pages). [cited by applicant]
Sowah, Robert A. et al. “Decision Support System (DSS) for Fraud Detection in Health Insurance Claims Using Genetic Support Vector Machines (GSVMs),” Hindawi Journal of Engineering, vol. 2019, Article ID 1432597, pp. 1-… [cited by applicant]
Sun, Xu et al. “Feature-Frequency-Adaptive On-Line Training For Fast and Accurate Natural Language Processing,” Computational Linguistics, vol. 40, No. 3, Sep. 1, 2014, pp. 563-586, DOI: 10.1162/COLI_a_00193. [cited by applicant]
Thesmar, David et al. “Combining The Power Of Artificial Intelligence With The Richness Of Healthcare Claims Data: Opportunities and Challenges,” PharmacoEconomics, vol. 37, pp. 745-752, Mar. 8, 2019. [cited by applicant]
Non-Final Rejection Mailed on Mar. 26, 2024 for U.S. Appl. No. 17/805,340, 19 page(s). [cited by applicant]
Advisory Action (PTOL-303) Mailed on Aug. 20, 2024 for U.S. Appl. No. 18/153,602, 3 page(s). [cited by applicant]
Notice of Allowance and Fees Due (PTOL-85) for U.S. Appl. No. 17/805,340, filed Jul. 26, 2024, 8 pages, United States Patent and Trademark Office, US. [cited by applicant]