IP Library › Granted Patent US 12,190,252
Granted Patent B2
US 12,190,252 · App. 17/141,775 · Granted Jan 7, 2025

Explainable unsupervised vector representation of multi-section documents

Inventors: Riccardo Mattivi (Dublin, IE); Peter Cogan (Dublin, IE)
Assignee: Optum Services (Ireland) Limited
G06N5/04G06F16/345G06F18/23213G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,252
App. No.
17/141,775
Granted
Jan 7, 2025
Kind
B2
Abstract

Embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for generating an inferred document representation for a multi-section document using a machine learning model. In accordance with one embodiment, a method is provided that includes: identifying a document corpus comprising the multi-section document and other multi-section documents; for each section of the document that is associated with a section type identifier: identifying a section batch that comprises common-type sections across the document corpus; and processing the section batch using the machine learning model to generate per-type section clusters for the section type identifier that comprise an inferred per-type section cluster for the current section; generating the inferred document representation based at least in part on each inferred per-type section cluster for a section of the document; and performing a prediction-based action based at least in part on the representation.

Claims (75)

1. A computer-implemented method for automated generation of an inferred document representation for a multi-section document to convey detailed content granularity of semantic characteristics within the multi-section document, the computer-implemented method comprising:

identifying, by one or more processors, a document corpus comprising the multi-section document and a plurality of other multi-section documents, wherein the multi-section document comprises a plurality of sections that are respectively associated with a plurality of document-wide section type identifiers;

identifying, by the one or more processors, a section batch from a current section of the plurality of sections that is associated with a current section type identifier from the plurality of document-wide section type identifiers, wherein:

(i) the section batch comprises a plurality of common-type sections that is associated with the current section type identifier,

(ii) the plurality of common-type sections comprises the current section and one or more other sections that are associated with the current section type identifier across the plurality of other multi-section documents, and

(iii) the current section type identifier corresponds to one or more semantic characteristics of the plurality of common-type sections;

inputting, by the one or more processors, the section batch to an unsupervised section clustering machine learning model to generate a plurality of per-type section clusters for the current section type identifier, wherein:

(i) each per-type section cluster of the plurality of per-type section clusters comprises a related section subset of the plurality of common-type sections, and

(ii) the plurality of per-type section clusters comprises an inferred per-type section cluster for the current section;

generating, by the one or more processors, the inferred document representation for the multi-section document based at least in part on a plurality of inferred per-type section clusters, including the inferred per-type section cluster, respectively generated for the plurality of sections; and

initiating, by the one or more processors, performance of one or more prediction-based actions based at least in part on the inferred document representation.

2. The computer-implemented method of claim 1 , further comprising:

identifying, by the one or more processors, a section schema associated with the document corpus, wherein: (i) the section schema describes a group of corpus-wide section type identifiers for the document corpus, and (ii) the group of corpus-wide section type identifiers comprise the plurality of document-wide section type identifiers.

3. The computer-implemented method of claim 2 , wherein:

the inferred document representation describes a plurality of per-type section cluster identifiers;

a per-type section cluster identifier of the plurality of per-type section cluster identifiers is associated with a corpus-wide section type identifier of the group of corpus-wide section type identifiers;

the per-type section cluster identifier for the corpus-wide section type identifier of the group of corpus-wide section type identifiers that is among the plurality of document-wide section type identifiers describes the inferred per-type section cluster for the current section of the plurality of sections that is associated with the corpus-wide section type identifier; and

the per-type section cluster identifier for the corpus-wide section type identifier of the group of corpus-wide section type identifiers that is not among the plurality of document-wide section type identifiers describes a default numerical value.

4. The computer-implemented method of claim 1 , further comprising:

processing, by the one or more processors, a per-type section cluster of the plurality of per-type section clusters using a document summarization machine learning model to generate a per-type section cluster summary for the per-type section cluster.

5. The computer-implemented method of claim 4 , wherein generating the per-type section cluster summary for the per-type section cluster of the plurality of per-type section clusters comprises:

processing the related section subset for the per-type section cluster using the document summarization machine learning model to generate the per-type section cluster summary.

6. The computer-implemented method of claim 4 , wherein performing the one or more prediction-based actions comprises:

causing presentation of a prediction output user interface, wherein: (i) the prediction output user interface describes a multi-section document summary for the multi-section document, and (ii) the multi-section document summary describes each the per-section type cluster summary for the per-type section cluster of the plurality of per-type section clusters.

7. The computer-implemented method of claim 1 , wherein performing the one or more prediction-based actions comprises:

causing presentation of a prediction output user interface, wherein the prediction output user interface describes the inferred document representation.

8. The computer-implemented method of claim 7 , wherein the prediction output user interface describes a multi-section document summary for the multi-section document.

9. A system for automated generation of an inferred document representation for a multi-section document to convey detailed content granularity of semantic characteristics within the multi-section document, the system comprising memory and one or more processors communicatively coupled to the memory, the one or more processors configured to:

identify a document corpus comprising the multi-section document and a plurality of other multi-section documents, wherein the multi-section document comprises a plurality of sections that are respectively associated with a plurality of document-wide section type identifiers;

identify a section batch from a current section of the plurality of sections that is associated with a current section type identifier from the plurality of document-wide section type identifiers, wherein:

(i) the section batch comprises a plurality of common-type sections that are associated with the current section type identifier,

(ii) the plurality of common-type sections comprises the current section and one or more other sections that are associated with the current section type identifier across the plurality of other multi-section documents, and

(iii) the current section type identifier corresponds to one or more semantic characteristics of the plurality of common-type sections;

input the section batch to an unsupervised section clustering machine learning model to generate a plurality of per-type section clusters for the current section type identifier, wherein:

(i) each per-type section cluster of the plurality of per-type section clusters comprises a related section subset of the plurality of common-type sections, and

(ii) the plurality of per-type section clusters comprises an inferred per-type section cluster for the current section;

generate the inferred document representation for the multi-section document based at least in part on a plurality of inferred per-type section clusters, including the inferred per-type section cluster, respectively generated for the plurality of sections; and

initiate performance of one or more prediction-based actions based at least in part on the inferred document representation.

10. The system of claim 9 , wherein the memory and the one or more processors are configured to:

identify a section schema associated with the document corpus, wherein: (i) the section schema describes a group of corpus-wide section type identifiers for the document corpus, and (ii) the group of corpus-wide section type identifiers comprise the plurality of document-wide section type identifiers.

11. The system of claim 10 , wherein:

the inferred document representation describes a plurality of per-type section cluster identifiers;

a per-type section cluster identifier of the plurality of per-type section cluster identifiers is associated with a corpus-wide section type identifier of the group of corpus-wide section type identifiers;

the per-type section cluster identifier for the corpus-wide section type identifier of the group of corpus-wide section type identifiers that is among the plurality of document-wide section type identifiers describes the inferred per-type section cluster for the current section of the plurality of sections that is associated with the corpus-wide section type identifier; and

the per-type section cluster identifier for the corpus-wide section type identifier of the group of corpus-wide section type identifiers that is not among the plurality of document-wide section type identifiers describes a default numerical value.

12. The system of claim 9 , wherein the one or more processors are further configured to:

process a per-type section cluster of the plurality of per-type section clusters using a document summarization machine learning model to generate a per-type section cluster summary for the per-type section cluster.

13. The system of claim 12 , wherein the one or more processors are further configured to:

process the related section subset for the per-type section cluster using the document summarization machine learning model to generate the per-type section cluster summary.

14. The system of claim 12 , wherein the one or more processors are further configured to:

cause presentation of a prediction output user interface, wherein: (i) the prediction output user interface describes a multi-section document summary for the multi-section document, and (ii) the multi-section document summary describes each per-section type cluster summary for the per-type section cluster of the plurality of per-type section clusters.

15. One or more non-transitory computer-readable storage media for automated generation of an inferred document representation for a multi-section document to convey detailed content granularity of semantic characteristics within the multi-section document, the one or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to:

identify a document corpus comprising the multi-section document and a plurality of other multi-section documents, wherein the multi-section document comprises a plurality of sections that are respectively associated with a plurality of document-wide section type identifiers;

identify a section batch from a current section of the plurality of sections that is associated with a current section type identifier from the plurality of document-wide section type identifiers, wherein:

(i) the section batch comprises a plurality of common-type sections that is associated with the current section type identifier,

(ii) the plurality of common-type sections comprises the current section and one or more other sections that are associated with the current section type identifier across the plurality of other multi-section documents, and

(iii) the current section type identifier corresponds to one or more semantic characteristics of the plurality of common-type sections;

input the section batch to an unsupervised section clustering machine learning model to generate a plurality of per-type section clusters for the current section type identifier, wherein:

(i) each per-type section cluster of the plurality of per-type section clusters comprises a related section subset of the plurality of common-type sections, and

(ii) the plurality of per-type section clusters comprises an inferred per-type section cluster for the current section;

generate the inferred document representation for the multi-section document based at least in part on a plurality of inferred per-type section clusters, including the inferred per-type section cluster, respectively generated for the plurality of sections; and

initiate performance of one or more prediction-based actions based at least in part on the inferred document representation.

16. The one or more non-transitory computer-readable storage media of claim 15 , further including instructions that, when executed by the one or more processors, cause the one or more processors to:

identify a section schema associated with the document corpus, wherein: (i) the section schema describes a group of corpus-wide section type identifiers for the document corpus, and (ii) the group of corpus-wide section type identifiers comprise the plurality of document-wide section type identifiers.

17. The one or more non-transitory computer-readable storage media of claim 16 , wherein:

the inferred document representation describes a plurality of per-type section cluster identifiers;

a per-type section cluster identifier of the plurality of per-type section cluster identifiers is associated with a corpus-wide section type identifier of the group of corpus-wide section type identifiers;

the per-type section cluster identifier for the corpus-wide section type identifier of the group of corpus-wide section type identifiers that is among the plurality of document-wide section type identifiers describes the inferred per-type section cluster for the current section of the plurality of sections that is associated with the corpus-wide section type identifier; and

the per-type section cluster identifier for the corpus-wide section type identifier of the group of corpus-wide section type identifiers that is not among the plurality of document-wide section type identifiers describes a default numerical value.

18. The one or more non-transitory computer-readable storage media of claim 15 , further including instructions that, when executed by the one or more processors, cause the one or more processors to:

process a per-type section cluster of the plurality of per-type section clusters using a document summarization machine learning model to generate a per-type section cluster summary for the per-type section cluster.

19. The one or more non-transitory computer-readable storage media of claim 18 , further including instructions that, when executed by the one or more processors, cause the one or more processors to:

process the related section subset for the per-type section cluster using the document summarization machine learning model to generate the per-type section cluster summary.

20. The one or more non-transitory computer-readable storage media of claim 18 , further including instructions that, when executed by the one or more processors, cause the one or more processors to:

cause presentation of a prediction output user interface, wherein: (i) the prediction output user interface describes a multi-section document summary for the multi-section document, and (ii) the multi-section document summary describes each per-section type cluster summary for the per-type section cluster of the plurality of per-type section clusters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2021
From: MATTIVI, RICCARDO; COGAN, PETER
To: OPTUM SERVICES (IRELAND) LIMITED
Reel/Frame 054816/0722 →
Continuity (1)
Related Publication 20220215274A1 · Jul 7, 2022
References Cited (13)
US 7818308B2 · Carus et al. · 2010 [cited by applicant]
US 10216715B2 · Broderick et al. · 2019 [cited by applicant]
US 10565234B1 · Sims · 2020 [cited by examiner]
US 10747958B2 · Nelson et al. · 2020 [cited by applicant]
US 20030115080A1 · Kasravi et al. · 2003 [cited by applicant]
US 20080109425A1 · Yih · 2008 [cited by examiner]
US 20180137107A1 · Buccapatnam Tirumala et al. · 2018 [cited by applicant]
US 20180365201A1 · Hunn et al. · 2018 [cited by applicant]
US 20190354584A1 · Wallenfelt · 2019 [cited by examiner]
US 20200327151A1 · Coquard · 2020 [cited by examiner]
US 20210365306A1 · Haldar · 2021 [cited by examiner]
WO WO2018006072A1 · 2018 [cited by examiner]
Ash, Elliott et al. “Unsupervised Extraction of Workplace Rights and Duties From Collective Bargaining Agreements,” In 2nd International Workshop on Mining and Learning in the Legal Domain (MLLD-2020)(virtual), Nov. 202… [cited by applicant]
Cited By (1)
US 12,524,605