IP Library Granted Patent US 11,093,687
Granted Patent B2
US 11,093,687 · App. 16/211,123 · Granted Aug 17, 2021

Systems and methods for identifying key phrase clusters within documents

Inventors: Max Kesin (Woodmere, NY); Hem Wadhar (New York, NY)
Assignee: Palantir Technologies Inc.
G06F40/106G06F3/0481G06F16/345G06F16/353G06F40/117G06F40/205
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,093,687
App. No.
16/211,123
Granted
Aug 17, 2021
Kind
B2
Abstract

Systems and methods are disclosed for key phrase clustering of documents. In accordance with one implementation, a method is provided for key phrase clustering of documents. The method includes obtaining a first plurality of documents based at least on a user input, obtaining a statistical model based at least on the user input, and obtaining, from content of the first plurality of documents, a plurality of segments. The method also includes identifying a plurality of clusters of segments from the plurality of segments, determining statistical significance of the plurality of clusters based at least on the statistical model and the content, and providing for display a representative cluster from the plurality of tokens, the representative cluster being determined based at least on the statistical significance. The method further includes determining a label for the representative cluster based at least on the plurality of clusters and the statistical significance.

Claims (44)

1. A computing system comprising:

computer-readable storage media storing instructions; and

one or more processors configured to execute the instructions to cause the computing system to:

obtain documents and a statistical model;

segment contents of the documents into segments;

determine statistical significances of the segments using the statistical model;

select, for each of the documents, a representative segment from the segments based on the determined statistical significances of the segments;

cluster the documents into clusters based at least in part on the representative segment selected for each document;

determine a label for each cluster;

identify a set of clusters, from the clusters, based on a user input, wherein the set of clusters includes a plurality of clusters, wherein the user input comprises at least one of: a string or a date range, and wherein the user input is related to the set of clusters; and

modify a graphical user interface to further include, for each of the clusters of the set of clusters:

an indication of the label associated with the cluster, and

an indication of the documents associated with the cluster;

wherein the clusters of the set of clusters and their respective associated documents are grouped and displayed in separate portions of the graphical user interface.

2. The computing system of claim 1 , wherein the selecting of the representative segment further comprises, for each of the documents, determining a quantity of the segments having the highest statistical significance values from among the segments of the document.

3. The computing system of claim 1 , wherein the documents are clustered into the clusters further based on identical or substantially identical representative segments.

4. The computing system of claim 1 , wherein the one or more processors configured to execute the instructions to cause the computing system to determine identical or substantially identical representatives segments based on at least one of an edit distance or a synonym.

5. The computing system of claim 4 , wherein the one or more processors configured to execute the instructions to cause the computing system to determine the identical or substantially identical representatives segments based on the edit distance that is based on a Levenshtein distance.

6. The computing system of claim 1 , wherein determining the label for each cluster further comprises determining the representative segment appearing most frequently in the cluster.

7. The computing system of claim 1 , wherein determining the label for each cluster further comprises determining a textual phrase different from the representative segments for the cluster.

8. The computing system of claim 7 , wherein the textual phrase is based in part on one or more of the representative segments for the cluster.

9. The computing system of claim 1 , wherein the user input comprises a date or date range.

10. The computing system of claim 1 , wherein the indication of the documents associated with the cluster includes contents of the documents, links to the documents, or a combination thereof.

11. A method performed by one or more processors, the method comprising:

obtaining documents and a statistical model;

segmenting contents of the documents into segments;

determining statistical significances of the segments using the statistical model;

selecting, for each of the documents, a representative segment from the segments based on the determined statistical significances of the segments;

clustering the documents into clusters based at least in part on the representative segment selected for each document;

determining a label for each cluster;

identifying a set of clusters, from the clusters, based on a user input, wherein the set of clusters includes a plurality of clusters, wherein the user input comprises at least one of: a string or a date range, and wherein the user input is related to the set of clusters; and

modifying a graphical user interface to further include, for each of the clusters of the set of clusters:

an indication of the label associated with the cluster, and

an indication of the documents associated with the cluster;

wherein the clusters of the set of clusters and their respective associated documents are grouped and displayed in separate portions of the graphical user interface.

12. The method of claim 11 , wherein the representative segment selected for each document is selected based at least in part on identical or substantially identical representative segments.

13. The method of claim 11 , wherein the representative segment selected for each document is selected based at least in part on synonym relationships between representative segments.

14. The method of claim 11 , wherein the representative segment selected for each document is selected based at least in part on an edit distance.

15. The method of claim 14 , wherein the edit distance is based on a Levenshtein distance.

16. The method of claim 11 , wherein determining the label for each cluster further comprises determining the representative segment appearing most frequently in the cluster.

17. The method of claim 11 , wherein determining the label for each cluster further comprises determining a textual phrase that is different from the representative segments in each cluster.

18. The method of claim 17 , wherein the textual phrase is based in part on one or more of the representative segments for one or more of the clusters.

19. The method of claim 11 , wherein the user input selects a date or date range.

20. The method of claim 11 , wherein the indications of the documents associated with the cluster includes contents of the documents, links to the documents, or a combination thereof.

Assignments (8)
ASSIGNMENT OF INTELLECTUAL PROPERTY SECURITY AGREEMENTS Recorded Jul 3, 2022
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0640 →
SECURITY INTEREST Recorded Jul 3, 2022
From: PALANTIR TECHNOLOGIES INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0506 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ERRONEOUSLY LISTED PATENT BY REMOVING APPLICATION NO. 16/832267 FROM THE RELEASE OF SECURITY INTEREST PREVIOUSLY RECORDED ON REEL 052856 FRAME 0382. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Aug 26, 2021
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 057335/0753 →
SECURITY INTEREST Recorded Jun 4, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 052856/0817 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2020
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 052856/0382 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
Reel/Frame 051713/0149 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 051709/0471 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2019
From: KESIN, MAX; WADHAR, HEM
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 048239/0900 →
Cited By (1)
US 12,436,988