IP Library Granted Patent US 9,535,974
Granted Patent B1
US 9,535,974 · App. 14/581,920 · Granted Jan 3, 2017

Systems and methods for identifying key phrase clusters within documents

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,535,974
App. No.
14/581,920
Granted
Jan 3, 2017
Kind
B1
Abstract

Systems and methods are disclosed for key phrase clustering of documents. In accordance with one implementation, a method is provided for key phrase clustering of documents. The method includes obtaining a first plurality of documents based at least on a user input, obtaining a statistical model based at least on the user input, and obtaining, from content of the first plurality of documents, a plurality of segments. The method also includes identifying a plurality of clusters of segments from the plurality of segments, determining statistical significance of the plurality of clusters based at least on the statistical model and the content, and providing for display a representative cluster from the plurality of tokens, the representative cluster being determined based at least on the statistical significance. The method further includes determining a label for the representative cluster based at least on the plurality of clusters and the statistical significance.

Claims (60)

1. An electronic device comprising:

a computer display;

one or more computer-readable storage media configured to store instructions; and

one or more processors configured to execute the instructions to cause the electronic device to at least:

obtain a first plurality of documents based at least in part on a user input;

obtain a statistical model based at least in part on the user input;

obtain, from content of the first plurality of documents, a plurality of segments;

determine statistical significance for the obtained plurality of segments based at least in part on the obtained statistical model;

determine, for each document in the first plurality of documents, representative segments from the obtained plurality of segments, the representative segments being determined based at least in part on the determined statistical significance;

cluster documents from the obtained first plurality of documents based at least in part on the determined representative segments;

receive a selection of a date range;

for a cluster of documents associated with a date within the date range, automatically associate a label with the cluster of documents based at least in party on the determined representative segments; and

display within a graphical user interface on the computer display a representation of the date range, the label, and contents of and/or links to documents in the cluster of documents.

2. The electronic device of claim 1 , wherein the user input identifies an entity, and the obtained statistical model was generated based on at least one of:

the first plurality of documents;

a second plurality of documents associated with the entity; and

a third plurality of documents associated with an industry associated with the entity.

3. The electronic device of claim 1 , wherein determining the label is automatically associated with the cluster of documents based at least in part on a frequency of appearances of the representative segments in the first plurality of documents.

4. The electronic device of claim 1 , wherein the one or more processors are further configured to execute instructions to cause the electronic device to:

receive a selection input associated with the cluster of documents; and

responsive to the selection input, provide for display contents of one or more documents within the cluster of documents.

5. A method performed by one or more processors, the method comprising:

obtaining a first plurality of documents based on at least a user input;

obtaining a statistical model based at least on the user input;

obtaining, from content of the first plurality of documents, a plurality of segments;

determining statistical significance for the obtained plurality of segments based at least on the obtained statistical model;

determining representative segments from the obtained plurality of segments for each document in the first plurality of documents, the representative segments being determined based at least in part on the determined statistical significance;

clustering documents from the obtained first plurality of documents based at least in part on the determined representative segments;

receiving a selection of a date range;

for a cluster of documents associated with a date within the date range, automatically associating a label with the cluster of documents based at least in party on the determined representative segments; and

providing for display within a graphical user interface a representation of the date range, the label, and at least one of contents of and links to documents in the cluster of documents.

6. The method of claim 5 , wherein the user input identifies an entity, and the obtained statistical model was generated based on at least one of:

the first plurality of documents;

a second plurality of documents associated with the entity; and

a third plurality of documents associated with an industry associated with the entity.

7. The method of claim 5 , wherein the automatically associating the label with the cluster of documents is further based on a frequency of appearances of the representative segments in the first plurality of documents.

8. The method of claim 5 further comprising:

receiving a selection input associated with the cluster of documents; and

responsive to the selection input, providing for display contents of one or more documents within the cluster of documents.

9. A non-transitory computer-readable medium storing a set of instructions that are executable by one or more electronic devices, each having one or more processors, to cause the one or more electronic devices to perform a method, the method comprising:

obtaining a first plurality of documents associated with a user input;

obtaining a statistical model associated with the user input;

obtaining, from content of the first plurality of documents, a plurality of segments;

determining statistical significance for the plurality of segments based at least on the statistical model;

determining, for each document in the first plurality of documents, representative segments from the plurality of segments, the representative segments being determined based at least in part on the statistical significance;

clustering documents from the first plurality of documents based at least in part on the representative segments;

receiving a selection of a date range;

for a cluster of documents associated with a date within the date range, automatically associating a label with the cluster of documents based at least in part on the determined representative segments; and

providing for display within a graphical user interface a representation of

the date range,

the label, and

contents of documents in the cluster of documents, or links to documents in the cluster of documents, or a combination thereof.

10. The non-transitory computer-readable medium of claim 9 , wherein the user input identifies an entity, and the statistical model was generated based on at least one of:

the first plurality of documents;

a second plurality of documents associated with the entity; and

a third plurality of documents associated with an industry associated with the entity.

11. The non-transitory computer-readable medium of claim 9 , wherein the automatically associating the label with the cluster of documents is further based on a frequency of appearances of the representative segments in the first plurality of documents.

12. The non-transitory computer-readable medium of claim 9 , the method further comprising:

receiving a selection input associated with the cluster of documents; and

responsive to the selection input, providing for display contents of one or more documents within the cluster of documents.

Assignments (8)
ASSIGNMENT OF INTELLECTUAL PROPERTY SECURITY AGREEMENTS Recorded Jul 3, 2022
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0640 →
SECURITY INTEREST Recorded Jul 3, 2022
From: PALANTIR TECHNOLOGIES INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0506 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ERRONEOUSLY LISTED PATENT BY REMOVING APPLICATION NO. 16/832267 FROM THE RELEASE OF SECURITY INTEREST PREVIOUSLY RECORDED ON REEL 052856 FRAME 0382. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Aug 26, 2021
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 057335/0753 →
SECURITY INTEREST Recorded Jun 4, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 052856/0817 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2020
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 052856/0382 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
Reel/Frame 051713/0149 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 051709/0471 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2015
From: KESIN, MAX; WADHAR, HEM
To: PALANTIR TECHNOLOGIES, INC.
Reel/Frame 034636/0298 →