IP Library Granted Patent US 10,162,887
Granted Patent B2
US 10,162,887 · App. 15/483,731 · Granted Dec 25, 2018

Systems and methods for key phrase characterization of documents

Inventors: Max Kesin (Woodmere, NY); Hem Wadhar (New York, NY)
Assignee: PALANTIR TECHNOLOGIES INC.
G06F17/30719
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,162,887
App. No.
15/483,731
Granted
Dec 25, 2018
Kind
B2
Abstract

Systems and methods are disclosed for key phrase characterization of documents. In accordance with one implementation, a method is provided for key phrase characterization of documents. The method includes obtaining a first plurality of documents based at least on a user input, obtaining a statistical model based at least on the user input, and obtaining, from content of the first plurality of documents, a plurality of segments. The method also includes determining statistical significance of the plurality of segments based at least on the statistical model and the content, and providing for display a representative segment from the plurality of segments, the representative segment being determined based at least on the statistical significance.

Claims (42)

1. An electronic device comprising:

one or more computer-readable storage media configured to store instructions; and

one or more processors configured to execute the instructions to cause the electronic device to:

receive a user input indicative of an entity and a search query;

identify a statistical model associated with the entity, wherein the statistical model is determined based on a first plurality of documents associated with the entity, the statistical model indicative at least of frequencies of one or more words within the first plurality of documents;

identify, responsive to the user input, a second plurality of documents, at least partially different than the first plurality of documents, corresponding to the search query and the indicated entity;

identify, for each of the second plurality of documents, one or more segments;

apply the identified statistical model to each of the identified segments to determine, for each of the second plurality of documents, a statistical significance of segments identified in the documents, the statistical significance indicative of frequencies of the one or more words in the segment compared to the frequencies of the one or more words indicated in the statistical model;

and provide for display at least a representative segment having a highest statistical significance and a link to the document containing the representative segment.

2. The electronic device of claim 1 , wherein the second plurality of documents is a subset of the first plurality of documents.

3. The electronic device of claim 1 , wherein the entity identifies a company, and the first plurality of documents are associated with an industry associated with the company.

4. The electronic device of claim 1 , wherein the user input comprises user's selection of an area on a stock chart, and wherein the first plurality of documents have publication dates corresponding to a date range that corresponds to the area on the stock chart.

5. The electronic device of claim 1 , wherein providing the representative segment for display comprises providing for display a phrase from the representative segment.

6. The electronic device of claim 1 , wherein the one or more processors are further configured to execute the instructions to cause the electronic device to:

after providing the representative segment for display, receive a selection input associated with the representative segment; and

responsive to the selection input, provide for display the contents of one or more of the second plurality of documents.

7. The electronic device of claim 1 , wherein statistical significance of segments is further based on a comparison of the statistical model and content of the second plurality of documents.

8. A method comprising:

by a computing system comprising a hardware computer processor and non-transitory storage medium storing software instructions,

receive a user input indicative of an entity and a search query;

identify a statistical model associated with the entity, wherein the statistical model is determined based on a first plurality of documents associated with the entity, the statistical model indicative at least of frequencies of one or more words within the first plurality of documents;

identify a second plurality of documents, at least partially different than the first plurality of documents, corresponding to the search query and the indicated entity;

identify, for each of the second plurality of documents, one or more segments;

apply the identified statistical model to each of the identified segments to determine, for each of the second plurality of documents, a statistical significance of segments identified in the documents, the statistical significance indicative of frequencies of the one or more words in the segment compared to the frequencies of the one or more words indicated in the statistical model; and

provide for display at least a representative segment having a highest statistical significance and a link to the document containing the representative segment.

9. The method of claim 8 , wherein the second plurality of documents is a subset of the first plurality of documents.

10. The method of claim 8 , wherein the entity identifies a company, and the first plurality of documents are associated with an industry associated with the company.

11. The method of claim 8 , further comprising:

after providing the representative segment for display, receiving a selection input associated with the representative segment; and

responsive to the selection input, providing for display the contents of one or more of the second plurality of documents.

12. The method of claim 8 , wherein statistical significance of segments is further based on comparing the statistical model and content of the second plurality of documents.

13. A non-transitory computer-readable medium storing a set of instructions that are executable by one or more electronic devices, each having one or more processors, to cause the one or more electronic devices to perform a method, the method comprising:

receiving a user input indicative of an entity and a search query;

identifying a statistical model associated with the entity, wherein the statistical model is determined based on a first plurality of documents associated with the entity, the statistical model indicative at least of frequencies of one or more words within the first plurality of documents;

identifying a second plurality of documents, at least partially different than the first plurality of documents, corresponding to the search query and the indicated entity;

identifying, for each of the second plurality of documents, one or more segments;

applying the identified statistical model to each of the identified segments to determine, for each of the second plurality of documents, a statistical significance of segments identified in the documents, the statistical significance indicative of frequencies of the one or more words in the segment compared to the frequencies of the one or more words indicated in the statistical model; and

providing for display at least a representative segment having a highest statistical significance and a link to the document containing the representative segment.

14. The non-transitory computer-readable medium of claim 13 , wherein the second plurality of documents is a subset of the first plurality of documents.

15. The non-transitory computer-readable medium of claim 13 , wherein the entity identifies a company, and the first plurality of documents are associated with an industry associated with the company.

16. The non-transitory computer-readable medium of claim 13 , wherein the user input comprises user's selection of an area on a stock chart, and wherein the first plurality of documents have publication dates corresponding to a date range that corresponds to the area on the stock chart.

17. The non-transitory computer-readable medium of claim 13 , wherein providing the representative segment for display comprises providing for display a phrase from the representative segment.

Assignments (8)
ASSIGNMENT OF INTELLECTUAL PROPERTY SECURITY AGREEMENTS Recorded Jul 3, 2022
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0640 →
SECURITY INTEREST Recorded Jul 3, 2022
From: PALANTIR TECHNOLOGIES INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0506 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ERRONEOUSLY LISTED PATENT BY REMOVING APPLICATION NO. 16/832267 FROM THE RELEASE OF SECURITY INTEREST PREVIOUSLY RECORDED ON REEL 052856 FRAME 0382. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Aug 26, 2021
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 057335/0753 →
SECURITY INTEREST Recorded Jun 4, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 052856/0817 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2020
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 052856/0382 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
Reel/Frame 051713/0149 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 051709/0471 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2017
From: KESIN, MAX; WADHAR, HEM
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 041950/0122 →
Continuity (2)
Continuation 14319765 · Jun 30, 2014
Related Publication 20170277780A1 · Sep 28, 2017