IP Library Granted Patent US 11,392,550
Granted Patent B2
US 11,392,550 · App. 16/548,803 · Granted Jul 19, 2022

System and method for investigating large amounts of data

Inventors: Geoffrey Stowe (San Francisco, CA); Chris Fischer (Somerville, MA); Paul George (New York, NY); Eli Bingham (New York, NY); Rosco Hill (Palo Alto, CA)
Assignee: PALANTIR TECHNOLOGIES INC.
G06F16/1744G06F11/2025G06F16/10G06F16/13G06F16/148G06F16/17G06F16/2365G06F16/248G06F16/24575G06F16/258G06F16/35G06F16/902G06F16/9535G06F17/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,392,550
App. No.
16/548,803
Granted
Jul 19, 2022
Kind
B2
Abstract

A data analysis system is proposed for providing fine-grained low latency access to high volume input data from possibly multiple heterogeneous input data sources. The input data is parsed, optionally transformed, indexed, and stored in a horizontally-scalable key-value data repository where it may be accessed using low latency searches. The input data may be compressed into blocks before being stored to minimize storage requirements. The results of searches present input data in its original form. The input data may include access logs, call data records (CDRs), e-mail messages, etc. The system allows a data analyst to efficiently identify information of interest in a very large dynamic data set up to multiple petabytes in size. Once information of interest has been identified, that subset of the large data set can be imported into a dedicated or specialized data analysis system for an additional in-depth investigation and contextual analysis.

Claims (42)

1. A computer-implemented method comprising:

receiving a stream of input data;

parsing the input data to identify boundaries of logical data entities in the stream of input data;

generating a block item comprising a key-value family identifier, a data block identifier, and a data block based on a logical data entity of the logical data entities, comprising:

compressing the data block; and

storing the key-value family identifier in association with a key-value pair comprising the data block identifier as a key of the key-value pair and the compressed data block as a value of the key-value pair;

creating and storing a parse item comprising the data block identifier and one or more parse tokens for indexing block items.

2. The computer-implemented method of claim 1 , further comprising:

receiving a search criterion;

using the parse item, determining that the one or more parse tokens match the search criterion;

using the data block identifier as a key to the key-value pair, identifying the compressed data block;

uncompressing the data block;

using the search criterion to identify one or more portions of the uncompressed data block; and

returning the one or more portions of the uncompressed data block as search results.

3. The computer-implemented method of claim 1 , further comprising storing a plurality of key-value pairs for a plurality of compressed data blocks, wherein each key-value pair of the plurality of key-value pairs is unique at least amongst all key-value pairs of the plurality of key-value pairs.

4. The computer-implemented method of claim 3 , wherein the plurality of key-value pairs comprises at least one million unique keys.

5. The computer-implemented method of claim 1 , wherein creating the parse item comprises extracting the one or more parse tokens from one or more of the logical data entities.

6. The computer-implemented method of claim 1 , wherein the parse item further comprises snippet identifying information.

7. The computer-implemented method of claim 6 , wherein the snippet identifying information is a byte offset into an uncompressed block and a byte length.

8. The computer-implemented method of claim 1 , wherein the parse tokens identify byte-sequential portions of the data block.

9. A system comprising:

one or more processors;

a memory storing instructions which, when executed by the one or more processors, causes performing:

receiving a stream of input data;

parsing the input data to identify boundaries of logical data entities in the stream of input data;

generating a block item comprising a key-value family identifier, a data block identifier, and a data block based on a logical data entity of the logical data entities, comprising:

compressing the data block; and

storing the key-value family identifier in association with a key-value pair comprising the data block identifier as a key of the key-value pair and the compressed data block as a value of the key-value pair;

creating and storing a parse item comprising the data block identifier and one or more parse tokens for indexing block items.

10. The system of claim 9 , wherein the instructions, when executed by the one or more processors, further cause performing:

receiving a search criterion;

using the parse item, determining that the one or more parse tokens match the search criterion;

using the data block identifier as a key to the key-value pair, identifying the compressed data block;

uncompressing the data block;

using the search criterion to identify one or more portions of the uncompressed data block; and

returning the one or more portions of the uncompressed data block as search results.

11. The system of claim 9 , wherein the instructions, when executed by the one or more processors, further cause performing storing a plurality of key-value pairs for a plurality of compressed data blocks, wherein each key-value pair of the plurality of key-value pairs is unique at least amongst all key-value pairs of the plurality of key-value pairs.

12. The system of claim 11 , wherein the plurality of key-value pairs comprises at least one million unique keys.

13. The system of claim 9 , wherein creating the parse item comprises extracting the one or more parse tokens from one or more of the logical data entities.

14. The system of claim 9 , wherein the parse item further comprises snippet identifying information.

15. The system of claim 14 , wherein the snippet identifying information is a byte offset into an uncompressed block and a byte length.

16. The system of claim 9 , wherein the parse tokens identify byte-sequential portions of the data block.

Assignments (7)
SECURITY INTEREST Recorded Jul 3, 2022
From: PALANTIR TECHNOLOGIES INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0506 →
ASSIGNMENT OF INTELLECTUAL PROPERTY SECURITY AGREEMENTS Recorded Jul 3, 2022
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0640 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ERRONEOUSLY LISTED PATENT BY REMOVING APPLICATION NO. 16/832267 FROM THE RELEASE OF SECURITY INTEREST PREVIOUSLY RECORDED ON REEL 052856 FRAME 0382. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Aug 26, 2021
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 057335/0753 →
SECURITY INTEREST Recorded Jun 4, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 052856/0817 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2020
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 052856/0382 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 051709/0471 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
Reel/Frame 051713/0149 →
Continuity (6)
Continuation 15824096 · Nov 28, 2017
Continuation 15446917 · Mar 1, 2017
Continuation 14961830 · Dec 7, 2015
Continuation 14451221 · Aug 4, 2014
Continuation 13167680 · Jun 23, 2011
Related Publication 20190384747A1 · Dec 19, 2019
Cited By (1)
US 12,445,517