IP Library Granted Patent US 10,318,630
Granted Patent B1
US 10,318,630 · App. 15/678,874 · Granted Jun 11, 2019

Analysis of large bodies of textual data

Inventors: Maxim Kesin (Woodmere, NY); Paul Gribelyuk (Jersey City, NJ)
Assignee: Palantir Technologies Inc.
G06F17/2715G06F17/2241G06F3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,318,630
App. No.
15/678,874
Granted
Jun 11, 2019
Kind
B1
Abstract

In various example embodiments, a textual identification system is configured to receive a set of search terms and identify a set of textual data based on the search terms. The textual identification system retrieves a data structure including textual identifications for the set of textual data and processes the data structure to generate a modified data structure. The textual identification system sums rows within the modified data structure and identifies one or more elements of interest. The textual identification system then causes presentation of the elements of interest in a first portion of a graphical user interface and the textual identifications for the set of textual data in a second portion of the graphical user interface.

Claims (71)

1. A method, comprising:

receiving one or more search terms within a graphical user interface;

identifying, by one or more processors of a machine, a set of textual data based on the one or more search terms;

retrieving a data structure including textual identifications for the set of textual data and an indication of one or more data elements within one or more text sets of the set of textual data;

processing, by the one or more processors, the data structure to generate a modified data structure, the modified data structure generated by reducing to text sets included in the set of textual data identified based on the one or more search terms;

summing rows, by the one or more processors, within the modified data structure, the rows including values for data elements included in each of the identified set of textual data;

identifying, by the one or more processors, one or more elements of interest within the set of textual data based on the summed rows of the modified data structure;

determining, by the one or more processors, a context of occurrence for each element of interest;

normalizing, by the one or more processors, the elements of interest by removing redundant elements of interest based on the context of occurrence of two or more elements of interest and generating a normalized set of elements of interest; and

causing presentation of the normalized set of elements of interest in a first portion of the graphical user interface and the textual identifications for the set of textual data in a second portion of the graphical user interface.

2. The method of claim 1 , wherein causing presentation of the elements of interest and the textual identifications further comprises:

causing presentation of at least a portion of a text set of the set of textual data in a third portion of the graphical user interface.

3. The method of claim 1 , wherein causing presentation of the elements of interest further comprises:

identifying an element type for each of the elements of interest; and

causing presentation of a visual indicator differentiating the elements of interest based on an element type.

4. The method of claim 1 , further comprising:

in response to determining the context of occurrence for each element of interest, generating a set of tokens for each element of interest, the set of tokens representing the context of occurrence;

identifying an overlap of two or more elements of interest based on the set of tokens for the two or more elements of interest; and

linking two or more elements of interest.

5. The method of claim 1 further comprising:

receiving a selection of a textual corpus from a set of textual corpora, each textual corpus of the set of textual corpora containing one or more text sets, the set of textual corpora identified from the selected textual corpus.

6. The method of claim 1 , wherein identifying the set of textual data further comprises:

accessing a set of textual corpora, each textual corpus of the set of textual corpora containing one or more text sets; and

dynamically partitioning the set of textual corpora to identify a textual corpus from the set of textual corpora containing the set of text sets associated with the one or more search terms.

7. A computer implemented system, comprising:

one or more processors; and

a processor-readable storage device comprising processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

receiving one or more search terms within a graphical user interface;

identifying a set of textual data based on the one or more search terms;

retrieving a data structure including textual identifications for the set of textual data and an indication of one or more data elements within one or more text sets of the set of textual data;

processing the data structure to generate a modified data structure, the modified data structure generated by reducing to text sets included in the set of textual data identified based on the one or more search terms;

summing rows within the modified data structure, the rows including values for data elements included in each of the identified set of textual data;

identifying one or more elements of interest within the set of textual data based on the summed rows of the modified data structure;

determining a context of occurrence for each element of interest;

normalizing the elements of interest by removing redundant elements of interest based on the context of occurrence of two or more elements of interest and generating a normalized set of elements of interest; and

causing presentation of the normalized set of elements of interest in a first portion of the graphical user interface and the textual identifications for the set of textual data in a second portion of the graphical user interface.

8. The system of claim 7 , wherein causing presentation of the elements of interest and the textual identifications further comprises:

causing presentation of at least a portion of a text set of the set of textual data in a third portion of the graphical user interface.

9. The system of claim 7 , wherein causing presentation of the elements of interest further comprises:

identifying an element type for each of the elements of interest; and

causing presentation of a visual indicator differentiating the elements of interest based on an element type.

10. The system of claim 7 , wherein the operations further comprise:

in response to determining the context of occurrence for each element of interest, generating a set of tokens for each element of interest, the set of tokens representing the context of occurrence;

identifying an overlap of two or more elements of interest based on the set of tokens for the two or more elements of interest; and

linking two or more elements of interest.

11. The system of claim 7 , wherein the operations further comprise:

receiving a selection of a textual corpus from a set of textual corpora, each textual corpus of the set of textual corpora containing one or more text sets, the set of textual corpora identified from the selected textual corpus.

12. The system of claim 7 , wherein identifying the set of textual data further comprises:

accessing a set of textual corpora, each textual corpus of the set of textual corpora containing one or more text sets; and

dynamically partitioning the set of textual corpora to identify a textual corpus from the set of textual corpora containing the set of text sets associated with the one or more search terms.

13. A processor-readable storage device comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

receiving one or more search terms within a graphical user interface;

identifying a set of textual data based on the one or more search terms;

retrieving a data structure including textual identifications for the set of textual data and an indication of one or more data elements within one or more text sets of the set of textual data;

processing the data structure to generate a modified data structure, the modified data structure generated by reducing to text sets included in the set of textual data identified based on the one or more search terms;

summing rows within the modified data structure, the rows including values for data elements included in each of the identified set of textual data;

identifying one or more elements of interest within the set of textual data based on the summed rows of the modified data structure;

determining a context of occurrence for each element of interest normalizing the elements of interest by removing redundant elements of interest based on the context of occurrence of two or more elements of interest and generating a normalized set of elements of interest; and

causing presentation of the normalized set of elements of interest in a first portion of the graphical user interface and the textual identifications for the set of textual data in a second portion of the graphical user interface.

14. The processor-readable storage device of claim 13 , wherein causing presentation of the elements of interest and the textual identifications further comprises:

causing presentation of at least a portion of a text set of the set of textual data in a third portion of the graphical user interface.

15. The processor-readable storage device of claim 13 , wherein causing presentation of the elements of interest further comprises:

identifying an element type for each of the elements of interest; and

causing presentation of a visual indicator differentiating the elements of interest based on an element type.

16. The processor-readable storage device of claim 13 , wherein the operations further comprise:

in response to determining the context of occurrence for each element of interest, generating a set of tokens for each element of interest, the set of tokens representing the context of occurrence;

identifying an overlap of two or more elements of interest based on the set of tokens for the two or more elements of interest; and

linking two or more elements of interest.

17. The processor-readable storage device of claim 13 , wherein identifying the set of textual data further comprises:

accessing a set of textual corpora, each textual corpus of the set of textual corpora containing one or more text sets; and

dynamically partitioning the set of textual corpora to identify a textual corpus from the set of textual corpora containing the set of text sets associated with the one or more search terms.

Assignments (8)
ASSIGNMENT OF INTELLECTUAL PROPERTY SECURITY AGREEMENTS Recorded Jul 3, 2022
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0640 →
SECURITY INTEREST Recorded Jul 3, 2022
From: PALANTIR TECHNOLOGIES INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0506 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ERRONEOUSLY LISTED PATENT BY REMOVING APPLICATION NO. 16/832267 FROM THE RELEASE OF SECURITY INTEREST PREVIOUSLY RECORDED ON REEL 052856 FRAME 0382. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Aug 26, 2021
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 057335/0753 →
SECURITY INTEREST Recorded Jun 4, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 052856/0817 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2020
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 052856/0382 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
Reel/Frame 051713/0149 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 051709/0471 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2018
From: KESIN, MAXIM; GRIBELYUK, PAUL
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 045120/0569 →
Continuity (1)
Provisional Application 62424844 · Nov 21, 2016
Cited By (1)
US 12,306,857