IP Library Granted Patent US 10,540,439
Granted Patent B2
US 10,540,439 · App. 15/488,708 · Granted Jan 21, 2020

Systems and methods for identifying evidentiary information

Inventors: Mahmoud Azmi Khamis (Burbank, IL); Bruce Golden (Chicago, IL); Rami Ikhreishi (Raleigh, NC)
Assignee: MARCA RESEARCH & DEVELOPMENT INTERNATIONAL, LLC
G06F17/271G06F17/277G06F17/2735G06F17/2775G06F17/2785
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,540,439
App. No.
15/488,708
Granted
Jan 21, 2020
Kind
B2
Abstract

Systems and methods for semantically analyzing digital information. A cognitive engine is configured to determine useful evidentiary information from large digital content data sets. Further, the cognitive engine can analyze or manipulate the evidentiary information to derive data needed to solve problems, identify issues, and identify patterns. The results can then be applied to any application, interface, or automation as appropriate.

Claims (82)

1. A system for analyzing digital content available via a networked resource, the system comprising:

a cognitive engine including a processor and an operably coupled memory, the memory comprising instructions that, when executed, causes the processor to implement:

a pre-reading of the digital content to determine a size of the digital content, a number of documents in the digital content, and an amount of processing needed to completely pre-process the digital content by

selecting a subset of the digital content,

processing the subset of the digital content using a subset of cognitive engine resources, the cognitive engine resources including a cogent information engine, a concept extraction engine, and an entity extraction engine, and

determining a benchmark amount of time to process the subset of the digital content;

a pre-processing of a full set of the digital content by selectively loading a plurality of cognitive engine instances, wherein the number of the plurality of cognitive engine instances that are loaded is based on a scaling of the benchmark amount of time and the size of the subset of the digital content from the pre-reading;

wherein the cogent information engine is configured to:

parse a document from the digital content,

identify a part of speech for every word in the document,

identify a subject word for all reference words in the document,

generate a parsing tree for the document to determine a sentence structure for every sentence in the document based on the parts of speech and the subject words,

determine a sentence meaning for every sentence in the document based on the sentence structure and a plurality of grammatical tests,

determine a weighting for each sentence wherein sentences having similar sentence meanings have similar weightings; and

output a subset of the sentences based on the weighting as cogent information of the document;

the concept extraction engine configured to determine noun phrases in the cogent information based on the part of speech identification, wherein each noun phrase is a digital content concept;

the entity extraction engine configured to:

identify a plurality of entities for each of the identified digital content concepts and one or more relations between the plurality of entities based on the part of speech identification, and

classify the plurality of entities; and

a pattern recognition engine configured to

determine a difference in relations between entities of the identified digital content concepts,

generate an output of the differences relative to a time marker, and

provide a conclusory action for the digital content.

2. The system for analyzing digital content of claim 1 , wherein the cognitive engine further comprises a question answering engine configured to:

receive an input question;

determine at least one noun phrase in the input question, wherein each noun phrase is an input concept;

perform an inference between the input concept and the digital content concepts; and

output one of the sentences as an answer to the input question based on the inference.

3. The system for analyzing digital content of claim 2 , wherein the question answering engine is further configured to output a citation to the outputted sentence relative to the digital content.

4. The system for analyzing digital content of claim 1 , wherein the digital content comprises a first document and a second document, and the cognitive engine further comprises a document comparison engine configured to:

receive a first sentence of the first document;

receive a second sentence of the second document;

determine at least one overlapping concept, overlapping entity, or overlapping relation between the first sentence and the second sentence;

perform an inference on the at least one overlapping concept, overlapping entity, or overlapping relation; and

output a difference based on the inference.

5. The system for analyzing digital content of claim 4 , wherein the document comparison engine is further configured to search a second networked resource for information related to a digital content concept.

6. The system for analyzing digital content of claim 5 , wherein the document comparison engine is further configured to identify information missing from the first document and the second document relative to the second networked resource.

7. The system for analyzing digital content of claim 1 , wherein the cogent information engine is configured to iteratively process all documents in the digital content.

8. The system for analyzing digital content of claim 1 , wherein determining the weighting for each sentence comprises using an inference to identify sentences having a most complete meaning according to a completeness value.

9. The system for analyzing digital content of claim 1 , wherein the one or more relations is an action between two entities.

10. The system for analyzing digital content of claim 1 , wherein the time marker is determined relative to a digital content document timestamp.

11. The system for analyzing digital content of claim 1 , wherein determining the amount of processing needed to analyze the digital content comprises:

conducting a benchmark processing for a subset of the digital content;

recording a time duration for the benchmark processing; and

determining a number of cognitive engine instances to launch based on the time duration.

12. The system for analyzing digital content of claim 1 , wherein determining the amount of processing needed to analyze the digital content comprises approximating a number of sentences in the digital content based on the size of the digital content.

13. A method for analyzing digital content available via a networked resource with a cognitive engine including a processor and an operably coupled memory, the method comprising:

pre-reading the digital content with the processor to determine a size of the digital content, a number of documents in the digital content, and an amount of processing needed to completely pre-process the digital content by

selecting a subset of the digital content,

processing the subset of the digital content using a subset of cognitive engine resources, the cognitive engine resources including a cogent information engine, a concept extraction engine, and an entity extraction engine, and

determining a benchmark amount of time to process the subset of the digital content;

pre-processing of a full set of the digital content by selectively loading a plurality of cognitive engine instances, wherein the number of the plurality of cognitive engine instances that are loaded is based on a scaling of the benchmark amount of time and the size of the subset of the digital content from the pre-reading;

wherein the cogent information engine is configured to:

read a document from the digital content;

identify a part of speech for every word in the document;

identify a subject word for all reference words in the document;

generate a parsing tree for the document to determine a sentence structure for every sentence in the document based on the parts of speech and the subject words;

determine a sentence meaning for every sentence in the document based on the sentence structure and a plurality of grammatical tests;

determine a weighting for each sentence, wherein sentences having similar sentence meanings have similar weightings;

output a subset of the sentences based on the weighting as cogent information of the document;

determine noun phrases in the cogent information based on the part of speech identification, wherein each noun phrase is a digital content concept;

identify a plurality of entities for each of the identified digital content concepts and one or more relations between the plurality of entities based on the part of speech identification;

classify the plurality of entities;

determine a difference in relations between entities;

generate an output of differences relative to a time marker; and

provide a conclusory action for the digital content.

14. The method for analyzing digital content of claim 13 , further comprising:

receiving, with the processor, an input question;

determining, with the processor, at least one noun phrase in the input question, wherein each noun phrase is an input concept;

performing, with the processor, an inference between the input concept and the digital content concepts; and

outputting, with the processor, one of the sentences as an answer to the input question based on the inference.

15. The method for analyzing digital content of claim 14 , further comprising outputting, with the processor, a citation to the outputted sentence relative to the digital content.

16. The method for analyzing digital content of claim 13 , wherein the digital content comprises a first document and a second document, and the method further comprises:

receiving, with the processor, a first sentence of the first document;

receiving, with the processor, a second sentence of the second document;

determining, with the processor, at least one overlapping concept, overlapping entity, or overlapping relation between the first sentence and the second sentence;

performing, with the processor, an inference on the at least one overlapping concept, overlapping entity, or overlapping relation; and

outputting, with the processor, a difference based on the inference.

17. The method for analyzing digital content of claim 13 , wherein the one or more relations are indexed based on at least one of the one or more relation, a related object, or a timestamp of the one or more relation.

18. The system for analyzing digital content of claim 1 , wherein the pre-reading of the digital content generates at least one reusable output used by the pre-processing.

19. The system for analyzing digital content of claim 18 , wherein the at least one reusable output is extracted cogent information, an extracted concept, an extracted entity or object, or an extracted entity relation.

20. The system for analyzing digital content of claim 12 , wherein the documents in the digital content are concatenated into a chunk of data, and wherein selectively loading the plurality of cognitive engine instances includes selecting a subset of the chunk of data from the digital content to be processed by each instance based on the number of sentences.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2017
From: KHAMIS, MAHMOUD AZMI; GOLDEN, BRUCE; IKHREISHI, RAMI
To: MARCA RESEARCH & DEVELOPMENT INTERNATIONAL, LLC
Reel/Frame 042506/0820 →
Continuity (2)
Provisional Application 62323118 · Apr 15, 2016
Related Publication 20170300470A1 · Oct 19, 2017