IP Library › Granted Patent US 10,452,781
Granted Patent B2
US 10,452,781 · App. 15/604,428 · Granted Oct 22, 2019

Data provenance system

Inventor: Vineet Verma (Hyderabad, IN)
Assignee: CA, Inc.
G06F17/2785G06F17/2765
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,452,781
App. No.
15/604,428
Granted
Oct 22, 2019
Kind
B2
Abstract

An electronic artifact is accessed which includes content of a particular type of media. Text is determined corresponding to the content and natural language processing is performed on the text to identify at least a subset of words in a statement within the text and determine meanings of each word in the subset of words. A context image is generated for the electronic artifact based on the natural language processing, where the context image includes a graph including nodes corresponding to the subset of words and the context image defines relationships between the subset of words.

Claims (41)

1. A method comprising:

accessing, from an index, an electronic artifact comprising content of a particular type of media;

automatically determining, using a data processor, text corresponding to the content;

performing natural language processing on the text, using the data processor, to identify at least a subset of words in a statement within the text and determine meanings of each word in the subset of words; and

generating a context image for the electronic artifact based on the natural language processing, wherein the context image comprises a graph comprising nodes corresponding to the subset of words, the context image comprises a syntax-free representation of the statement, the context image comprises the subset of words but less than all words in the statement, and the context image defines relationships between the subset of words.

2. The method of claim 1 , wherein the natural language processing comprises determining that a first word in the statement comprises a key term representing a topic of the statement and further comprises determining that at least a second word in the statement comprises an attribute of the key term, wherein the subset of words comprises the first and second words and the context image defines that the second word is an attribute of the first word.

3. The method of claim 1 , further comprising:

determining a plurality of statements in the electronic artifact from the natural language processing; and

generating a plurality of context images for the electronic artifact corresponding to each of the plurality of statements based on the natural language processing.

4. The method of claim 3 , further comprising generating an aggregate context image for the electronic artifact comprising the plurality of context images.

5. The method of claim 1 , wherein the context image comprises a first context image and the method further comprises comparing the first context image with a plurality of other context images corresponding to a plurality of other electronic artifacts in a corpus to determine a degree of similarly between content of the first context image and a second context image in the plurality of other context images, wherein the second context image is associated with a second electronic artifact in the plurality of electronic artifacts.

6. The method of claim 5 , further comprising determining that the second electronic artifact is a source of the statement based on the degree of similarity.

7. The method of claim 6 , further comprising generating an annotation for association with the first electronic artifact to indicate that the second electronic artifact is the source of the statement.

8. The method of claim 5 , wherein the plurality of context images are included in an index of context images, and the method further comprises adding the first context image to the index of context images and defining, within the index of context images, a relationship between the first and second context images based on the degree of similarity between the first and second context images.

9. The method of claim 5 , wherein the second electronic artifact comprises content of a different, second type of media.

10. The method of claim 1 , wherein the particular type comprises a text document.

11. The method of claim 1 , wherein the particular type comprises an image and determining the text comprises determining text present within the image.

12. The method of claim 1 , wherein the particular type comprises a video and determining the text comprises determining one of speech in audio of the video or text included in an image within the video.

13. The method of claim 1 , wherein the particular type comprises audio content and determining the text comprises determining speech in the audio content and converting the speech to text.

14. The method of claim 1 , further comprising:

determining a language of the text; and

translating the subset of words from the language into a common language for use in the context image.

15. A non-transitory computer readable medium having program instructions stored therein, wherein the program instructions are executable by a computer system to perform operations comprising:

identifying digital media of a particular type;

determining text statements from content of the digital media;

performing natural language processing on the text statements to:

identify a first word in a particular one of the text statements as a key term in the particular text statement, wherein the key term represents a topic of the particular text statement; and

identify a set of second words in the particular text statement representing attributes of the topic;

generating a context image for the statement, wherein the context image comprises a graph comprising nodes corresponding to the first word and the set of second words, the context image comprises a syntax-free representation of the statement, the context image comprises the first word and set of second words but less than all words in the statement, and defining relationships between the nodes to indicate that the set of second words represent attributes of the topic represented by the first word; and

determining a similarity score for the particular text statement based on a comparison of the context image with a plurality of other context images generated from other digital media.

16. A system comprising:

a data processing apparatus;

a memory element storing data comprising an electronic artifact;

a text extractor, executable by the data processing apparatus to determine a text statement from content of the electronic artifact;

a natural language processor, executable by the data processing apparatus to assess the text statement to:

determine meanings of a set of words included in the text statement;

identify a first word in the set of words as a key term in the text statement, wherein the key term represents a topic of the text statement; and

identify a set of second words in the text statement representing attributes of the topic; and

a context image generator, executable by the data processing apparatus to generate a context image for the text statement, wherein the context image comprises a graph comprising nodes corresponding to the first word and the set of second words, the context image comprises a syntax-free representation of the text statement, the context image comprises the first word and set of second words but less than all words in the text statement, and defining relationships between the nodes to indicate that the set of second words represent attributes of the topic represented by the first word.

17. The system of claim 16 , further comprising a search tool to identify a set of other context images similar to the context image and determine a relationship between the electronic artifact and a set of other electronic artifacts corresponding to the set of other context images based on similarities between the set of other context images and the context image of the electronic artifact.

18. The system of claim 17 , wherein the search tool is to search a corpus of context images generated at least in part by the context image generator from a plurality of other electronic artifacts.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2017
From: VERMA, VINEET
To: CA, INC.
Reel/Frame 042498/0434 →
Continuity (1)
Related Publication 20180341631A1 · Nov 29, 2018
Cited By (3)
US 12,443,572 US 12,620,252 US 12,718,258