IP Library Granted Patent US 8,793,264
Granted Patent B2
US 8,793,264 · App. 11/879,992 · Granted Jul 29, 2014

Determining a subset of documents from which a particular document was derived

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,793,264
App. No.
11/879,992
Granted
Jul 29, 2014
Kind
B2
Abstract

Embodiments of the present invention pertain to determining a subset of documents from which a particular document was derived. According to one embodiment, similarity measurements indicating similarities between contents of documents are received. A subset of the documents that the particular document was derived from is determined based on dates the documents were created and the similarity measurements without requiring document tracking information to be associated with the documents to determine the subset.

Claims (42)

1. A method of determining a subset of documents from which a particular document was derived, the method comprising:

receiving similarity measurements indicating similarities between contents of documents; and

determining the subset of the documents that the particular document was derived from based on dates the documents were created and the similarity measurements without requiring document tracking information to be associated with the documents to determine the subset.

2. The method as recited by claim 1 , further comprising:

ranking documents associated with the subset based on how similar each document in the subset is to the particular document.

3. The method as recited by claim 1 , wherein the determining of the subset of the documents further comprises:

creating a forward-in-time graph based on the similarity measurements and the dates the documents were created.

4. The method as recited by claim 3 , wherein the creating of the forward-in-time graph further comprises:

associating arrows that point from older documents to newer documents for each document pair associated with the forward-in-time graph.

5. The method as recited by claim 3 , wherein the creating of the forward-in-time graph further comprises:

creating the forward-in-time graph that has path lengths from the particular document to other documents that are less than or equal to a user specified maximum path length value.

6. The method as recited by claim 3 , wherein the determining of the subsets of the documents further comprises:

creating a reverse-in-time graph based on the forward-in-time graph.

7. The method as recited by claim 6 , wherein the determining of the subsets of the documents further comprises:

creating an initial matrix of probabilities based on the reverse-in-time graph;

multiplying the initial matrix of probabilities by itself, to create a plurality of matrixes wherein each matrix is considered a generation, until a matrix of probabilities is achieved, wherein the values from the matrix of probabilities is similar to the values from the previous generation matrix; and

determining the subset from the matrix of probabilities.

8. A computer system for determining a subset of documents from which a particular document was derived, the system comprising:

document-similarity-receiver configured for receiving similarity measurements indicating similarities between contents of documents; and

provenance-based-on-document-content-determiner configured for using dates the documents were created and the similarity measurements to determine a subset of the documents that a particular document was derived from, wherein the subset of documents is determined without requiring document tracking information to be associated with the documents.

9. The computer system of claim 8 , wherein the provenance-based-on-document-content-determiner is configured for creating a forward-in-time graph based on the similarity measurements and the dates the documents were created, wherein the forward-in-time graph has path lengths from the particular document to other documents that are less than or equal to a user specified maximum path length value.

10. The computer system of claim 9 , wherein the provenance-based-on-document-content-determiner is configured for creating a reversed-in-time graph based on the forward-in-time graph, wherein one or more similarity values are normalized to determine a probability that one document associated with the reversed-in-time graph resulted in another document associated with the reversed-in-time graph.

11. The computer system of claim 10 , wherein the provenance-based-on-document-content-determiner is configured for determining a Markov chain based on the reversed-in-time graph wherein documents associated with the reversed-in-time graph represent nodes of the Markov chain and probabilities determined based on normalized similarity values are associated with transitions between the nodes.

12. The computer system of claim 11 , wherein the provenance-based-on-document-content-determiner is configured for determining an initial matrix of probabilities based on the Markov chain and multiplying the initial matrix by itself, to create a plurality of matrixes wherein each matrix is considered a generation, until a matrix of probabilities is achieved, wherein the values from the matrix of probabilities is similar to the values from the previous generation matrix.

13. The computer system of claim 12 , wherein documents are represented by the rows and columns of the initial matrix and each probability is associated with a particular row and column pair of the initial matrix.

14. The computer system of claim 12 , wherein the subset of documents is determined based on the matrix of probabilities.

15. A computer-usable storage medium having instruction embodied therein that when executed cause a computer system to perform a method for determining a subset of documents from which a particular document was derived, the method comprising

receiving similarity measurements indicating similarities between contents of documents, wherein the documents are unstructured documents that do have document tracking information; and

using the similarity measurements to determine a subset of the documents that a particular document was derived from whereby analyzing the unstructured documents is enabled.

16. The computer-usable storage medium of claim 15 , wherein the computer-readable program code embodied therein causes a computer system to perform the method, and wherein further comprising:

ranking documents associated with the subset based on how similar each document in the subset is to the particular document.

17. The computer-usable storage medium of claim 15 , wherein the computer-readable program code embodied therein causes a computer system to perform the method, and wherein the determining of the subset of the documents further comprises:

creating a forward-in-time graph based on the similarity measurements and the dates the documents were created.

18. The computer-usable storage medium of claim 17 , wherein the computer-readable program code embodied therein causes a computer system to perform the method, and wherein the creating of the forward-in-time graph further comprises:

creating the forward-in-time graph that has path lengths from the particular document to other documents that are less than or equal to a user specified maximum path length value; and

creating a reverse-in-time graph based on the forward-in-time graph.

19. The computer-usable storage medium of claim 18 , wherein the computer-readable program code embodied therein causes a computer system to perform the method, and wherein the determining of the subsets of the documents further comprises:

creating an initial matrix of probabilities based on the reverse-in-time graph;

multiplying the initial matrix of probabilities by itself, to create a plurality of matrixes wherein each matrix is considered a generation, until a matrix of probabilities is achieved, wherein the values from the matrix of probabilities is similar to the values from the previous generation matrix; and

determining the subset from the matrix of probabilities.

20. The computer-usable storage medium of claim 15 , wherein the computer-readable program code embodied therein causes a computer system to perform the method, and wherein the method further comprises:

determining an order that documents associated with the subset were used to derive the particular document based on how similar each document in the subset is to the particular document.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2022
From: OT PATENT ESCROW, LLC
To: VALTRUS INNOVATIONS LIMITED
Reel/Frame 061244/0298 →
PATENT ASSIGNMENT, SECURITY INTEREST, AND LIEN AGREEMENT Recorded Jan 26, 2021
From: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP; HEWLETT PACKARD ENTERPRISE COMPANY
To: OT PATENT ESCROW, LLC
Reel/Frame 055269/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2015
From: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 037079/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2007
From: DEOLALIKAR, VINAY
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 019607/0660 →