IP Library Granted Patent US 9,189,482
Granted Patent B2
US 9,189,482 · App. 13/672,064 · Granted Nov 17, 2015

Similar document search

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,189,482
App. No.
13/672,064
Granted
Nov 17, 2015
Kind
B2
Abstract

Described herein are methods for finding substantially similar/different sources (files and documents), and estimating similarity or difference between given sources. Similarity and difference may be found across a variety of formats. Sources may be in one or more languages such that similarity and difference may be found across any number and types of languages. A variety of characteristics may be used to arrive at an overall measure of similarity or difference including determining or identifying syntactic roles, semantic roles and semantic classes in reference to sources.

Claims (58)

1. A method for comparing documents, the method comprising:

associating, by a processor, a respective weight with each of a plurality of information types including text-based information, graphical information, audio information, or video information;

identifying for each of the documents, by the processor, one or more segments each corresponding to one of the plurality of information types; and

estimating, by the processor, a similarity value between a first document of the documents and a second document of the documents by comparing each segment of the first document with a segment of the second document that corresponds to a same information type and combining results of the comparison based on the respective associated weights.

2. The method of claim 1 , wherein the method further comprises:

identifying a set of source sentences in each of the documents; and

constructing a language-independent semantic structure (LISS) for each source sentence of the identified set of source sentences, wherein said estimating the similarity value between the first and the second documents includes a comparison of the LISS's of the first document with the LISS's of the second document.

3. The method of claim 1 , wherein the method further comprises:

performing optical character recognition (OCR) on the one or more segments that include text in the form of graphical information.

4. The method of claim 2 , wherein the method further comprises:

displaying through a user interface a portion of the first document and a portion of the second document;

identifying fragments of the first and the second documents that are related to said similarity value; and

aligning in the user interface the identified fragments with respect to each other and with respect to the user interface.

5. The method of claim 4 , wherein the method further comprises:

aligning identified fragments of the first document and the second document based on a comparison of the LISS's of sentences of the identified fragments of the first and the second documents.

6. The method of claim 1 , wherein the method further comprises:

displaying the estimated similarity value through a user interface.

7. The method of claim 6 , wherein said displaying the similarity value includes:

identifying, through the user interface, the first document and the second document if their similarity value exceeds a similarity threshold value; or

identifying, through the user interface, the first document and the second document if their similarity value fails to exceed the similarity threshold value.

8. The method of claim 6 , wherein said displaying the estimated similarity value between the first and second documents includes:

highlighting, underlying, or changing a font of respective pairs of similar portions of text-based information in the first and the second documents, wherein each of the respective pair of similar portions of text-based information includes a portion in the first document and a portion in the second document that are associated with a respective similarity value that exceeds a similarity threshold value or fails to exceed the similarity threshold value.

9. The method of claim 2 , wherein the method further comprises generating semantic class tokens of the sets of sources sentences, and wherein the similarity value is based on comparing the semantic classes of tokens of the set of source sentences of the first with the semantic class tokens of the set of source sentences of the second documents.

10. The method of claim 2 , wherein the method further comprises:

indexing each source sentence of said set of source sentences.

11. The method of claim 2 , wherein the method further comprises:

indexing the LISS's for each of said set of source sentences.

12. The method of claim 10 , wherein said estimating a similarity value between the first and second documents is based on comparing the indexes of fragments of text-based information.

13. The method of claim 2 , wherein the set of source sentences of the first document are in a different language than the set of source sentences of the second document.

14. The method of claim 2 , wherein the method further comprises:

indexing semantic classes of said set of source sentences, and wherein said estimating the similarity value is based on comparing the indexes of semantic classes of said set of source sentences.

15. The method of claim 11 , wherein said estimating the similarity value is based on comparing indexes of the language-independent semantic structures (LISS's) of the sets of source sentences.

16. The method of claim 15 , wherein the indexes of the language-independent semantic structures (LISS's) of the set of source sentences include indexes of lexical features.

17. The method of claim 15 , wherein the indexes of the language-independent semantic structures (LISS's) of the set of source sentences include indexes of grammatical features.

18. The method of claim 15 , wherein the indexes of the language-independent semantic structures (LISS's) of the set of source sentences includes indexes of syntactical features.

19. The method of claim 15 , wherein the indexes of the language-independent semantic structures (LISS's) of the set of source sentences includes indexes of semantic features.

20. The method of claim 3 , wherein the method further comprises:

converting each of the segments that correspond to non-text-based information to one or more text-based source sentences thereby creating a text-based equivalent block for each of the segments that correspond to non-text-based information;

identifying source sentences in each of the segments that correspond to text-based information; and

generating for each text-based segment or text-based equivalent block a language-independent semantic structure (LISS) of said source sentences.

21. The method of claim 1 , wherein the method further comprises:

determining a logical structure of each of the documents, and wherein said estimating the similarity value is further based on said logical structures.

22. An electronic device for comparing documents, the electronic device comprising:

a processor;

a display in electronic communication with the processor; and

a memory in electronic communication with the processor and the display, the memory configured with instructions to perform a method by the processor, the method including:

associating a respective weight with each of a plurality of information types including text-based information, graphical information, audio information, or video information;

identifying for each of the documents one or more segments each corresponding to one of the plurality of information types; and

estimating a similarity value between a first document of the documents and a second document of the documents by comparing each segment of the first document with a segment of the second document that correspond to a same information type and combining results of the comparison based on the respective associated weights.

23. The electronic device of claim 22 , wherein the method further comprises:

identifying a set of source sentences in each of the documents; and

constructing a language-independent semantic structure (LISS) for each source sentence of the identified set of source sentences, wherein said estimating the similarity value between the first and second documents includes a comparison of the LISS's of the first document with the LISS's of the second document.

24. The electronic device of claim 22 , wherein the method further comprises:

performing optical character recognition (OCR) on the one or more segments that include text in the form of graphical information.

25. The electronic device of claim 23 , wherein the method further comprises:

displaying through a user interface a portion of the first document and a portion of the second document;

identifying fragments of the first and the second documents that are related to said similarity value; and

aligning in the user interface the identified fragments with respect to each other and with respect to the user interface.

Assignments (5)
SECURITY INTEREST Recorded Aug 14, 2023
From: ABBYY INC.; ABBYY USA SOFTWARE HOUSE INC.; ABBYY DEVELOPMENT INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 064730/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: ABBYY PRODUCTION LLC
To: ABBYY DEVELOPMENT INC.
Reel/Frame 059249/0873 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR DOC. DATE PREVIOUSLY RECORDED AT REEL: 042706 FRAME: 0279. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 25, 2017
From: ABBYY INFOPOISK LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 043676/0232 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2017
From: ABBYY INFOPOISK LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 042706/0279 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2012
From: DANIELYAN, TATIANA; ZUEV, KONSTANTIN
To: ABBYY INFOPOISK LLC
Reel/Frame 029440/0731 →