CROWD-SOURCED EXCLUSION OF SMALL MATCHES IN DIGITAL SIMILARITY DETECTION
The present invention relates to systems that search documents and highlight occurrences of text found in previously published documents, publications, Internet websites and electronic documents. In particular, the present invention relates to originality assessment of a variety of documents (e.g., student papers, college admissions essays, PhD theses, magazines, newspapers, and book publications).
1 . A system for document analysis, comprising a processor and software configured a) generate a anti-source mask of a submitted original work by removing undesired match text from said submitted original work, and b) generate a similarity report of said submitted original work by identifying text in a match sources text found in said submitted original work.
2 . The system of claim 1 , wherein said undesired match text is stored and retrieved as a hash or as individual strings of text.
3 . The system of claim 2 , wherein said software is further configured to generate a text exclusion hash of removed text by the steps of a) receiving a plurality of undesired match text submitted by users; and b) generating a text exclusion hash of undesired matches from said plurality of undesired match text.
4 . The system of claim 1 , wherein said submitted original work is selected from the group consisting of student papers, college admissions essays, PhD theses, magazines, newspapers, book publications and software code.
5 . The system of claim 1 , wherein said system further comprises a processor and software configured to facilitate review or mark-up of said original work.
6 . The system of claim 1 , wherein said plurality of undesired match text comprises 50 or more text sections.
7 . The system of claim 1 , wherein said plurality of undesired match text comprises 1000 or more text sections.
8 . The system of claim 1 , wherein said plurality of undesired match text comprises 10,000 or more text sections.
9 . The system of claim 3 , wherein said software is configured for updating said text exclusion hash with new undesired match text.
10 . The system of claim 1 , wherein said system is further configured to display said similarity report.
11 . A method for document analysis, comprising:
a) generating an anti-source mask of a submitted original work by removing undesired match text from said submitted original work; and
b) generating a similarity report of said submitted original work by identifying text in a match sources text found in said submitted original work.
12 . The system of claim 11 , wherein said undesired match text is stored and retrieved as a hash or as individual strings of text.
13 . The method of claim 12 , further comprising the step of generate a text exclusion hash of said removed text by a) inputting a plurality of undesired match texts from users into a computer processor comprising computer software; and b) generating a text exclusion hash from said plurality of undesired match text.
14 . The method of claim 11 , wherein said submitted original work is selected from the group consisting of student papers, college admissions essays, PhD theses, magazines, newspapers, book publications and software code.
15 . The method of claim 11 , wherein said method further comprises review or mark-up of said original work.
16 . The method of claim 11 , wherein said plurality of undesired match text comprises 50 or more text sections.
17 . The method of claim 11 , wherein said plurality of undesired match text comprises 1000 or more text sections.
18 . The method of claim 11 , wherein said plurality of undesired match text comprises 10,000 or more text sections.
19 . The method of claim 12 , further comprising the step of updating said text exclusion hash with new undesired match text.
20 . The method of claim 11 , further comprising the step of displaying said similarity report.