IP Library › Granted Patent US 9,842,103
Granted Patent B1
US 9,842,103 · App. 15/095,488 · Granted Dec 12, 2017

Machined book detection

Inventors: Mitsuo Takaki (Bellevue, WA); Divya Mahalingam (Seattle, WA); David Gordon Leatham (Kirkland, WA); David Rezazadeh Azari (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G06F17/278G06F17/274G06F17/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,842,103
App. No.
15/095,488
Granted
Dec 12, 2017
Kind
B1
Abstract

A system and method for determining whether a textual work submitted for publishing is machine generated or non-machine generated by identifying and quantifying various aspects of the textual work and comparing those aspects to known works. For example, the system and method may identify aspects of a textual work, including, a relationship between the sentences within the textual work, a writing style of the author of the textual work, a grammatical structure of the sentences within the textual work, a quality of the textual work, and other aspects of the textual work. Upon determining that the textual work is machine generated the textual work may be rejected for publishing.

Claims (56)

1. A computer-implemented method, comprising:

storing, in computer memory, a textual work;

generating an N-gram of N words of the textual work;

generating a 2-dimensional plot based, at least in part, on the N-gram;

storing, in computer memory, data corresponding to at least one 2-dimensional representation of at least one pre-determined machine generated work;

processing the 2-dimensional plot and the at least one 2-dimensional representation to obtain a score indicative of a correlation between characteristics of the textual work and characteristics of the at least one pre-determined machine generated work; and

determining, based at least in part on the score, that the textual work comprises at least a portion of machine generated grammatically unintelligible text.

2. The computer-implemented method of claim 1 , further comprising:

storing pre-determined shape descriptors corresponding to the at least one pre-determined machine generated work;

comparing a shape descriptor of the 2-dimensional plot to the pre-determined shape descriptors;

calculating a center of mass and the shape descriptor of the 2-dimensional plot; and

determining further, based at least in part on the comparing, that the textual work is machine generated.

3. The computer-implemented method of claim 2 , in which the comparing includes comparing the center of mass to pre-determined centers of the pre-determined machine generated works.

4. The computer-implemented method of claim 2 , further comprising searching an index of a database of the pre-determined shape descriptors for a pre-determined shape descriptor having substantially a same shape as the shape descriptor of the textual work.

5. The computer-implemented method of claim 1 , in which generating the N-gram includes identifying a number of times each of the N words appears in one or more sentences of the textual work.

6. The computer-implemented method of claim 1 , further comprising:

parsing the textual work; and

identifying at least one of verbs, nouns, pronouns, adjectives, adverbs, prepositions, conjunctions, and interjections in each sentence of the textual work.

7. The computer-implemented method of claim 6 , further comprising:

applying one or more grammatical rules to each sentence of the textual work; and

computing a score for the textual work based on the application of the one or more grammatical rules.

8. A computing device, comprising:

at least one processor;

a memory device including instructions operable to be executed by the at least one processor to perform a set of actions, configuring the at least one processor:

to receive a textual work;

store, in computer memory, the textual work;

to generate an N-gram of N words of the textual work;

to generate a 2-dimensional plot based, at least in part, on the N-gram in the N-dimensional space;

to store, in computer memory, data corresponding to at least one 2-dimensional representation of at least one pre-determined machine generated work;

to process the 2-dimensional plot and the at least one 2-dimensional representation to obtain a score indicative of a correlation between characteristics of the textual work and characteristics of the at least one pre-determined machine generated work; and

to determine, based at least in part on the score, that the textual work comprises at least a portion of machine generated grammatically unintelligible text.

9. The computing device of claim 8 , further comprising configuring the at least one processor to:

store pre-determined shape descriptors corresponding to the at least one pre-determined machine generated work;

compare a shape descriptor of the 2-dimensional plot to the pre-determined shape descriptors;

calculate a center of mass and the shape descriptor of the 2-dimensional plot; and

determine further, based at least in part on the comparing, that the textual work is machine generated.

10. The computing device of claim 9 , in which the processor is configured to compare the center of mass to pre-determined centers of the pre-determined machine generated works.

11. The computing device of claim 9 , further comprising configuring the at least one processor to search an index of a database of the pre-determined shape descriptors for a pre-determined shape descriptor having substantially a same shape as the shape descriptor of the textual work.

12. The computing device of claim 8 , further comprising configuring the at least one processor to generate the N-gram by identifying a number of times each of the N words appears in one or more sentences of the textual work.

13. The computing device of claim 8 , further comprising configuring the at least one processor to:

parse the textual work; and

identify at least one of verbs, nouns, pronouns, adjectives, adverbs, prepositions, conjunctions, and interjections in each sentence of the textual work.

14. The computing device of claim 13 , further comprising configuring the at least one processor to:

apply one or more grammatical rules to each sentence of the textual work; and

compute a score for the textual work based on the application of the one or more grammatical rules.

15. A computer-implemented method of identifying machine generated text, comprising:

receiving a textual work;

storing, in computer memory, the textual work;

parsing a portion of the textual work and identifying at least one of verbs, nouns, pronouns, adjectives, adverbs, prepositions, conjunctions, and interjections in each sentence of the portion;

representing, in the computer memory, the parsed portion of the textual work as a 2-dimensional representation of the identified at least one of verbs, nouns, pronouns, adjectives, adverbs, prepositions, conjunctions, and interjections;

applying one or more grammatical rules to the 2-dimensional representation;

calculating a confidence score corresponding to the application of the one or more grammatical rules; and

determining, based at least in part on the confidence score, that the textual work comprises at least a portion of machine generated grammatically unintelligible text.

16. The computer-implemented method of claim 15 , in which the calculating of the confidence score includes assigning a score relating to grammatical correctness for each sentence of the portion based on application of the one or more grammatical rules.

17. The computer-implemented method of claim 16 , in which the calculating includes computing an overall score for the portion based one or more sentence scores, wherein the overall score is a combination of the score for each of the sentences.

18. The computer-implemented method of claim 15 , in which the applying of the one or more grammatical rules includes applying one or more of an improper verb clustering rule, an improper verb tense rule, and improper noun clustering rule, an improper sentence ending rule, and a foreign word occurrence rule.

Continuity (1)
Continuation 13720195 · Dec 19, 2012